Dissertation
Enhancing nucleotide sequence representation for functional, taxonomy and phenotype prediction
Doctor of Philosophy (Ph.D.), Drexel University
May 2026
DOI:
https://doi.org/10.17918/00011403
Abstract
Modern sequencing technologies generate vast amounts of nucleotide sequence data across genomics, metagenomics, and biotechnology. However, the central challenge has shifted from data generation to biological interpretation, particularly when sequences are fragmented, noisy, weakly labeled, or derived from organisms poorly represented in reference databases. Conventional pipelines based on alignment, k-mer matching, or task-specific feature engineering remain essential, but they often struggle under novelty, distribution shift, and incomplete sequence context. This dissertation addresses these challenges by developing foundation-model-based methods for learning transferable nucleotide sequence representations that support functional, taxonomic, and phenotype-related prediction. The dissertation introduces a progression of representation-learning systems for biological sequences. First, MetaBERTa is developed as a fragment-scale nucleotide foundation model pretrained on diverse prokaryotic genomes to support taxonomic classification and sequence-level representation learning. Second, Carmania extends nucleotide modeling to long genomic contexts using transition-regularized self-supervision, improving the ability to capture sequence organization beyond short fragments. Third, Scorpio introduces hierarchical contrastive adaptation to reshape embedding geometry across gene and taxonomic levels, enabling retrieval, classification, confidence estimation, and biological interpretation across full genes, short fragments, simulated reads, promoter sequences, and antimicrobial resistance tasks. Finally, Nevermore extends this framework toward multimodal biological modeling by integrating protein and ligand representations for target-conditioned, database-grounded molecular optimization. Together, these contributions advance nucleotide sequence modeling from isolated supervised pipelines toward reusable biological representation systems. The results demonstrate that learned embeddings capture meaningful biological structure, improve downstream prediction, support interpretation under uncertainty, and provide a foundation for future multimodal models connecting genomic sequences with proteins, molecules, and phenotype-related objectives. This work contributes to scalable and interpretable biological artificial intelligence systems for genomics, metagenomics, and computational discovery.
Metrics
1 Record Views
Details
- Title
- Enhancing nucleotide sequence representation for functional, taxonomy and phenotype prediction
- Creators
- Mohammad Saleh Refahi
- Contributors
- Gail L. Rosen (Advisor)
- Awarding Institution
- Drexel University
- Degree Awarded
- Doctor of Philosophy (Ph.D.)
- Publisher
- Drexel University
- Number of pages
- xxi, 146 pages
- Resource Type
- Dissertation
- Language
- English
- Academic Unit
- College of Engineering (1970-2026); Electrical (and Computer) Engineering (1970-2026); Drexel University
- Other Identifier
- 991022188673804721