Deep-learning approaches to predicting molecular phenotypes using sequence-to-function models
Skip to main content
eScholarship
Open Access Publications from the University of California

UC Berkeley

UC Berkeley Electronic Theses and Dissertations bannerUC Berkeley

Deep-learning approaches to predicting molecular phenotypes using sequence-to-function models

Abstract

Noncoding sequences coordinate gene regulation through transcription factor binding and haplotype-specific expression, but translating regulatory sequence into quantitative predictions at the level of individual molecules and individual people remains challenging. This dissertation addresses two aspects of that problem using genomic deep learning, with a shared focus on how data representation shapes model behavior. The first project asks whether representing DNA as both nucleotide sequence and local structure changes what a transcription-factor binding model can learn and explain. In Chapter 2, I develop DeepShape, a convolutional neural network that augments one-hot DNA sequence with five DNA structural attributes, including minor groove width, propeller twist, helical twist, roll, and electrostatic potential, to predict transcription factor (TF) binding across 919 regulatory targets in 148 cell types. Shape features contribute prediction accuracy independently of sequence motifs, and DeepLIFT-based attribution applied to shape inputs recovers structural binding signatures not apparent from sequence alone, providing a new axis of interpretability for genomic deep learning models. The second project asks how representing expression data at haplotype resolution changes sequence-to-expression modeling. In Chapter 3, I fine-tune the Enformer model for allele-specific expression (ASE) prediction on phased allelic counts from the Genotype-Tissue Expression project (GTEx) whole blood, obtained using phASER (phasing and allele-specific expression from RNA sequencing), with raw-count targets and mask-based uninformative-donor filtering as design choices motivated by the structure of phased ASE data. The trained model captures structure in allelic-imbalance magnitude on held-out genes, but does not recover the direction of imbalance. These results characterize the reach and current limits of Enformer-based ASE fine-tuning and motivate next-generation architectures that make haplotype-specific variation more explicit. In Chapter 1, I establish the conceptual framework for both contributions. I review how regulatory sequences shape gene expression, summarize the role of DNA shape readout in transcription factor binding, and trace the development of sequence-to-expression models from early convolutional architectures to recent long-range models such as Enformer, Borzoi, and AlphaGenome. I discuss the cross-individual evaluation gap that has emerged as these models have been applied to personal genome prediction, recent attempts to close it through fine-tuning on personal expression data and through allele-specific learning, and the modality gap between chromatin and expression prediction that frames the limitations Chapter 3 confronts directly. In summary, my work characterizes how the choice of input representation, including the addition of DNA shape channels and the use of phased diploid haplotype inputs, changes what genomic deep learning models can learn and explain about regulatory sequence. These contributions clarify both the reach and the current limits of sequence-to-function modeling for transcription factor binding and haplotype-resolved expression, and motivate next-generation architectures that make structural and haplotype-specific variation more explicit.