Skip to main content
eScholarship
Open Access Publications from the University of California

UC San Diego

UC San Diego Electronic Theses and Dissertations bannerUC San Diego

Integrating Multi-Omics Knowledge with Metabolic Modeling: Enabling Interoperability Across Computational Biology Tools

Abstract

Cells continually rebalance limited transcriptional and proteomic resources to survive and grow across environments, but the analytical (omics-scale) and mechanistic (model-based) tools that explain this balance often operate in isolation. This dissertation bridges that divide by integrating knowledge-enriched multi-omics data analytics with proteome-constrained genome-scale modeling to enable interoperable discovery and prediction. First, we show that the bacterial proteome is modularized in a manner that mirrors transcriptomic modules: matched proteome and transcriptome module components share gene content and condition-dependent activities, revealing how transcriptional and post-translational regulation shape proteome composition. Crucially, these modules enable inference of absolute proteome allocation directly from transcriptomic signals, quantifying resource trade-offs without requiring proteomics in every condition. Second, we integrate experimental data with Metabolism and Expression (ME) models to dissect respiratory plasticity in engineered and evolved electron transport system variants, identifying an “aerobicity stimulon” that captures coordinated shifts in energy metabolism and proteome allocation across aerobic states. Third, we apply this interoperable workflow across diverse case studies, oxidative stress, high-cell-density physiology, extracellular respiration, and hierarchical substrate choice, to quantify condition-specific proteome reallocations and connect mutations, regulation, and phenotype. Finally, we encapsulate these methods in COBRAme.org, the first fully web-based platform for proteome-constrained ME modeling, which separates a lightweight user interface from scalable cloud workers to eliminate software barriers and standardize reproducible analyses. Collectively, the results demonstrate that experimental data and proteome sector derived constraints measurably improve mechanistic predictions, advancing a practical path from big-data signal discovery to explanatory, predictive models. This work provides both conceptual and software infrastructure for interoperable, knowledge-enriched modeling that connects sequencing-scale measurements to genome-scale physiology.