- Main
Computer-Aided Drug Discovery: Generative Molecular Design with Large Language Models and Structure-Based Benchmarking
- Sun, Kunyang
- Advisor(s): Head-Gordon, Teresa
Abstract
The discovery of new small-molecule therapeutics remains a notoriously time-consuming and capital-intensive endeavor, constrained by the vastness of synthesizable chemical space and the limitations of traditional screening approaches. Generative artificial intelligence has fundamentally reshaped this landscape by enabling the direct sampling of molecules with desired properties, yet persistent challenges in synthesizability, data quality, and translational validation continue to limit the practical impact of computational methods. This dissertation addresses these challenges across three interconnected pillars: Large Language Model (LLM)-based generative molecular design, rigorous structure-based benchmarking, and prospective validation in competitive drug discovery settings.In Chapters 2 and 3, I develop two LLM-driven generative frameworks that address core bottlenecks in molecular design. SynLlama, described in Chapter 2, tackles the synthesizability problem by fine-tuning an open-weight LLM to generate synthetic pathways that decompose arbitrary molecules into commercially available building blocks assembled via known and robust chemical reactions. By constraining generation to this synthesizable chemical space, every proposed candidate comes with an actionable synthesis route, bridging the gap between in silico design and practical wet-lab execution. In Chapter 3, I introduce LinkLlama, a complementary framework that reframes linker design as a conditional generation task to solve a persistent bottleneck in fragment-based drug discovery. By leveraging natural language conditioning rather than explicit geometric training objectives, LinkLlama produces chemically reasonable linkers with high spatial fidelity across diverse scales, from simple fragment joining and scaffold hopping to the demanding design of linkers for proteolysis targeting chimera (PROTAC). Together, these models form a versatile LLM-based generative suite where natural language prompts enables flexible multi-objective molecular design with natural language guidance.In Chapters 4 and 5, I establish reliable structural foundations for training and evaluating computational drug discovery methods. Chapter 4 introduces HiQBind, a comprehensive open-source workflow that systematically identifies and corrects structural errors in protein-ligand complexes. By integrating tools such as OpenMM and PDBFixer with native annotations from the PDB, the pipeline produces over 30,000 high-quality, non-covalent protein-ligand structures with accurately assigned bond orders, adjusted protonation states, and refined protein structures, directly resolving long-standing community concerns regarding the reproducibility of proprietary curation methods. In Chapter 5, I introduce Kin-ConfBench, a conformation-aware benchmark that evaluates cofolding models based on their ability to capture the conformational landscape of human kinases. KinConfBench reveals a critical tension in current models that, while they often achieve strong pocket-centric metrics, they frequently generate structures that adopt incorrect regulatory conformations for a given ligand, exposing a pervasive issue of apo-state memorization. By establishing a protein-centric benchmark beyond pocket and ligand geometries, KinConfBench provides an explicit, clinically relevant standard for structure prediction tools. Together, both works lay the groundwork for the community to move forward with more robust and physically-correct model training and evaluations.In Chapter 6, I describe my participation in the ASAP-Polaris-OpenADMET Antiviral Drug Discovery Challenge, a prospective, blinded community challenge that tasks participants with predicting crystal binding poses for SARS-CoV-2 and MERS-CoV Main Protease ligands. Rather than deploying deep learning models, I investigate to what extent classical docking with AutoDock Vina can be enhanced through fragment-guided sampling and scoring. This approach ranks 6th overall and 2nd among non-deep-learning methods, demonstrating that traditional docking paired with careful scoring refinement remains competitive with modern machine learning approaches. Furthermore, the results reveal that molecular docking retains a distinct advantage as a robust hypothesis generator due to its ability to freely sample diverse poses, whereas current cofolding models often exhibit constrained sampling capabilities, suggesting that the community can still greatly benefit from classical methods even as the field shifts toward learned structure prediction architectures.Collectively, these efforts demonstrate that the advanced reasoning and sampling capabilities of Large Language Models, combined with rigorous structural benchmarking and prospective experimental validation, can produce practical, competitive tools for real-world drug discovery. By ensuring synthesizability at the point of generation, enforcing structural integrity and conformational correctness in evaluation datasets, and validating predictions in blinded community settings, this work establishes a coherent framework that bridges the critical gap between computational molecular design and successful translational outcomes.