Efficient binding free energy calculations for diverse chemical space in drug discovery
- Caldaruse, Ana-Maria
- Advisor(s): Mobley, David L.
Abstract
Small-molecule drug discovery often begins with a search through large numbers of candidate molecules. A key concern is finding the ones that bind a disease-related protein tightly enough to alter its biological function. Because evaluating each candidate experimentally is slow and costly, computational methods that predict binding strength have become central to modern drug discovery. Fast approximate methods can rank very large libraries, but their predictions are unreliable. Physics-based binding free energy calculations are far more accurate. They are also computationally expensive, which limits how many compounds they can evaluate. To manage that cost, pharmaceutical research relies heavily on relative binding free energy calculations. These compare each molecule to a closely related one, so they work only when the two molecules are structurally similar. That requirement has largely confined them to the later stage of discovery, where an already promising compound is refined. It has left out earlier stages, where candidate compounds are small, weakly binding, and structurally diverse.This dissertation applies and extends a computational method, Separated Topologies, that removes this restriction and enables accurate binding predictions between structurally dissimilar molecules. I first demonstrate that the method can reliably rank small molecular fragments, including fragments that occupy entirely different regions of a protein. I then integrate it with active learning, a machine learning strategy that decides which compounds merit the expensive calculation. This lets me search a large and diverse compound library while running the free energy calculations on only a small fraction of it. Finally, I couple the method with a second, more efficient technique called multi-site $\lambda$-dynamics so that the overall approach makes the best use of limited computational resources. Collectively, this work extends accurate, physics-based screening to a stage of drug discovery where it was previously impractical. The broader aim is to find promising drug candidates faster and at lower cost.