An Explorer’s Toolkit for Traversing Chemical Space: Making Decisions under the Weight of Uncertainty in Drug Discovery
- Nandkeolyar, Aakankschir
- Advisor(s): Mobley, David L
Abstract
Drug discovery proceeds in stages, and each stage asks a different question ofcomputation. Early on, the task is to find promising starting points in a space far too large to examine exhaustively, and the methods used are fast and only loosely predictive. Later, once a chemical series is established, the question narrows to how much better one analogue is than another, and answering it demands a physical estimate at far greater cost. Beneath both sits the question of whether the molecule being modelled is the species that actually exists in solution. This dissertation develops computational methods at each of these stages. Chapters 2 and 3 address hit identification in make-on-demand combinatoriallibraries, which now enumerate billions of products and cannot be scored exhaustively. Both introduce TACTICS, an adaptive Thompson Sampling framework that treats each reagent as an arm of a multi-armed bandit and, unlike existing methods that explore on fixed schedules, measures at every cycle how well each reaction component's reagents can be distinguished and routes the remaining budget toward the components that remain ambiguous. Chapter 2 evaluates the approach on ligand-based screening against an established twenty-library benchmark; Chapter 3 carries it to structure-based docking and builds the systematic benchmark that did not previously exist, spanning four reaction chemistries, two docking protocols, and three protein families. In both settings the adaptive methods recover more top-scoring compounds than the established ones, with the advantage widening as the evaluation budget tightens. Chapter 4 turns to lead optimization, where relative binding free energycalculations are organized into perturbation networks whose thermodynamic inconsistencies reveal error. It replaces the standard maximum likelihood correction, which returns a single number and redistributes error across well-converged and poorly converged edges alike, with Bayesian estimators that return a posterior over per-ligand free energies and can identify unreliable edges rather than averaging over them. Chapter 5 examines the physical properties on which the preceding chaptersdepend, reporting the SAMPL8 blind challenge assessment of community predictions of pKa and logD for multiprotic compounds. It shows that a measured pKa does not reveal which protonation transition produced it, so two methods can agree closely on a number while disagreeing entirely about the chemistry, and that the ranking of methods depends on how predictions are paired with experiment.