- Main
Evaluating Offline Reinforcement Learning for Vasopressor Management in the ICU
- Chun, Allen
- Advisor(s): Zhu, Yuhua
Abstract
This thesis presents an analysis of offline reinforcement learning (RL) methods for vasopressor management in the intensive care unit (ICU) using retrospective clinical data. Vasopressor administration is defined as a sequential decision-making problem, and the problem is formulated as a finite-horizon Markov Decision Process (MDP). The research examines three different approaches: Behavior Cloning (BC), Conservative Q-Learning (CQL), and Batch-Constrained Q-Learning (BCQ). Policy behavior is evaluated using mean arterial pressure (MAP). All three techniques yield clinically plausible policies that increase the usage of vasopressors in the presence of lower MAP values and reduce their usage in the presence of higher MAP values. BC approximates clinician behavior fairly accurately, CQL offers comparable performance but with more regularization to prevent overestimation, and BCQ yields systematic and conservative divergences. In addition to this, the study explores off-policy evaluation through the application of Weighted Importance Sampling (WIS) and Fitted Q-Evaluation (FQE), which further reveal substantial variability and overlapping confidence intervals across methods. These findings suggest that, while empirical evaluation remains challenging, offline reinforcement learning provides a solid foundation for future exploration in clinical decision support systems.