Interpretable and robust statistical learning for clinical decisions and health-system design
- Jiang, Muyan
- Advisor(s): Aswani, Anil
Abstract
Healthcare analytics increasingly supports communication audits, individualized treatment, payment reform, and real-time quality and safety monitoring. Across these settings, methods must be interpretable so stakeholders can understand and act on model outputs, and robust so conclusions remain reliable under heterogeneity, dependence, confounding, and strategic behavior. This dissertation develops a coherent set of statistical learning tools built around a common principle: introducing explicit structure that turns complex data into transparent, auditable decision objects and supports valid inference and monitoring in real-world healthcare settings. The dissertation is arranged to mirror a lifecycle for trustworthy deployment. First, I develop an aspect-based NLP measurement framework for pediatric progress notes that quantifies family-centered communication and enables evaluation of note-sharing policies such as OpenNotes, with inference procedures designed for dependent comparisons. Second, I propose an interpretable semiparametric model for personalized rules with continuous treatments and binary outcomes. The method separates patient-specific risk from treatment-effect heterogeneity into two clinically legible scores, linked by a flexible nonparametric component that yields visualizable dose-response surfaces and actionable individualized recommendations. Third, I study end-of-life payment design through a principal–agent formulation and derive incentive-compatible mixed contracts combining fee-for-service and pay-for-performance, with transparent parameters that clarify how optimal reimbursement depends on patient heterogeneity and provider risk preferences. Finally, I develop robust monitoring and detection methods for deployed dynamical systems under drift or adversarial manipulation by combining physically consistent observer design with dependence-aware statistical testing, providing a general approach to trustworthy detection when standard independence assumptions fail. Together, these contributions unify measurement, individualized decision-making, incentive design, and monitoring within a single methodological agenda for interpretable and robust learning in complex healthcare systems.