- Main
Human-Centered Evaluation of Ambient Artificial Intelligence Clinical Documentation
- Guo, Yawen
- Advisor(s): Zheng, Kai
Abstract
Ambient artificial intelligence (AI) documentation tools convert patient–clinician conversations into draft clinical notes that clinicians review and modify before signing in the electronic health record (EHR). These systems are increasingly adopted as a response to documentation burden and clinician burnout, but real-world effectiveness depends on the amount of clinician revision required for AI drafts and on persistent barriers that limit uptake and sustained use. This dissertation presents a human-centered, multi-level evaluation of ambient AI documentation conducted within the University of California, Irvine Health (UCI Health).First, the implementation of ambient AI documentation was evaluated using a mixed-methods approach that paired EHR time-efficiency measures with pre- and post-implementation survey responses from outpatient clinicians at UCI Health across multiple specialties. Results indicated that time spent on documentation decreased, while note length increased during the three-month period following implementation. Survey findings aligned with these trends, reflecting improved perceived efficiency and reduced documentation burden.Second, clinicians’ edits of ambient AI drafts were quantified at scale using AI draft and clinician-finalized note pairs. Across notes that contained one or more ambient AI–generated note sections, 84.4% were revised by clinicians before signing. Editing intensity was highest in the Assessment and Plan section, while the tendency to accept drafts without edits varied primarily across individual clinicians rather than across specialties.Third, a qualitative content analysis was conducted using a manually annotated corpus of sentence-level edits derived from draft and final note pairs. The analysis identified recurring patterns in clinicians’ revisions to AI drafts, including adding specialty-specific details, calibrating certainty to reflect available evidence and diagnostic framing, and converting transcript-like narratives into professional documentation. Using this annotated dataset, few-shot prompting of large language models (LLMs) was evaluated to assess the feasibility of scalable automated edit measurement across five target categories. Results indicated that LLM-assisted classification was most reliable when the revision could be inferred from the local before-and-after text alone and less consistent when accurate classification required broader clinical context.Finally, semi-structured interviews with clinicians who use ambient AI in routine practice were conducted to understand the rationale behind editing AI-drafted notes. Clinicians described editing as a deliberate process to ensure accuracy, align drafts with clinical reasoning and documentation standards, and meet billing and medico-legal requirements. Interviews further underscored the need for specialty-aware customization and workflow support for efficient draft review and verification. Recommendations emphasized system-level optimization across model behavior, individual personalization, workflow design, and EHR integration.Collectively, these findings highlight that effective ambient AI implementation must align with clinical documentation standards, support specialty-specific workflows, and enable clinicians to efficiently verify, refine, and structure notes. By clarifying how ambient AI is used in routine practice, what revisions clinicians make, and why those revisions occur, this dissertation provides practical evidence and guidance to inform implementation, continuous monitoring, and iterative design improvements, ultimately aiming to reduce documentation burden while maintaining high-quality clinical records.