Skip to main content
eScholarship
Open Access Publications from the University of California

What Makes an LLM Response Read as Empathic? A Multi-Dimensional Study of Empathic Alignment

Creative Commons 'BY' version 4.0 license
Abstract

As large language models (LLMs) are increasingly deployed in emotional support and counseling-style dialogue, what makes their responses read as empathic remains an open question. Drawing on prior work in empathic communication, we examine four communicative dimensions of perceived empathy in LLM-generated responses: specificity, emotional reflection, affective word choice, and diversity. Across five instruction-tuned LLMs, emotional reflection (explicitly acknowledging and mirroring feelings) was the primary bottleneck across these metrics, while the other three dimensions clustered near ceiling. A preference-based learning approach that targeted reflection improved metric-based empathic quality on the four-dimensional composite. However, human raters showed no reliable preference between baseline and DPO-optimized responses, suggesting that these metric-based gains preserved, rather than enhanced, perceived response quality. We read this gap as a scope characterization: this metric set captures optimizable aspects of empathic response generation, but does not exhaust the cues humans use when judging empathy.