- Main
The Grunt Work of Conversational Grounding: Minimal Responses in Multimodal Dialogue
Abstract
Conversational grounding, the establishment of mutual understanding between participants, is fundamental to studies of dialogue. However, methods of grounding multimodal information remain understudied. We investigate how minimal vocalizations (conversational grunts like hm and ooh) encode grounding of verbal information versus observed events in collaborative tasks. Acoustic analysis of 497 grunts from natural dialogues revealed that token selection is strongly specialized by information source. Backchannels like mm-hm predominantly acknowledge speech, while reactive tokens like oh preferentially responded to events. Event-following grunts also exhibit longer duration and more open vowel production at the corpus level, though these differences are largely a consequence of which tokens speakers select rather than independent prosodic modulation. These findings establish observable events as informational contributions requiring explicit acknowledgment, and identify lexical selection as the primary mechanism by which speakers differentiate grounding sources in co-situated task interaction.