Skip to main content
eScholarship
Open Access Publications from the University of California

Comparing LLM and Human Responses to Human- and AI-labeled Partners During Naturalistic Conversation

Creative Commons 'BY' version 4.0 license
Abstract

Large language models (LLMs) increasingly interact with humans and other AI systems, raising questions about whether they adjust behavior based on partner identity. In a companion study, humans showed behavioral differences when conversing with partners labeled as human versus AI. Here, we extend this investigation to LLMs. We simulated 2,000 conversations where GPT-3.5-turbo "participants" (N = 50) engaged with partners labeled as either human or AI. All partners were actually identical LLMs, isolating label effects from actual differences. We analyzed transcripts using linguistic measures paralleling human data. LLMs showed robust label effects: more questions, politeness, interpersonal discourse markers, and mental state language with human-labeled partners; more words and positive sentiment with AI-labeled partners. Hedging showed category-specific patterns. Only the discourse marker "like" showed similar patterns across the LLM and human studies. These divergent patterns suggest LLMs have learned partner-type associations, but their social behavior differs fundamentally from humans.