Skip to main content
eScholarship
Open Access Publications from the University of California

Hearing Speech or Doing Inference? Diagnosing Speech Perception in Speech Large Language Models

Creative Commons 'BY' version 4.0 license
Abstract

Speech LLMs with native audio input can produce fluent text from raw acoustic signals, but it remains unclear whether this reflects speech perception as such. I use the sine-wave speech paradigm to probe responses across free description, forced-choice discrimination with silent-audio controls, and open-ended transcription without alternatives. A qualitative split emerges across models in behavior. One class shows little sensitivity to the sine-wave speech, failing to use the signal in free description and transcription while maintaining high forced-choice accuracy even under silence. A second class shows limited stimulus sensitivity, with instruction-dependent improvements that disappear when acoustic input is removed, resembling a weak analogue of human perceptual reorganization in part. I suggest that these differences track access to speech-relevant acoustic information at the front-end rather than downstream language modeling.