Skip to main content
eScholarship
Open Access Publications from the University of California

Argument Evaluation Strategies in Human and Machine Reasoning: LLMs Resemble Sound Deliberate (But Not Intuitive) Thinkers

Creative Commons 'BY' version 4.0 license
Abstract

Humans routinely evaluate arguments to decide which reasons warrant belief or action. Large language models (LLMs) are increasingly used as evaluators, yet little is known about how their argument evaluation compares to human evaluation. In this study we examined how humans and LLMs evaluate arguments supporting answers to classic reasoning problems. Human participants rated arguments either under time pressure and cognitive load or without constraints, allowing us to situate LLM evaluations relative to intuitive and deliberate human judgments. Arguments varied in response correctness and justification type. Modeling argument ratings as a function of argument validity, surface explicitness, and alignment with the evaluator's own response revealed that intuitive human evaluations were dominated by belief-consistency bias, whereas deliberate evaluations –especially among correct responders– prioritized justification validity. In contrast, LLMs generally showed stable, validity-centered evaluative strategies, closely resembling deliberate human evaluations from correct responders and diverging from intuitive human judgments.