Skip to main content
eScholarship
Open Access Publications from the University of California

Cognitively Grounded Benchmark Generation for Robust Compositional Reasoning in Large Language Models

Creative Commons 'BY' version 4.0 license
Abstract

Large language models (LLMs) can write fluent mathematics, yet their reasoning collapses when proofs require verifiable composition across many dependent steps. We propose a cognitively grounded protocol for generating robust benchmark datasets in advanced mathematics by operationalizing seven human-inspired dimensions of expert practice: concept formation, dualization, negative knowledge, transfer, invariance control, lemma synthesis, and counterexample search. Rather than toy tasks, we construct meta-prompts from research-grade problem fragments and evaluate complete solution traces, auditing faithfulness and invariant preservation step by step. Under matched conditions, four state-of-the-art systems show a global breaking degree of nore than 90% on stress tests, with failures concentrated in long-horizon planning, premise selection, lemma construction, and counterexample discovery. We conclude with dataset design principles that maximize diagnostic power and reproducibility, and with training levers—rationale SFT, process supervision with reward models, and stepwise preference learning—that directly target step-level correctness.