Skip to main content
eScholarship
Open Access Publications from the University of California

When to Think Deep? Resource-Rational Metacognitive Control for Adaptive Inference in LLM-Based Legal QA

Creative Commons 'BY' version 4.0 license
Abstract

Legal question-answering systems based on large language models (LLMs) encounter reasoning demands that vary significantly across queries. Consequently, static retrieval-and-reasoning pipelines frequently result in inefficient resource allocation. In this paper, we introduce ALEX, a cost-aware adaptive inference architecture grounded in resource-rational accounts of metacognitive control. The proposed framework implements a metacognitive controller that monitors query complexity via multi-dimensional linguistic and confidence cues, and subsequently routes tasks to hierarchical workflows within an expected utility framework. Experimental results on legal benchmarks demonstrate that ALEX achieves superior accuracy while maintaining an efficient trade-off between accuracy and effort, consistent with the principles of bounded rationality. Furthermore, analyses of routing behavior and error correction driven by verification reveal how metacognitive monitoring can mitigate resource misallocation in LLM-based legal QA.