Skip to main content
eScholarship
Open Access Publications from the University of California

Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models

Creative Commons 'BY' version 4.0 license
Abstract

How large language models (LLMs) align with the neural representation and computation of human language is a central question in cognitive science. Using representational geometry as an explanatory lens, we addressed this by tracking entropy, curvature, and fMRI encoding scores throughout Pythia (70M--1B) training. We identified a geometric modularization where layers self-organize into stable low- and high-complexity clusters. The low-complexity module, characterized by reduced entropy and curvature, consistently better predicted human language network activity. This alignment followed heterogeneous spatiotemporal trajectories: rapid and stable in temporal regions (AntTemp, PostTemp), but delayed and dynamic in frontal areas (IFG, IFGorb). Crucially, reduced curvature remained a robust predictor of model--brain alignment even after controlling for training progress, an effect that strengthened with model scale. These results link training-driven geometric reorganization to temporal-frontal functional specialization, suggesting that representational smoothing facilitates neural-like linguistic processing.