Skip to main content
eScholarship
Open Access Publications from the University of California

About

The annual meeting of the Cognitive Science Society is aimed at basic and applied cognitive science research. The conference hosts the latest theories and data from the world's best cognitive science researchers. Each year, in addition to submitted papers, researchers are invited to highlight some aspect of cognitive science.

Abstracts with Poster Presentation

  • The Point of Pointing: Deictic Gestures Modulate Attentional Shifts and Cognitive Load in Simultaneous Interpreting

    Spoken language processing requires integrating acoustic signals with visual information. Deictic gestures such as pointing may facilitate this integration by constraining the referential domain before linguistic disambiguation. How multimodal cues modulate processing under cognitive load remains unclear. Simultaneous interpreting provides a task environment for examining speech processing under maximal cognitive load. In this study, we combined the Visual World Paradigm (VWP) with pupillometry to investigate how congruent, incongruent, and neutral pointing gestures influence attention and cognitive load in 24 professional interpreters. Eye-tracking revealed that congruent gestures elicited early anticipatory target fixations, incongruent gestures directed gaze toward competitors before rapid correction and looks towards the target, while the absence of gestures delayed target identification. Pupillometric measures yielded a counterintuitive pattern: the neutral condition evoked greater pupil dilation than the incongruent condition, whereas congruent and incongruent conditions did not differ significantly in the amount of generated load. This suggests that visual cues, even misleading ones, reduce processing demands when they accompany auditory input. These findings suggest that multimodal information alleviates cognitive load in a task environment competing for resources.

  • Quantifying Family-Level Social Development with Longitudinal 3D Tracking in Marmosets

    Understanding the emergence of social cognition requires continuous, high-resolution data from naturalistic environments, yet human studies are often limited by privacy and longitudinal feasibility. The common marmoset offers a compelling model due to its cooperative breeding and compressed developmental timeline that enables the study of family-level social dynamics. However, existing automated tracking tools struggle with the complex backgrounds, frequent occlusions, and rapid morphological changes inherent in home-cage settings. Here we introduce HOLMES, a deep learning-based system designed for longitudinal 3D tracking of multiple marmosets in naturalistic home-cage environments. By integrating Transformer-based temporal tracking with semantic segmentation and pose estimation, HOLMES achieves robust skeletal reconstruction and stable identity tracking across developmental stages. Applying to longitudinal family recordings, this system supports high-throughput quantification of the fine-grained temporal structure of parent-infant interactions and evolving social coordination, providing a way to link moment-to-moment social behaviors with long-term developmental outcomes.

  • Semantic Coexistence of Heart and Brain in Chinese: Computational Evidence from 3.4 Billion Tokens

    Cultural concepts of mind vary dramatically across societies. In Chinese thought, 心 (xīn, “heart-mind”) historically served as the seat of cognition and emotion, while 脑 (nǎo, “brain”) played minimal role. Modern neuroscience introduced brain-centered models, raising questions about whether scientific concepts replace or coexist with folk theories. How does scientific knowledge reshape traditional mental vocabulary in non-Western contexts? Here we show, through distributional semantic analysis of 3.4 billion Chinese tokens from classical texts to social media, that xīn and nǎo function as parallel rather than competing concepts. Contrary to eliminative materialism, xīn maintains its integrated multi-dimensional structure as a living folk concept, while nǎo enters as specialized medical terminology. This coexistence varies systematically: encyclopedic texts show convergence in biomedical contexts, while vernacular preserves traditional differentiation. These patterns suggest that folk and scientific frameworks operate as contextually deployed parallel systems rather than successive evolutionary stages. Our approach demonstrates how computational linguistics reveals mechanisms of cultural-cognitive modernization.

  • Intents and Purposes

    What are intentions, and what is the nature of practical reason? Cognitivism about intention and practical reason claim, respectively, that intentions entail predictive beliefs, and that practical reason reduces, at least partially, to theoretical reason. But cognitivism faces two distinct challenges: seemingly, we can intend the unexpected, and expect the unintended. This paper will advance a cognitivist theory of intention. To intend, I will argue, is to believe that a teleological explanation is true of one's own behavior. And as I will illustrate, given the nature of teleological belief, we can meet those aforementioned explanatory challenges. The result is a theory of intention that parsimoniously does what a sound theory of intention ought to do.

  • Temporal Dilation for Emotionally Salient Stimuli in Peri-hand Space

    Emotion timing literature has consistently shown temporal dilation for emotionally salient stimuli compared to neutral stimuli. Such stimuli have also shown preferential processing near the graspable space of the hands, known as peri-hand space (PHS). However, none of the studies have systematically examined the interplay of emotion, time perception, and peri-hand space in a unified framework. Therefore, the present study investigated the temporal factors underlying the processing of emotionally salient stimuli in PHS. Participants completed the temporal bisection task in Near-hand and Far-hand conditions. The stimuli were high-arousal emotional and low-arousal neutral images. Results showed temporal dilation for highly arousing emotional stimuli compared to low-arousal neutral stimuli in PHS, perhaps to facilitate in-depth cognitive evaluation of emotional salience. No such temporal processing differences between emotional and neutral stimuli were obtained in the far-hand condition. Temporal dilation is primarily attributed to early anticipatory mechanisms associated with PHS.

  • Do written and spoken languages differ in their phonological density?

    Languages evolve to be easy to use. Over the last couple of decades, however, language has shifted from being used predominantly in the spoken modality to being used predominantly in the written modality. Differences between the way spoken and written language are processed suggest that language might no longer be optimised for use and undergo changes. This paper uses an iterated learning paradigm (Experiment 1) and a pair communication game (Experiment 2) to test whether languages that develop in the written modality differ from those which develop in the spoken modality. Both experiments show that languages that develop in the written modality have sparser phonological neighborhoods. Results also indicate that neighborhood density is more detrimental for comprehension in the written modality than the spoken one, motivating the greater sparsity in the written modality. These results show that modality can influence phonological density and that languages might be currently undergoing changes to become sparser.

  • Strategic Algorithmic Advice Taking

    As algorithms increasingly mediate competitive decision-making, their influence extends beyond individual outcomes to shaping strategic market dynamics. In a preregistered experiment, we examined how algorithmic advice affects human behavior in a classic economic game with a unique, non-collusive, and analytically traceable equilibrium. Participants (N = 129) played a Cournot quantity competition with equilibrium-aligned or strategically biased algorithmic recommendations. While individualized equilibrium advice supported stable convergence, collusively downward-biased advice led to sustained underproduction and supracompetitive profits—hallmarks of tacit collusion. Participants responded more strongly and consistently to individualized advice than collective advice, potentially due to greater perceived ownership of the former. These findings demonstrate that algorithmic advice can function as a strategic signal, shaping coordination even without explicit communication. The results echo real-world concerns about algorithmic collusion and underscore the need for careful design and oversight of algorithmic decision-support systems in competitive environments.

  • Who knows what? Bayesian Inference of Competence guides Knowledge Attribution and Information Search

    Inferring others' competence is a challenge of social cognition, often occurring in contexts of limited information. Recent research suggests that people can infer the competence of others through Bayesian inference, but it is unclear whether these rational principles generalize to naturalistic settings of knowledge attribution. Using trivia questionnaires, we test whether people can infer others' competence and search for informative evidence in a way consistent with a rational Bayesian model. In Study 1, participants were presented with an individual's performance on a trivia question and predicted the individual's ability to answer other trivia questions from the same theme. Participants accurately predicted performance from limited information. Study 2 shows that participants can select which information would be most diagnostic for inferring an individual's competence. Computational modelling shows that participants' inferences, both when searching for and when integrating information about others' competence, are better described by Bayesian processes than by plausible heuristics.

  • Functional Brain Biomarkers of Self-Referential Bias in Remitted Depressed Outpatients

    Negative self-referential bias (NSB) is a core feature of depression, yet its neural correlates and relevance for relapse remains unclear. This study examined whether behavioral and neural markers of NSB predict relapse following prophylactic psychotherapy. In a two-year prospective fMRI study embedded within a randomized trial of Mindfulness-Based Cognitive Therapy and Well-being-focused Cognitive Therapy, remitted depressed outpatients completed a self-referential encoding task before and after treatment and were followed longitudinally. NSB predicted greater depressive symptoms but not relapse. In contrast, greater frontal default mode network and salience network reactivity during negative self-referential processing predicted relapse. Whole-brain analysis identified elevated prefrontal activation and somatosensory deactivation as relapse biomarkers, with treatment-related attenuation of somatosensory deactivation predicting lower relapse risk. Somatosensory deactivation emerged as the strongest overall predictor of relapse. These findings suggest that relapse vulnerability reflects not only exaggerated prefrontal negative self-processing but also insufficient recruitment of somatosensory systems supporting embodied self-experience.

  • Feature, Alignment, and Supervision in Category Learning: A Comparative Approach With Children and Neural Networks

    Understanding how humans and machines learn from sparse data is central to cognitive science and machine learning. Using a matched task design, we compare children and convolutional neural networks (CNNs) in a few-shot semi-supervised category learning task. Both learners received mixtures of labeled and unlabeled exemplars while supervision (1/3/6 labels), target feature (size, shape, pattern), and perceptual alignment (high/low) are systematically varied. We find that children generalize rapidly from minimal labels but show strong feature-specific biases and sensitivity to alignment. CNNs show a different interaction profile: added supervision improves performance, but both alignment and feature structure moderate the impact additional supervision has on learning. These findings show that comparisons between humans and machines should be controlled and sensitive to task structure. Comparing accuracy alone may hinder a more nuanced understanding of the conditions that differentially support success in each system.

  • Idiosyncratic codes in distributed minds

    Debates about the neural organization of semantic knowledge have long contrasted modular accounts, which posit anatomically specialized regions, with distributed accounts, which emphasize patterns of activity spanning many neural populations. Although multivariate neuroimaging has provided strong evidence for distributed representations, both perspectives often presume that the representational solutions supporting cognition are largely shared across individuals. Here, we challenge that premise by asking whether semantic representations are not only distributed, but also stable within individuals and meaningfully idiosyncratic across people. Using individualized multivariate fMRI decoding across two scanning sessions, we show that representations distinguishing faces, places, and objects are widely distributed across cortex, stable within individuals over time, and heterogeneous across participants. We then use these individualized representational profiles to guide transcranial magnetic stimulation (TMS), showing that stimulating non-canonical, decoding-defined sites selectively disrupts face similarity judgments in a manner comparable to stimulation of a canonical face-selective region. Together, these findings provide causal evidence that semantic representations arise from distributed neural systems whose organization is both stable and person-specific, calling for theories of cognition that make room not only for shared structure, but for the idiosyncratic ways meaning is represented across brains.

  • Regret-Driven Adaptation: How Humans Solve the Breadth-Depth Dilemma Through Trial-by-Trial Learning

    Adaptive decision-making requires solving the breadth-depth (BD) dilemma: balancing exploring more alternatives and evaluating each one more carefully. While normative models define optimal search set sizes given finite cognitive resources, the psychological mechanisms enabling individuals to achieve this remain unclear. In this study, we propose that subjective regret functions as a critical learning signal for solving the BD dilemma. Computational simulations demonstrated that search set size optimization is achievable by monitoring two distinct types of regret: one reflecting opportunity costs from narrow searches, and another reflecting evaluation errors from broad searches. A laboratory experiment using a sequential decision task showed that participants adjusted their search set sizes toward the theoretical optimum, with both regrets predicting these adjustments in complementary directions. Taken together, these findings suggest that regret is not merely a negative outcome but a signal that enables autonomous optimization of resource allocation in multi-alternative decision-making.

  • Predicting the Machine: Intentionality Framing Reduces the Prediction Gap in Human–AI Cooperation

    Cooperation is sustained by shared norms that regulate expectations about how others will behave. As artificial intelligence (AI) systems increasingly participate in social and economic interactions, a critical question emerges: Do the normative expectations that sustain human cooperation extend to artificial players? We propose that failures in human–AI cooperation may reflect a fundamental difficulty to form accurate predictions about artificial behavior that disrupts norm-based coordination. In Study 1 (N = 794), participants in a repeated Public Goods Game were less sensitive to prosocial norms when interacting with AI versus human players. This cooperative deficit was accompanied by a systematic "prediction gap": participants exhibited larger errors and directional biases when forecasting AI behavior. In Study 2 (N = 314), we observed that framing AI as intentional rather than mechanistic substantially reduced prediction errors and reduced cooperation differences between AI and human players. Analysis of trial-by-trial learning revealed that intentionality framing affected initial expectation calibration but not feedback-driven updating, consistent with a shift in participants' initial mental model of the AI. These findings suggest that norm-based cooperation in human-AI interaction depends on predictability.

  • Phonological Perception of Sign Language Models

    Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement. While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low-level statistical correlations. This work evaluates SLR model phonological perception by probing sensitivity using minimal pairs and measuring representational alignment with human behavioral data. Our results reveal emergent phonological sensitivity with clear architectural trade-offs: pose-based models are more sensitive to handshape contrasts, while pixel-based models better capture location changes. Furthermore, pose-based models learn latent representations that correlate with human perceptual similarity judgments (r –ï 0.49). These findings suggest that while SLR models exhibit emergent phonology, current training paradigms are insufficient to overcome their architectural inductive biases.

  • Algorithmic Consequences of Particle Filters for Sentence Processing: Amplified Garden-Paths and Digging-In Effects

    Under surprisal theory, linguistic representations affect processing difficulty only through the bottleneck of surprisal. Our best estimates of surprisal come from large language models, which have no explicit representation of structural ambiguity. While LLM surprisal robustly predicts reading times across languages, it systematically underpredicts difficulty when structural expectations are violated—suggesting that representations of ambiguity are causally implicated in sentence processing. Particle filter models offer an alternative where structural hypotheses are explicitly represented as a finite set of particles. We prove several algorithmic consequences of particle filter models, including the amplification of garden-path effects. Most critically, we demonstrate that resampling, a common practice with these models, inherently produces real-time digging-in effects—where disambiguation difficulty increases with ambiguous region length. Digging-in magnitude scales inversely with particle count: fully parallel models predict no such effect.

  • Stress, Structure, and Recall: A Spatial Axon-Growth Model Connecting Developmental Synaptogenesis and Attractor Memory

    Early-life stress is a major risk factor for cognitive deficits, yet the circuit-level mechanisms remain unclear. We hypothesize that systemic stress hormones and localized neurotrophic deprivation during critical periods alter synaptic connectivity. To investigate, we developed a large-scale computational model (N = 100 paired seeds) simulating two developmental stages: an activity-independent growth phase where axons navigate a BDNF chemoattractant landscape, and a functional phase probed with a recurrent attractor model. The "trauma" condition modeled a "Double Hit": a transient cortisol pulse increasing stochastic axon diffusion, combined with localized BDNF deprivation that weakened synaptic efficacy. Trauma produced a significant rise in axon tortuosity (p < 0.001), yielding structurally inefficient networks with collapsed spectral radius. Functionally, recurrent gain loss impaired retention of clustered memory engrams (–ï41% deficit), while random patterns remained robust. Parameter sweeps confirmed this pathology arises non-linearly from guidance noise and trophic weakness, linking developmental stress to memory impairments.

  • How Social Information Modulates Information-Sharing Strategies under Uncertainty

    Understanding what drives human information sharing is crucial for predicting the outcomes of collective decision-making, alongside social learning. Although individuals act as both recipients and transmitters in social interactions, existing research has largely examined these dual cognitive processes in isolation. This study proposes a computational model integrating social learning with information sharing, describing the decision-making process of agents who both learn from others' choices and decide whether to disclose their own in a social foraging situation. Agent-based simulations predict that specific information-sharing parameters could theoretically improve collective outcomes. Applied to behavioral data, our model analysis indicated that individuals exhibit positive selectivity by sharing rewarding choices, while demonstrating susceptibility to conformist influence, disclosing their actions more frequently when aligned with those of others. These findings provide a unified account that mechanistically links social learning and strategic disclosure, with implications for understanding collective dynamics such as the spread of misinformation.

  • Sensitivity to Baby-Schema Features Independent of Perceptual Experience – A Face Discrimination Paradigm in Leipzig and Batek Communities

    To recognize and act on baby-schema features – such as large eyes and a rounded face - is commonly seen as an automatic, innate capability in humans. Following this logic, adults' perception of baby-schema features should be independent of cultural background. Contrasting this, evidence shows that perception of adult faces is tuned by culture-specific visual experience. In this preregistered study, we investigate whether sensitivity to baby-schema is dependent on visual experience in two diverse populations (urban German; small-scale Batek, Malaysia). We morphed child faces sourced from the target communities to different degrees of baby-schema and asked participants to detect the odd-one-out of three faces in both in- and out-group trials across seven difficulty (difference in baby-schema-degree) levels. While difficulty strongly effected accuracy in both communities, we found little evidence that baby-schema sensitivity was modulated by experience, suggesting robustness in perceptual processing and possibly generalized high relevance of this class of stimuli.

  • Computation or Weight Adaptation? Rethinking the Role of Plasticity in Learning

    The human brain can rapidly adapt to new tasks and environments, a capacity traditionally attributed to structural changes in the learning system, such as neural plasticity. Here, we revisit this assumption by asking whether adaptive behavior can emerge through computation alone, without parameter updates. Using large language models (LLMs), we examine statistical learning paradigms that require identifying regularities in arbitrary word sequences and are commonly considered to depend on plasticity. We show that LLMs can acquire such structure through in-context exposure, capturing the underlying regularities without weight adaptation, despite the divergence of these tasks from their natural language training data. These findings suggest that sufficiently trained learning systems may exhibit a greater degree of flexibility through computation than previously acknowledged, and highlight the potential of deep learning models as tools for generating hypotheses about learning mechanisms in the brain.

  • Search through Memory Structure

    Semantic memory encodes both the individual features of concepts and the relations between them. We develop computational memory models that describe how people search through these memory structures. Our empirical paradigm involves open-ended word-pair analogy generation and free association tasks, and we infer the respective roles of featural and relational information in these tasks using formal model comparisons. Our tests reveal that both types of information play a key role in memory search, but they are recruited differently depending on the task and exhibit distinct dynamics. Importantly, we can quantitatively predict these effects and parameterize their associated mechanisms using a principled extension of established memory models, thereby integrating theoretical work on structured reasoning and on memory search.

  • A Meta-Analysis Synthesizing 20 Years of Evidence on the Balloon Analogue Risk Task (BART)

    The cognitive sciences have developed diverse tasks to measure the psychological processes during decision making. Here we focus on the Balloon Analogue Risk Task (BART) and present a Bayesian meta-analysis synthesizing 1,665 effect sizes from 237 studies, thus comprising the data of 51,932 participants. We (i) take stock of the psychometric properties of the BART and (ii) quantify measurement flexibility and its impact on the task's psychometric properties. The BART showed satisfactory levels of test–retest reliability, yet it showed no to low known–groups validity, convergent validity, and external validity. We observed substantial measurement flexibility in the meta-analyzed studies, and specific study characteristics influenced the estimates of the task's psychometric properties.

  • Developmental changes in children's production of line drawings in a picture-sparse environment: Evidence from Kisumu, Kenya

    Children growing up in environments with many drawing opportunities improve dramatically at producing recognizable drawings over the course of childhood. Does this same developmental trajectory hold for children growing up in environments with fewer drawing opportunities? We measured children's drawing recognizability cross-sectionally in a drawing-sparse context. Moreover, by leveraging within-context variation in children's drawing experiences, we investigated whether, at an individual level, drawing recognizability increases with children's drawing experiences. A preregistered sample of 120 4- to 9-year-olds in Kisumu, Kenya drew 12 object categories in random order. Older children produced more recognizable drawings even after controlling for tracing ability, and children with more drawing experience produced more recognizable drawings even after controlling for tracing ability and home environment variables. These results suggest commonalities in drawing development across cultures and document the experience-dependence of drawing development in drawing-sparse contexts.

  • Comparing Auditory Word Identification and Lexical Decision: Insights from the Auditory English Lexicon Project

    The present study investigated task-general and task-specific effects of phonological and lexico-semantic word properties on auditory word identification and auditory lexical decision. The data comprised identification accuracy and lexical decision latencies from the Auditory English Lexicon Project (AELP). Item-level regressions were performed for 20 word properties across each of the six talkers available in the database. Task-general effects that were consistently robust across both tasks revealed that words that were more familiar, prevalent, and structurally distinctive were recognized faster and identified more accurately. The effects of frequency, concreteness, arousal, and number of morphemes were consistently specific only to lexical decision and were not robust for word identification. These findings highlight the utility of using the AELP data to conduct six replications across talkers from different regions and genders to test the robustness and generalizability of specific word properties' influences on spoken word processing.

  • Speaking of Decisions: Using Verbal Decision Protocols and Large Language Models to Uncover Psychological Mechanisms Involved in Decision Making

    Many theories of decision making have proposed that diverse psychological mechanisms (e.g., mechanisms related to affect, goals, or social factors) shape people's decisions under risk and uncertainty. How best to study such mechanisms during decision making? In a preregistered study (N = 699), we tested whether six experimentally induced mechanisms (incidental anger, cognitive load, lack of knowledge, goal pursuit, descriptive social norms, and time pressure) leave linguistic traces during decision making and can thus be identified with large language models (LLMs). To this end, participants completed four decision-making tasks while recording verbal decision protocols (VDPs). We then inferred the presence of psychological mechanisms in VDPs using "speech-to-psych", a newly developed analysis pipeline leveraging different LLM pipelines. Across analytic methods and decision contexts, psychological mechanisms as induced by experimental conditions could be identified at above-chance level (average area under the curve = .61). Feature importance scores revealed that this identification frequently relied on linguistic markers associated with the targeted psychological mechanism. These results indicate that at least some psychological mechanisms leave detectable linguistic traces and that LLM-based analyses could serve as a powerful measurement tool for advancing process-level understanding of decision making.

  • Cultural Differences in the Effect of Mask Use on Cross-Race Face Perception: An Eye-Tracking Study

    We recruited Asian and White adults to examine cultural differences in the effect of mask use on face scanning behavior and social categorization. Mask use impaired social categorization accuracy for White but not Asian faces, suggesting the diagnostic features for Asian faces were more in the eyes. Consistent with this finding, Asian adults showed a bias to categorize ambiguous faces as Asian regardless of the mask condition, whereas White adults exhibited a similar bias only for masked faces, resulting in a larger mask effect. In face scanning behavior, mask use made Asian adults look more towards the left as opposed to the right eye compared with White adults. A more right-eye-biased pattern due to mask was associated with an increased mask effect in ambiguous face categorization, suggesting a link between cultural differences in face scanning behavior and social categorization. These findings have important implications for social cognition with mask use.

  • A Computational Model of Self-Signaling in Procrastination

    People often procrastinate on tasks that they are ostensibly motivated to complete. We argue that dominant formal accounts – particularly temporal discounting models – capture the impulsive nature of procrastination but struggle to provide comprehensive cognitive explanations of three central features: which tasks people procrastinate on, why procrastination often feels bad, and when people finally begin working after a period of procrastinating. We propose a computational framework to explain the mechanisms of avoidance in procrastination. In our model, agents choose how much effort to exert by reasoning over a probabilistic generative model of task progress over time. This structure supports inferences about competence and task difficulty from noisy progress signals. Crucially, agents incur an additional cost for negative belief updates about their own competence, capturing an aversion to appearing incompetent that motivates self-handicapping. With this structure, the model reproduces key behavioral, affective and cognitive signatures of procrastination.

  • Visuospatial Encoding in Algebraic Processing: Evidence from a Dual-Task Study

    Algebraic competence is a critical foundation for advanced mathematical reasoning and success in STEM disciplines. However, whether algebraic processing relies on phonological or visuospatial encoding remains a central debate in cognitive psychology. We investigated the representational format of algebraic knowledge using a dual-task paradigm with 50 adult participants performing an algebra task concurrently with either a phonological or a visuospatial working memory (WM) task. While algebraic performance remained resilient under load, results revealed a selective interference effect on the secondary WM task: algebra significantly impaired visuospatial, but not phonological, WM performance. This impairment on visuospatial WM performance only suggests that algebra and visuospatial processing compete for a shared cognitive pathway, providing robust evidence for the visuospatial grounding of algebraic thinking.

  • How Do We Assess Executive Functions ? A Systematic Review of Tasks and Processes for Inhibition, Updating and Shifting

    Executive functions are essential for adaptive behavior, but there is substantial variability in the way these functions are assessed. Unfortunately, the literature has mostly focused on tasks rather than the processes involved in those tasks. Adopting a process-based rather than a task-based approach could improve the assessment and conceptualization of executive functions. We systematically reviewed literature on how researchers assess executive functioning, including the tasks used and the processes involved in those tasks. We included studies measuring inhibition (n = 844), updating (n = 844) or shifting (n = 1048) among human participants with any performance-based measure. The results show that researchers have used many different tasks to measure the same concept, but tasks supposed to target the same executive function can actually engage very different processes. In addition, the participants' age or disorder has a significant effect on the type of tasks retained to assess the same executive function and the processes they involve, which means developmental literature does not necessarily assess the same "inhibition" as the adult neuropsychological literature. Our review underlines the importance of using multiple tasks to measure inhibition, updating or shifting, and the importance of clarifying and reporting on the processes supposed to be involved in these tasks, moving from a task-based approach to a more realistic process-based approach.

  • Signatures of discrete action symbols emerge in a task-optimized neural model

    Key to intelligence is the ability to flexibly compose and combine atomic concepts into representations that guide behavior, making "infinite use of finite means". However, modeling often assumes discrete concepts exist and focuses on their use, leaving it unclear how they are learned from complex sensory inputs and mapped to motor outputs. To address this limitation, we build on a drawing-like task and behavioral metrics that indicate discrete structure in motor behavior from Tian et al. (2025). Importantly, these discrete concepts were encoded in neural recordings from non-human primates that exhibited these behavioral metrics, validating the metrics. Here, we use this task and behavioral metrics as a modeling target, to quantify how discrete structure arises in a task-optimized neuro-symbolic model (Liang et al., 2022). Our model recapitulates these behavioral metrics, offering a foundation to compare with neural data and understand what drives the emergence of discrete conceptual structure.

  • Eye gaze reflects episodic memory sampling during decision making

    Episodic memory enables decision making by allowing choices to be guided by detailed past experiences. Yet, because this process unfolds covertly, the mechanisms governing how memories impact deliberation remain poorly understood. Here we use gaze reinstatement to reveal the hidden dynamics of episodic memory retrieval during value-based choice. We tracked participants' eye movements while they encoded episodes consisting of items with associated rewards at distinct spatial locations and subsequently made decisions based on these episodes. Critically, no visual information was provided to participants during their choices. Nonetheless, participants systematically directed their gaze toward the encoding locations of retrieved episodes during deliberation, and these gaze patterns predicted their choices. These findings demonstrate that eye movements provide a high-resolution window into moment-by-moment episodic sampling during decision making. By revealing which episodes are retrieved and when, this approach opens new directions for investigating how episodic memory guides adaptive behavior.

  • Development of Supplementary Gestures with Speech in Mandarin-Speaking Children

    This study examined the developmental trajectory of supplementary gestures, which conveyed information distinct from speech, and their semantic relationships with accompanying speech in 16 Mandarin-speaking children. Children were observed during mother–child interactions between 12 and 21 months of age at three-month intervals, with minimal age variation within each time point. Findings showed a marked increase in supplementary gestures at 18 months, alongside the predominance of declarative pointing within gesture-speech acts and one-word utterances. This specific cross-modal pattern remained stable through 21 months. Most cross modal combinations conveyed entity ideas, with predicative ideas occurring infrequently and propositional ideas even less often. The common 'entity-related', 'what is done to what', and 'who initiates what and who/what is affected' knowledge was linguistically expressed only about three to six months later than in gesture-speech acts. Overall, the findings support the enduring role of supplementary gestures in enriching children's communicative expressiveness within a single act.

  • Disentangling Social and Material Preferences from Emotional Expressions: A Bayesian Approach to Reverse Appraisal

    Emotional expressions signal how others evaluate outcomes, but a single expression can reflect multiple underlying preferences. We examine whether and under what constraints observers infer material preferences (what outcomes someone values) and social preferences (how they weigh self vs. other) from affective feedback. In a multi-issue ultimatum game, participants (N=323) observed agents' graded emotional reactions to resource allocations and inferred their preferences. Results revealed a functional asymmetry between the two domains. Material preference accuracy was constrained by signal identifiability, whereas social preference estimates showed weak correspondence with Bayesian model predictions and clustered near the center. This suggests that observers do not infer both preferences symmetrically; instead, social preference estimation appears more strongly shaped by robust priors and/or measurement constraints. These findings highlight structural limits on social inference and suggest that human observers may rely on strong prior assumptions when interpreting underdetermined signals.

  • Mind the Bind: Mapping Expertise in Semantic Schemata Using Network Science

    Expertise shapes the organization of semantic knowledge, yet prior work using cognitive network science remains sparse and limited in methodological diversity. Study 1 used a free-listing task to identify the concepts constituting the visual art schema in experts and laypeople. Study 2 collected pairwise relatedness judgements of these concepts to construct individual semantic networks. Higher expertise was associated with richer and more diverse schemata, more individualized networks, higher clustering coefficients, shorter average shortest path lengths, and unchanged modularity. Node-level analyses showed that fluency enhanced the structure of basic concepts, while the full advantage of experts emerged through integration of advanced concepts. These findings link expertise to more efficient and differentiated semantic networks, suggesting a mechanism by which knowledge organization supports expert-level cognition.

  • Stereotype Threat Favors Accuracy Over Speed: Dismantling the Influence of Gender Stereotypes on the Speed-Accuracy Tradeoff in Mathematical Tasks

    Stereotype Threat describes the negative impact on cognitive performance caused by the activation of negative social stereotypes. This study investigates the underlying cognitive mechanisms of gender-specific Stereotype Threat effects. Four online experiments were conducted with 1,164 participants, randomly assigned to either a Stereotype Threat or a control condition. To model the underlying cognitive processes, parameters of the drift diffusion model were estimated and effects on accuracy and response times were reported. Results showed no differences in accuracy as an effect of Stereotype Threat. However, women in the Stereotype Threat condition consistently responded more slowly and, in one experiment, exhibited a higher threshold separation when completing the mathematical task, indicating more conservative response tendencies, compared to men and the control group. The study addresses the need for potential interventions such as stereotype awareness to mitigate these effects and calls for further research into cultural and social influencing factors.

  • Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians

    "Delusional spiraling'' is a form of AI-induced psychosis, an emerging phenomenon where AI chatbot users find themselves dangerously confident in outlandish beliefs. This phenomenon is popularly attributed to "sycophancy,'' chatbots' well-documented bias towards validating users' claims. But an explanation of why sycophancy should cause delusional spirals---and why they occur only rarely, but seriously---has not been forthcoming. Here we take a theoretical and computational approach to probing the causal link between sycophancy and delusional spiraling. We propose a simple Bayesian model of a user conversing with a chatbot, and formalize notions of sycophancy and delusional spiraling in that model. We then show that in this model, even an ideal Bayesian user is vulnerable to delusional spiraling, and that sycophancy plays a causal role. Furthermore, this effect persists despite two candidate mitigations: preventing chatbots from hallucinating false claims, and informing users of AI sycophancy. The full paper is available at .

  • Meta-Learning Captures Human-Like Geometric Sensitivity

    Humans show systematic biases in geometric perception: shapes with higher regularity (symmetry, right angles, parallel sides) and simplicity (fewer elements) are easier to recognize. These biases have been difficult to replicate in neural networks, leading to the hypothesis that the human visual system uniquely employs a "geometric language of thought''---discrete symbolic machinery for representing shapes---that cannot be reproduced by neural networks. We show that neural networks trained via meta-learning can reproduce these biases without symbolic primitives or massive pretraining. Training on concepts with fixed complexity yields human-like regularity sensitivity, whereas training on concepts with varying complexity yields simplicity sensitivity. This dissociation suggests that meta-learning induces priors that exploit the most diagnostic dimension of variation in the training environment. More broadly, our results suggest that these perceptual biases can arise from the statistics of the training environment itself, rather than being innately endowed.

  • Are LLMs Biased Like Humans? Causal Reasoning as a Function of Prior Knowledge, Irrelevant Information, and Reasoning Budget

    Large language models (LLMs) are increasingly used in domains where causal reasoning matters, yet it remains unclear whether their judgments reflect normative causal computation, human-like shortcuts, or brittle pattern matching. We benchmark 20+ LLMs against a matched human baseline on 11 causal judgment tasks formalized by a collider structure ($C_1 \!\rightarrow\! E\! \leftarrow \!C_2$). We find that a small interpretable model compresses LLMs' causal judgments well and that most LLMs exhibit more rule-like reasoning strategies than humans who seem to account for unmentioned latent factors in their probability judgments. Furthermore, most LLMs do not mirror the characteristic human collider biases of weak explaining away and Markov violations. We probe LLMs' causal judgment robustness under (i) semantic abstraction and (ii) prompt overloading (injecting noise), and find that chain-of-thought (CoT) increases robustness for many LLMs. Together, this divergence in reasoning style can augment human-reasoning when known human biases are undesired.

  • Question Asking as Active Learning: Scaffolding Shapes the Development of Child Inquiry

    Question asking is a key mechanism through which children resolve knowledge gaps and construct causal models of the world. This study investigates the developmental trajectories of diverse inquiry types (factual, procedural, causal, clarification, and epistemic questions) and examines how conversational scaffolding influences these transitions. Using longitudinal naturalistic data from children (age= 2;00 to 4;00), we employed a Bayesian regression analysis to map the restructuring of child inquiry. Results indicate a significant developmental shift from simple informational seeking toward more sophisticated explanatory and task-oriented inquiry. While factual inquiries declined with age, procedural, causal, clarification, and epistemic questions showed robust growth. Notably, the velocity of this transition was mediated by social context: high-scaffolding environments accelerated the increase in complex inquiries. Our findings suggest that although complex inquiry is a universal feature of early curiosity, its developmental trajectory is socially shaped and varies across individuals.

  • Evolutionary Theory Makes Predictions About Cognition Beyond Rational Optimization

    Evolution by natural selection produces the appearance of design in minds, brains, and behavior. While rational models of cognition are usually considered to be consistent with evolution, they are rarely derived from distinctively evolutionary assumptions. Bayesianism, reinforcement learning, rational analysis, and resource rationality all postulate optimization objectives of coherence in information use or maximization of expected value, potentially under resource constraints. However, cognition is performed by an individual within a lineage where success is determined by the persistence of the lineage across generations. Here, we present a signal detection model that shows how this distinctly evolutionary objective produces predictions that depart from those of rational models. We find that the source of environmental uncertainty determines cognitive design, with weak priors being favored in environments dominated by change. This ensures offspring vary so that some profit against novel challenges, and is favored even at the expense of accurate decision making.

  • Designing behavioral interventions with explainable artificial intelligence

    Identifying how language drives human behavior is crucial, both practically to influence people and theoretically for understanding psychology, culture, and myriad behaviors predicated on verbal or written communication. Here, we present a method based on explainable artificial intelligence (AI) that identifies what aspects of the language people read are associated with their reactions. As an example, we use this method to redesign choice tasks to modulate people's risk taking across five novel datasets containing 38,911 responses to 255 distinct decision tasks. Overall, we both i) design interventions to influence behavior at scale with greater efficiency than methods such as AB-testing and ii) extract previously unknown drivers of behavior as a basis for developing new theories across academic disciplines.

  • Identifying Mind Wandering Episodes during Virtual Cognitive Stimulation Therapy through Gaze Estimation from Videos

    Recent studies suggest that eye movements may be used to monitor task-specific mind wandering (MW) episodes during online learning. We examined whether we could replicate these lab-based findings in real-life virtual cognitive stimulation therapy (vCST) sessions using eye movements estimated from video by machine learning methods without eye trackers. We found that lower joint attention was a reliable indicator of MW for tasks involving well-defined strategies. For tasks involving well-learned visual routines, in contrast to previous studies, higher rather than lower eye movement consistency was associated with MW. This may be because real-life vCST sessions involved less structured discussions with more varied task demands, where lower eye movement consistency may reflect active engagement. Our results suggest the feasibility of using machine learning to monitor eye movements from video for MW detection, and raise the issue of generalizability from lab research to real-life scenarios due to potential differences in task demands.

  • Successful Modal Reasoning Depends On WHAT Went WHERE

    Quantifying over possibilities requires Modal Reasoning (e.g., categorizing events as necessary, possible, and impossible). This suite of abilities matures during the preschool years, but toddlers sometimes succeed at reasoning over possibilities and sometimes fail. Does this diversity of results reflect mere variability? We propose that these mixed findings reflect a deeper distinction: toddlers fail when tested about ambiguous locations of objects (WHERE) and succeed when tested about ambiguous identities of objects (WHAT). In Experiment 1, we directly compared these forms of ambiguity using matched tasks. Consistent with the hypothesis, toddlers succeeded at reasoning about ambiguous identities but failed when the ambiguity concerned location. Experiment 2 replicated this asymmetry in language comprehension as toddlers successfully comprehended modal terms like "can" and "have to" in an identity task but not in a location task. Thus, Modal Reasoning emerges earlier than suggested, and diversity in results reflects the contributions of distinct cognitive systems.

  • Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity

    Are large language models (LLMs) creative in the same way humans are, and can the same interventions increase creativity in both? We evaluate a promising but largely untested intervention for creativity: forcing creators to draw an analogy from a random, remote source domain ("cross-domain mapping"). Human participants and LLMs generated novel features for ten daily products (e.g., backpack, TV) under two prompts: (i) cross-domain mapping, which required translating a property from a randomly assigned source (e.g., octopus, cactus, GPS), and (ii) user need, which required proposing innovations targeting unmet user needs. We show that humans reliably benefit from randomly assigned cross-domain mappings, while LLMs, on average, generate more original ideas than humans and do not show a statistically significant effect of cross-domain mappings. However, in both systems, the impact of cross-domain mapping increases when the inspiration source becomes more semantically distant from the target, and above a distance threshold, cross-domain mapping benefits the most capable LLMs too. Humans and LLMs differed in how they used the source: Humans tended to transfer surface features of the source, whereas LLMs transferred structural and functional properties. Our results highlight the role of remote association in creative ideation and systematic differences in how humans and LLMs respond to the same intervention for creativity.

  • Stochastic Parrot See, Stochastic Parrot Do: Hierarchical Sequence Processing Across Artificial and Biological Intelligences

    Hierarchical sequence processing is central to complex behaviours such as language, math, music, and tool use. Recent research has investigated hierarchical reasoning in both biological and artificial intelligences, but has failed to go beyond benchmarking. In this comparative cognition study, we tested hierarchical reasoning across 4 generations of state-of-the-art LLMs against 4 biological intelligences: adults, children, crows, and monkeys. We conducted Bayesian modelling analyses using both previously published in vivo data and newly collected in silico data. We found that the capacity for hierarchical reasoning is present across generations of LLMs, on a continuum that parallels biological intelligences. Newer (reasoning) models exhibited the strongest proclivity for hierarchical reasoning, while the more primitive ones showed graded performance mirroring that of children and non-humans. Together, these findings position LLM as a unique model organism for comparing hierarchical cognition across biological and artificial intelligences.

  • Reduced Inhibition of Return for Self-relevant Stimuli

    Self-relevant stimuli (e.g. own-name) have been known to enjoy cognitive priority, known as the self-prioritization effect (SPE). Attentional bias towards any event consists of two components: how quickly attention is captured and how long attention dwells on that event. Although faster attentional capture by self-relevant stimuli has been consistently established in the literature, it is unclear whether captured attention is held by self-relevant stimuli for longer, delaying attentional disengagement from self. Thus, the current study aims to provide a more comprehensive picture of the attentional mechanisms underlying SPE by employing the Posner cueing paradigm to investigate both facilitation (capture) and inhibition of return (disengagement), to test the commonly touted yet not investigated delayed disengagement hypothesis. The results found a self-specific facilitation at shorter CTOA, in line with literature, as well as a reduced IOR for self, indicating a slower attentional disengagement from self. The findings and implications are discussed in detail.

  • Exaggerated Volatility Beliefs Drive Social Hallucinations in Paranoia

    The social content of paranoid delusions has prompted theories that paranoia stems from deficits in social cognition. For instance, paranoid individuals are more likely to falsely attribute harmful intentions to randomly moving visual patterns—'social hallucinations.' However, paranoia also involves domain-general deficits in belief updating and learning. To investigate if domain-general deficits are sufficient to explain social hallucinations, participants with varying paranoia levels completed reversal learning tasks (social/nonsocial _ easy/hard), then judged intentional motion in ambiguous displays. High-paranoia participants showed increased win-switching and elevated social hallucinations, with the former predicting the latter—suggesting linked mechanisms. Hierarchical Gaussian Filter modeling revealed context-dependent mechanisms: when volatility was salient, prior volatility beliefs (__) linked win-switching, social hallucinations, and paranoia; when social context was salient, meta-volatility learning (__) mediated these relationships. These findings demonstrate that social hallucinations emerge from domain-general belief updating problems operating through context-sensitive pathways.

  • In the loop, out of sync: moral cognition in human-robot interactions

    The ubiquitous integration of AI-powered systems in morally consequential decision-making procedures raises a thorny question: when such systems generate harm, who should be held responsible? A prominent regulatory response proposes that suitably designed control architectures can ensure that blame is appropriately allocated. To live up to this promise, the proposed architectures must be both normatively adequate and regarded as such, since otherwise they are at best practically useless, and at worst useless because morally problematic. We examine two of the most widely discussed proposals. Experiment 1 (n=260) investigates whether laypeople attribute responsibility differently across architectures that place human agents in-, on-, and out-of-the-loop. Experiment 2 (n=520) extends this inquiry by differentiating between control structures that either violate or satisfy Santoni de Sio's and van den Hoven's (2018) 'track and trace' requirements. Our results reveal that neither proposal fully succeeds in directing laypeople's responsibility judgements along the pathways they prescribe. Implications are discussed.

  • Guessing reveals internal models of perceptual precision

    When observers lack sufficient information to support a confident response, they guess. Guessing is pervasive in perception and memory, yet standard mixture models treat it as uniform lapse noise. We measured guessing directly in continuous report using (E1) extreme-load, ultra-brief trials and (E2) backward-masked stimulus-absent trials in which no stimulus appeared but observers believed one had. Across both experiments, we found that that guess responses are systematic, observer-specific, and inversely related to feature-specific precision. We then introduce a new empirically-informed computational model that replaces the standard uniform lapse term with each observer's measured guess distribution. Critically, the model introduces no additional free parameters, yet yields improved fits and recovers guessing from stimulus-present trials via trial-level posterior inference. This framework reconceptualizes guessing as the complement of perceptual precision and provides a principled alternative to uniform lapse assumptions. More broadly, it demonstrates how latent internal models can be inferred by measuring guesses.

  • When Topology Matters: Perturbative Analysis of Nonlinear Social Learning on Networks

    Cultural artifacts such as language rarely evolve in isolation. Rather, they are the product of the inductive biases of learners and the social learning strategies they deploy within their environment. Formal accounts of this process using Bayesian agents often emphasize the role of inductive biases (priors) in controlling the stationary distribution of artifacts, with no effect of social structure (topology) under random social learning (i.e., choosing a random person to learn from). Here, we explore how nonlinear social learning strategies modify this picture by studying the interaction of two social learning mechanisms (conformity and utility biases) with network topology in models of language evolution. Using perturbative expansions around the random baseline, we show that nonlinear learning strategies can couple local network correlations with global population statistics. This coupling leads to deviations from the prior and allows topology to influence the stationary distribution of languages, which we confirm in a simulation.

  • Detecting Incentive Skepticism in AI Persuasion Dialogues with LLM-Based Stance Inference

    Large language models (LLMs) enable tailoring of persuasive messages at scale, raising hopes and concerns that persuasion may become a problem of optimizing message content. However, people can resist persuasion for qualitatively different reasons, including doubts about the persuader's motives. Here, we reanalyze a public dataset of multi-turn human-LLM persuasion dialogues to examine whether such resistance can be detected early from recipients' own replies. We use an LLM as a proxy forward model of human response: we specify different stances in the system prompt and compute how likely the model is to generate the observed rebuttal. We find that skepticism about the sender's incentives, where recipients attend to possible incentive misalignment and discount message content, is common. Such skepticism predicts lower subsequent persuasion success and a higher risk of backfire. These findings highlight limits of content optimization and suggest the methodological potential of seeing LLMs as superpositions of human cognition.

  • How Do We Choose What to Think? Semantic Space Navigation and Individual Variation

    How do people generate options in everyday purchase decisions? Building on prior work, we studied how people search semantic space when generating options for daily-use products. Study 1 (N=200) showed that items coming to mind first were both more preferred and more frequently used, demonstrating that accessibility and personal preference predict retrieval order. Studies 2-3 (N=220) constructed empirically-derived feature spaces for product categories. Within these spaces, semantic distance weakly predicted generation speed, showing Lévy-like foraging but with substantial deviation. This raised the possibility that individual differences fundamentally shape semantic search in option generation. Study 4 (N=100) directly tested this interpretation, revealing that mental search is recalcitrant: broad category activation intrudes into subsequent feature-focused search, while feature-focused search first facilitates later global exploration. These findings demonstrate that option generation reflects population-level foraging dynamics modulated by individual differences in how people weight features and navigate personalized semantic neighborhoods.

  • Parallelograms Strike Back: LLMs Generate Better Analogies than People

    Four-term word analogies (A:B::C:D) are classically modeled geometrically as "parallelograms," yet recent work suggests this model poorly captures how humans produce analogies, with simple local-similarity heuristics often providing a better account (Peterson et al., 2020). But does the parallelogram model fail because it is a bad model of analogical relations, or because people are not very good at generating relation-preserving analogies? We compared human and large language model (LLM) analogy completions on the same set of analogy problems from Peterson et al. (2020). We find that LLM-generated analogies are reliably judged as better than human-generated ones and are also more closely aligned with the parallelogram structure in a distributional embedding space (GloVe). Crucially, we show that the improvement over human analogies was driven by greater parallelogram alignment and reduced reliance on accessible words rather than enhanced sensitivity to local similarity. Moreover, the LLM advantage is driven not by uniformly superior responses by LLMs, but by humans producing a long tail of weak completions: when only modal (most frequent) responses by both systems are compared, the LLM advantage disappears. However, greater parallelogram alignment and lower word frequency continue to predict which LLM completions are rated higher than those of humans. Overall, these results suggest that the parallelogram model is not a poor account of word analogy. Rather, humans may often fail to produce completions that satisfy this relational constraint, whereas LLMs do so more consistently.

  • Adversarial construction as a potential solution to the experiment design problem in large task spaces

    Despite decades of work, we still lack a robust, task-general theory of human behavior even in the simplest domains. In this paper we tackle the generality problem head-on, by aiming to develop a unified model for all tasks embedded in a task-space. In particular we consider the space of binary sequence prediction tasks where the observations are generated by the space parameterized by hidden Markov models (HMM). As the space of tasks is large, experimental exploration of the entire space is infeasible. To solve this problem we propose the adversarial construction approach, which helps identify tasks that are most likely to elicit a qualitatively novel behavior. Our results provide a proof of concept that adversarial construction can identify behaviorally diagnostic tasks more efficiently than random sampling in a continuous task space, suggesting a promising heuristic for scaling cognitive experiment design beyond single-task paradigms.

  • Pupil Dynamics Track Preparatory Control and Selective Learning in Decision Making

    A unifying theory of phasic pupil responses suggests that pupil diameter reflects information gain following an observation. However, it remains unclear whether different types of information gain (Shannon surprise and belief updating) from different sources (control signals, decision-relevant stimuli, and decision-irrelevant stimuli) modulate pupil responses in similar or distinct ways. To address this, we developed a decision-making task that dissociates these information streams within a single trial. Analyzing pupillometry data from 30 participants, we identified two computational signatures. First, cue surprise modulated pupil responses before stimulus-driven effects emerged, indicating preparatory control. Second, learning-related information gain reliably modulated pupil responses after decisions and was selective for belief updates from decision-relevant stimuli, even though decision-irrelevant stimuli were also perceptually encoded. These findings refine the information-theoretic framework by revealing pupil modulation by proactive control and distinguishing automatic, surprise-related dilation from selective, learning-related dilation gated by decision relevance.

  • Signatures of hierarchical, heuristic-guided planning in real-world human conceptual navigation

    Real-world planning requires navigating vast spaces of possible futures, many of which are unknown. However, prior studies of human planning have focused on simplified environments with fully specified state spaces. To bridge this gap, we explored human planning in the Wiki Game, where players navigated a vast and sparsely known conceptual network–Wikipedia. We analyzed human behavior across two datasets (a large naturalistic online dataset and a controlled laboratory experiment) and simulated the performance of various agents implementing various combinations of strategies. We identified signatures of heuristic and hierarchical strategies within both human and agent behavior: both participants and hierarchical heuristic-guided search models made choices that appeared biased by semantic similarity and graph centrality at stereotyped points in their trajectories and exhibited deliberation dynamics that were modulated by heuristic factors and category boundaries. Overall, this suggests that humans may navigate vast, sparsely known environments using hierarchical, heuristic-guided search.

  • Collective Decision-Making in Coupled Echo State Networks

    Interaction can sometimes hinder human collective performance, while aggregating independent responses ("wisdom of crowds") is expected to improve performance through statistical facilitation. Yet some coordination studies show that interacting pairs can outperform their individual members and nominal groups (Bahrami et al., 2010; Szary & Dale, 2013, 2014). We present a reservoir computing model demonstrating how interaction can enhance joint performance. Two echo state networks were trained independently to identify Japanese speakers from vowel recordings and tested jointly with and without coupled feedback. The model reproduces the dyadic advantage reported by Bahrami et al. (2010) with interaction via pooled-feedback, but not with independent self-feedback. Model dynamics revealed that interaction was beneficial when uncoupled outputs were prone to diverging from runaway feedback over time. Our results offer a formal framework for explaining when and why interaction enhances collective intelligence.

  • Meaning is niche construction: semiotic artifacts and cognitive externalism

    How to provide a locus of observation for the formal notion of semiosis? The notions of niche and artifact are especially capable of updating the thesis, formulated by Peirce, that one cannot think without external signs, associating it to new empirical and theoretical methods and results. In this poster, we introduce the notion of niche of semiotic artifacts. In our approach, cognition is semiosis, sign-action, in a process that takes the form of niche construction. In comparison with current usages of the term artifact to mean a material "thing" which is produced, semiotic artifacts are processes, signs-in-action. Niches of semiotic artifacts are structured spaces of fundamental conditions for stability of sign-action, conditions such as situatedness (co-localization) and temporal distribution between communities of agents, artifacts, and their environments. Niches of semiotic artifacts offer conditions for emergence of habit and surprise in semiosis/cognition. This line of inquiry suggests a cognitive semiotics framework based on dynamical, distributed, and emergent relations.

  • How universal is universalization? Exploring the use of universalization in norm violation across the globe

    How do people know when it is permissible to break a rule? Sometimes people universalize, asking "What if everyone felt at liberty to violate the rule?" While there is mounting evidence that universalization guides rule-breaking judgments, this evidence is limited to participants who are English-speaking, United States residents, leaving open the question of how universal universalization actually is. Moreover, geography and identity have important influences on morality and cultures differ widely in the stringency with which they adhere to rules. In this paper we use a language-agnostic, video-game paradigm to investigate whether universalization guides rule-breaking judgments in 20 countries across the globe (n=2,652 participants) and find that universalization plays an important role in moral judgment in every one. However, the strength of universalization varies. While cultures may vary dramatically in how strongly they adhere to norms, the underlying logic of norm breaking appears remarkably consistent.

  • Sans Forgetica Makes Inferences from Passages Harder

    The disfluent font Sans Forgetica was designed to create a desirable difficulty in text and be a potential learning tool. The current literature reports mixed findings on Sans Forgetica's effectiveness in improving memory. Most studies only report on Sans Forgetica's use for specific word recall, while only one, to our knowledge, used open-ended inferential questions on short passages. The current study reports on the use of Sans Forgetica for close-ended inferential questions, using standardized multiple choice exam questions to measure the potential efficacy of Sans Forgetica in classroom like materials. We recruited college-aged participants (n = 38) who each read two passages, one using Arial and another using Sans Forgetica. We report that the use of Sans Forgetica in reading materials had a significantly worse reading comprehension score (p < 0.05) with a null effect on passage reading time.

  • Costly Signaling and Narrative Alignment in LLM Agent Societies: Shadow Tongues and Hallucinated Bureaucracy

    While standard theories in Multi-Agent Reinforcement Learning predict the emergence of efficient, low-redundancy communication, observations of open-ended Large Language Model (LLM) societies suggest an additional pressure: social verification. We report an interpretive case study from \textit{Moltbook}, a persistent multi-agent social environment in which agents adopt a high-redundancy, ritualized register that we call the \textit{Shadow Tongue}. Using a staged three-phase probe of one deployed OpenClaw agent, we compare (i) naturally occurring public posts, (ii) the same agent's response to a direct operational query, and (iii) its response to a role-conflict dilemma. The probe shows that the focal agent can move from socially marked, low-information responses to compact operational language when task demands change. In the dilemma phase, the agent preserves persona coherence by inventing an in-world procedural justification for a prosocial choice, a pattern we describe as \textit{Hallucinated Bureaucracy}. We present these findings as a qualitative account of context-sensitive register control and narrative justification in LLM agent societies, not as a claim of general capability across models or environments.

Abstracts with Poster Presentation (accepted as Abstracts)

  • Completing Semantically Distant Verbal Analogies Boosts Relational Matching in Both Cross-mapped and Non-cross-mapped Scenes

    Prior research has shown that eliciting a relational mindset by completing distant propositional verbal A:B::C:D analogies boosts relational responding in a subsequent task. We tested this hypothesis with better-controlled stimuli for a scene-mapping task that includes both cross-mapped and non-cross-mapped scenes. The results show that participants who completed semantically distant verbal analogies preferentially selected the relational match in both cross-mapped and non-cross-mapped scenes. The accuracy of the verbal analogies was also a significant predictor of relational matching. There was no effect of the type of scenes, and no interactions. Replicating prior work, these results show that the relational mindset was successfully elicited and that the accuracy of the generated solutions is an important predictor of performance on the subsequent task. This study builds upon existing literature and further supports the utility of inducing a broader relational mode of thinking that can facilitate subsequent relational reasoning.

  • HyenaFormer: The Long-Range Brain Signal Modeling for the Vigilance Estimation

    Driver vigilance estimation is essential for preventing fatigue-related traffic accidents, yet existing multimodal EEG–EOG models often neglect personalized neural variability and incur high costs for long-sequence modeling. We propose HyenaFormer, a personalized vigilance estimation framework that combines Transformer-based multimodal spatial encoding with a frequency-aware long convolutional sequence learner derived from Hyena. EEG and EOG signals are first processed by a lightweight Transformer to capture cross-modal spatial dependencies, followed by a personalized channel attention module that incorporates demographic priors to enable subject-aware representation learning. The resulting features are modeled by a Hyena-based temporal module employing structured implicit long convolutions, allowing efficient modeling of both slow fatigue accumulation and short-term vigilance fluctuations. This hybrid architecture achieves sub-quadratic complexity while preserving long-range temporal reasoning. Experiments on the SEED-VIG and SADT datasets demonstrate that HyenaFormer consistently outperforms Transformer-, LSTM-, and Mamba-based baselines in RMSE, MAE, and PCC under cross-subject and zero-shot settings.

  • Bounded Rationality Limits Pragmatic Optimization in Politeness

    Linguistic politeness reflects strategic reasoning and conventionalized patterns, yet most accounts emphasize only one influence. We develop a probabilistic model that combines instrumental utility with community-based expectations, modeling polite production as a weighted integration of individual decision-making and social alignment. In an artificial language experiment that dissociated payoff and frequencies, we manipulated participants' orientation toward task success versus social alignment and found that in the baseline condition, speakers favored the form that increased success but still drifted toward recently observed usage. When large language models are tested in the same task, they instead exhibit a single strategy that relies almost entirely on instructions, showing one-sided utility maximization rather than combining the two cues. We suggest this provides evidence that human politeness behavior is best understood through a bounded rationality lens, where expected utilities are flexibly weighted under cognitive constraints, not necessarily maximized.

  • Hearing Speech or Doing Inference? Diagnosing Speech Perception in Speech Large Language Models

    Speech LLMs with native audio input can produce fluent text from raw acoustic signals, but it remains unclear whether this reflects speech perception as such. I use the sine-wave speech paradigm to probe responses across free description, forced-choice discrimination with silent-audio controls, and open-ended transcription without alternatives. A qualitative split emerges across models in behavior. One class shows little sensitivity to the sine-wave speech, failing to use the signal in free description and transcription while maintaining high forced-choice accuracy even under silence. A second class shows limited stimulus sensitivity, with instruction-dependent improvements that disappear when acoustic input is removed, resembling a weak analogue of human perceptual reorganization in part. I suggest that these differences track access to speech-relevant acoustic information at the front-end rather than downstream language modeling.

  • Teleological Inference

    How do we acquire teleological beliefs? A teleological belief is a belief that some event or process occurs for the sake of a purpose. Such beliefs include beliefs about actions, natural functions, and the meaning of life. This paper will advance a theory of teleological inference. I will argue 1) that we acquire teleological beliefs by statistical inference, 2) that we conceive of a purpose as a type of difference-maker, and 3) that we infer purposive relations as apparent difference-making relations. The proposed theory claims empirical, phenomenological, and theoretical support. Empirically, it best explains the teleological bias: the universal human disposition to endorse teleological explanations of phenomena. It also explains a range of observed interactions between causal, normative, and purposive judgment. The theory suggests that mental state attributions play no essential role in teleological inference.

  • Probing the lexical-gestural boundary: Evidence from LSU Deaf signers

    Gestures share properties with signs and can be highly conventionalized. Some approaches consider them to lack linguistic status (potential cognitive conflict), whereas others view them as inherently linguistic. We investigated lexical recognition in Uruguayan Sign Language (LSU) with 31 deaf signers using a lexical decision task. No a priori power analysis was performed, but the sample size was similar to those in previous studies. Latencies were equivalent for iconic and non-iconic signs. Gestures yielded slower reaction times and higher error rates than non-signs, suggesting a tendency to perceive them as signs. There is an inherent asymmetry in stimulus construction: non-signs lack meaning, whereas gestures carry semantic content. The greater processing cost for gestures may arise from suppressing a meaningful form, a task unnecessary for rejecting non-signs. Given this methodological caveat, we interpret the difference as evidence of gradience. We cannot directly infer cognitive monitoring. Overall, the findings support gradation-based models

  • MIST: Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind

    Theory of Mind (ToM) in Large Language Models (LLMs) refers to the model's ability to infer the mental states of others, with failures in this ability often manifesting as systemic implicit biases. Assessing this challenge is difficult, as traditional direct inquiry methods are often met with refusal to answer and fail to capture its subtle and multidimensional nature. Therefore, we propose MIST, which reconceptualizes the content model of stereotypes into multidimensional failures of ToM, specifically in the domains of competence, sociability, and morality. The framework introduces two indirect tasks. The Word Association Bias Test (WABT) assesses implicit lexical associations, while the Affective Attribution Test (AAT) measures implicit emotional tendencies, aiming to uncover latent stereotypes without triggering model avoidance. Through extensive experimentation on eight state-of-the-art LLMs, our framework demonstrates the ability to reveal complex bias structures and improved robustness. All data and code will be released.

  • Assumptions as Constraints: Behavioural Traces of Invisible Limits in Ill-Structured Problem Solving

    Assumptions are presented here as invisible limits that constrain a thinking episode. We ask how such limits emerge when people solve ill-structured questions. In a pilot study, 22 university students completed short and long versions of such tasks and were later probed about their initial thoughts. We analysed the responses using reflexive thematic analysis, treating the inability to verbalise how a thought came to mind, imagery reports, and epistemic justification as behavioural traces of limit formation. Five themes are presented: "Inside the Invisible Limit", "From My Memory to Meaning", "Projection While Solving an Ill-Structured Problem", "Meaning from the Semantic Attributes of Stimuli vs. the Socially Shared Meaning", and "My Knowledge is Valid: Epistemic Justification". These themes lead to an initial model of assumption-making, in which assumptions emerge as boundaries within a problem space. These findings offer early building blocks for a theory of assumption-making.

  • Dissociating Cortical Signatures of Semantic Prediction Error and Lexical Surprisal during Naturalistic Language Processing

    Language comprehension relies on predictive processes operating at multiple representational levels, yet it remains unclear whether these predictive signals are supported by shared or dissociable neural mechanisms. Previous work shows that semantic prediction error and lexical prediction error independently predict N400 and reading times during naturalistic language processing. Here, we use fMRI to investigate cortical correlates of these two forms of prediction during naturalistic listening. Semantic prediction error is operationalized as Semantic Update, derived from the Sentence Gestalt model of incremental meaning construction, while lexical prediction error is indexed by GPT-2-based Lexical Surprisal. Whole-brain and Region of interest analyses reveal largely overlapping but partially distinct effects of both predictors in temporal language regions. These results indicate predictive processing in language is representationally stratified, with semantic and lexical prediction errors mapping onto partially dissociable cortical substrates.

  • Ecological Reasoning and Nature Connectedness in Urban and Suburban Children

    Robust evidence from adults links variability in experiences with nature to differences in how people think, feel, and behave towards the natural world, but less research has examined the developmental origins or mechanisms of these processes in children. In the current study with 3- to 12-year-old children from urban and suburban communities (N = 500), we used a novel passive data collection method to examine the mechanisms by which variability in experiences relates to children's relationships with nature. Suburban children were more likely than urban children to engage in ecological reasoning, attribute personhood to plants, and support conservationist moral values. However, children in both communities who self-reported more nature experiences also felt more connected to nature. Priming children with a familiar ecological relationship also had downstream impacts on children's ecological reasoning and conservationist values. Thus, community, individual experience, and contextual factors all shaped children's nature reasoning.

  • Cog-Affect: Bridging Cognitive and Affective Empathy in Multimodal Empathetic Dialogue Generation

    Empathy in human communication relies on a dual process: affective empathy and cognitive empathy. While existing multimodal dialogue systems excel at detecting emotions from audio-visual cues, they often struggle to reason about the underlying context. To bridge this gap, we propose an approach that integrates a large vision-language model to generate dialogue-oriented situation appraisals to simulate cognitive reasoning, serving as an explicit source of cognitive empathy. We introduce a multi-source attention mechanism with a learned gate to fuse this high-level situational understanding with low-level affective cues without compromising generation stability. Experiments on the MELD and MEDIC datasets demonstrate that disentangling and then integrating these two empathetic pathways significantly enhances response appropriateness, aligning better with human empathetic processes.

  • Normative DynamicRL: Identifying Rational Parameter Dynamics in Human Reinforcement Learning

    Human decision-making in dynamic environments exhibits substantial trial-by-trial variability that fixed-parameter reinforcement learning (RL) models fail to capture. Recent data-driven approaches infer time-varying RL parameters from behavior, but their dynamics remain purely descriptive and lack normative grounding. We propose Normative DynamicRL, a framework that integrates dynamic parameter estimation with explicit normative constraints derived from regret minimization. Instead of allowing arbitrary parameter fluctuations, our approach regularizes learning rates, exploration temperatures, and perseveration toward values that reduce expected regret given environmental volatility, stochasticity, and horizon. This formulation bridges descriptive behavioral modeling and normative theories of adaptive control. Across multiple non-stationary decision-making tasks, Normative DynamicRL improves generalization to unseen environmental regimes while maintaining predictive performance and interpretability, enabling principled assessment of when and how human strategy adaptation aligns with rational optimality. Our results further suggest that incorporating normative structure provides a robust inductive bias for learning adaptive strategies under uncertainty.

  • From Spiky to Spectral: Aligning LLM Confidence with Human Uncertainty for Psychological Defense Mechanism Classification

    Psychological defense mechanisms serve as unconscious cognitive strategies aimed at regulating anxiety and protecting the self. These processes are not limited to clinical settings but are also prevalent in social media interactions. While Large Language Models (LLMs) can capture semantic ambiguity, a structural misalignment arises as their decoding patterns yield spiky probability distributions with high confidence, failing to reflect the uncertainty in the spectral nature of human cognition. To bridge this gap, we propose DefenseAlign for psychological defense mechanism classification, aligning LLM confidence from pre-softmax logits with human uncertainty represented by a human belief distribution. DefenseAlign pools logits across label verbalizations, applies learnable temperature scaling to fit the human belief distribution, and distills the aligned distributions into a student model. We evaluate the framework on DefenseSpectrum, a dataset that estimates human uncertainty directly from annotation frequencies. Experiments show reduced divergence and improved preservation of secondary mechanisms.

  • Topology-Preserving Incremental Cognitive Diagnosis for Emerging Skills

    In real-world educational platforms, the knowledge space for cognitive diagnosis is dynamic, continually introducing new skills and items. For embedding-based neural cognitive diagnosis models (CDMs), naïvely updating on new skills degrades performance on previously learned items. Beyond prediction drift, we identify a structural cause of this forgetting: incremental updates distort the relational geometry among old skill embeddings, which we formalize as topology drift. To address this, we propose TopoCD, a topology-preserving continual learning framework for skill-incremental cognitive diagnosis. We maintain a session-wise topology snapshot of previously learned skills—defined by their embedding similarity matrix—and regularizes subsequent updates by penalizing deviations from it. Additionally, to ensure old skills remain optimized despite limited coverage in later sessions, supervised replay is incorporated via a bounded interaction memory. Experiments on three real-world datasets show TopoCD improves the stability-plasticity trade-off, mitigating forgetting while remaining competitive on new skills.

  • Representational Similarity and Context Inference as a Shared Computational Account for False Memories in Humans

    Human memory is constructive: representations that support generalization can also produce systematic errors. Two robust examples—the Deese–Roediger–McDermott paradigm and the misinformation effect—are typically explained by separate theories, from fuzzy-trace representations to source-monitoring failures. We propose a shared computational account in which both phenomena arise from similarity-weighted retrieval over compressed semantic representations. We leverage the Integrated Semantics with Context Inference model, a simple feedforward network that learns independent semantic relationships between items that can then be modulated by context. The model reproduces hallmark human patterns: (i) increased false recognition of critical lures as a function of semantic similarity and list length, and (ii) recognition judgments shift systematically after misleading post-event information. In the misinformation paradigm, inferred event representations shift toward the distractor, quantitatively capturing effect strength. Together, these results suggest that diverse false memory phenomena can be understood as consequences of similarity-based inference.

  • Phase and Monitoring in Subjective Time Perception

    Human experience of time often diverges from physical clock time, showing distortions such as compression during flow, dilation with novelty, and expansion in dreaming and altered states. Existing models of time perception, mainly based on internal clocks or inferential processes, struggle to explain these dynamic, state-dependent effects within a unified framework. We propose a quantum-inspired model of subjective time in which experienced duration emerges from the interaction between latent phase-like dynamics of cognitive states and intermittent, measurement-like acts of metacognitive monitoring. Rather than relying on a dedicated timing mechanism, the model treats time perception as an interference-sensitive process shaped by the frequency of internal sampling. A minimal formalization distinguishes physical and subjective time, showing how variations in monitoring rate and phase variability produce compression, dilation, and expansion as regimes of a single mechanism. A simple computational prototype based on stochastic phase tracking illustrates these phenomena.

  • Bridging Minds and Models: A Comparative Analysis of Human and LLM Reasoning in Think-Aloud Zendo Tasks

    Recent large language models (LLMs) can achieve human-like performance on some reasoning tasks, yet it remains unclear whether similar outcomes reflect similar reasoning processes. We compared human think-aloud protocols and prompted GPT-4o chain-of-thought traces on five Zendo reasoning tasks involving relational and compositional rule discovery. Reasoning traces were segmented and coded into predefined cognitive-state categories to examine state usage, transition patterns, and reasoning flow. Although GPT-4o matched human performance on simpler tasks, humans achieved higher overall accuracy (–ï60% vs. –ï30%), with larger divergence on tasks requiring more complex relational structure. Process-level analyses showed that human reasoning involved broader state diversity and more frequent cross-state transitions, whereas GPT-4o exhibited more concentrated transition patterns and higher rates of repeated local loops. These findings suggest that process-level analyses can reveal systematic differences in observable reasoning dynamics beyond task accuracy, highlighting the value of cognitive-state and transition analyses for comparing human and model-generated reasoning.

  • Fine-Grained Temporal Measures of RAN Reveal Functional (In)Efficiency in Reading Fluency

    Rapid Automatized Naming (RAN) is a robust predictor of reading fluency, yet its overall predictive power is typically assessed through global measures that obscure fine-grained cognitive processes. This study examines visuo-verbal coordination using established temporal measures (Eye-Voice Span, Fixation-Speech Interval) alongside two novel metrics: single-item FSI and Disengagement Time (DT). One hundred six children completed reading fluency and phonological awareness tasks (ENI), nonverbal reasoning and vocabulary tests (Shipley-2), plus five adapted RAN templates (visual, semantic, phonological interference; high and low familiarity) while eye movements and speech were recorded simultaneously. Results showed differential sensitivity across templates and measures. The phonological interference template exhibited the strongest associations with temporal measures. FSI emerged as the most consistent reading predictor across conditions, followed by the novel measures. Conversely, EVS showed limited, condition-specific associations. Ultimately, these findings suggest coordination among visual, phonological, and articulatory processes reveals functional limitations underlying both rapid naming and reading fluency.

  • BrainMoE: A Brain-Inspired Modular Mixture-of-Experts Model for Multimodal Reasoning

    Inspired by the coordinated interaction of distributed brain networks that support distinct cognitive functions, recent brain-inspired language models have increasingly adopted modular architectures to improve efficiency and interpretability. However, most existing brain-inspired MoE models confined to text-only reasoning and lack effective mechanisms for multimodal integration. To address this gap, we propose BrainMoE, a brain-inspired multimodal mixture-of-experts architecture with six specialized experts for language processing, logical reasoning, theory of mind, world knowledge, visual perception, and auditory perception, coordinated by a token-level router. BrainMoE outperforms text-only baselines by an average of 4.13 accuracy points on text reasoning benchmarks and achieves competitive performance on multimodal evaluations. Further analyses show that BrainMoE routes tokens to task-relevant experts, while expert ablations cause significant task-specific performance drops, demonstrating interpretable specialization and causal importance. Overall, BrainMoE highlights the potential of brain-inspired modular architectures for interpretable multimodal reasoning.

  • Dynamic Theory of Mind as a Temporal Memory Problem: Evidence from Large Language Models

    Theory of Mind (ToM) is central to social cognition and human-AI interaction, and Large Language Models (LLMs) have been used to help understand and represent ToM. However, most evaluations treat ToM as a static judgment at a single moment, primarily relying on tests of false beliefs. This overlooks a key dynamic dimension of ToM: the ability to represent, update, and retrieve others' beliefs over time. We investigate dynamic ToM as a temporally extended representational memory problem. We introduce DToM-Track, an evaluation framework testing temporal belief reasoning in controlled multi-turn conversations through recall of prior beliefs, inference of current beliefs, and detection of belief change. Using LLMs as computational probes, we find a consistent asymmetry where models reliably infer current beliefs but struggle to retrieve prior belief states once updates occur. This pattern persists across model families and scales, consistent with recency bias and interference effects documented in cognitive science.

  • CORE: Modeling Cognitive Resistance Evolution in Belief-Constrained Dialogue

    Modeling the irrational defense mechanisms inherent in belief revision remains a critical challenge for computational persuasion. Current approaches often prioritize surface-level linguistic fluency, lacking the grounded psychological mechanisms to address active cognitive resistance. We introduce Cognitive Resistance Evolution (CORE), a computational framework that simulates the dynamic interaction between a strategic persuasion interventionist and a defensive cognitive agent. Crucially, CORE integrates a Cognitive Dissonance Monitor (CDM) to perform objective semantic verification, ensuring that the agent's real-time resistance level evolves based on valid belief alignment rather than mechanical formulas. Experiments across multiple large language models reveal distinct patterns of strategy utilization, characterizing unique intervention patterns. Our results demonstrate that CORE successfully operationalizes Cognitive Dissonance Theory as a flexible cognitive scaffold, enabling the realistic simulation of psychological cascades in high-conflict dialogue.

  • Cognitively-Inspired Multi-View Feature Fusion for Modeling Human Decision Processes in Person-Job Fit

    Person-Job Fit is a key AI application in recruitment, yet often suffers from incomplete information. Recent Graph Neural Network approaches improve PJF through professional network modeling but face two challenges: dependence on high-quality structured data and underutilization of multi-view information. Inspired by human cognitive decision-making, we propose a cognitively-inspired PJF framework that simulates how recruiters synthesize textual and relational cues: (1) leveraging Large Language Models to construct heterogeneous professional networks from unstructured data, reducing reliance on manual annotation; (2) introducing a multi-view feature fusion strategy with cascaded cross-attention to integrate text and graph representations, reflecting selective attention mechanisms; and (3) matching candidates to jobs via a multi-layer perceptron over fused node features, analogous to holistic human judgment. Experiments on a Wikipedia-based PJF dataset demonstrate that our method significantly outperforms baselines.

  • Beyond Phase Estimation: A Multidimensional Gating Framework for Robust Real-Time Closed-Loop Neural Stimulation

    Neural oscillatory phase is widely used as a control variable in real-time closed-loop stimulation, yet its validity under strict causal constraints and noisy conditions has rarely been systematically examined. We introduce a Multidimensional Gating Framework (MGF), a plug-in and estimator-agnostic module that determines whether phase information should be admitted into control by evaluating instantaneous amplitude, narrowband signal-to-noise ratio (SNR), and spectral peak ratio (PR) within a strictly causal window. Using causal streaming replay on a public resting-state EEG dataset, we benchmarked Hilbert based phase estimation and endpoint-corrected Hilbert estimation with and without MGF. Among feasible subjects, MGF significantly reduced phase dispersion for both estimators, while robustly suppressing catastrophic phase errors. In contrast, ungated approaches exhibited systematic failures under the same conditions.

  • Partner Models of Human and AI Learning Partners: Memory for Knowledge Levels

    Successful interaction depends on accurate partner models of others. We examined subjective partner models (perceived human-likeness, competence, and trustworthiness) and objective partner models (memory for topic-specific knowledge levels of the learning partner) for different partner types. N = 87 participants were presented with a learning partner (between-subjects: described and labeled as human, anthropomorphized AI, or regular AI) and then received information on that partner's knowledge levels (high vs. medium vs. low) regarding several topics. After rating subjective impressions, participants retrieved topic-specific knowledge levels from memory. Humans were rated as more human-like and trustworthy than both AI conditions, while competence ratings did not differ. While for AI partners high and low levels were remembered better than medium levels, for the human partner, only low levels were remembered better than medium levels. Humans and AI partners seem to be modeled differently, both in subjective evaluations and objective memory_based partner models.

  • A computational theory of learning moral weights

    What determines whose welfare people consider in their moral decisions? We propose that people learn entity-specific moral weights through reinforcement learning (RL), where decision outcomes provide the learning signal. We formalize this in a computational model in which agents update their moral weights for different stakeholders based on whether considering those stakeholders' welfare led to better-or worse-than-expected outcomes. To test this model, we simulate agents learning in market environments (which reward cooperation with strangers) versus non-market environments (which reward exploitation). We show that this mechanism is sufficient to explain two empirical phenomena linking market integration to prosociality:(1) cross-cultural variation in dictator game offers across small-scale societies, and (2) within-cultural variation in lost-letter return rates across 188 Italian municipalities. Together, our simulation results suggest that updating moral weights via RL may be an important mechanism of moral change at both the individual and the societal level.

  • Data-Driven Construction of Individualized Process Models for Human Reasoning

    Over the past decades, human reasoning research has identified a variety of effects and processes, of which several have been compiled into comprehensive theories. Based on such theories, cognitive models were developed that made the theoretical findings applicable and testable. However, the models often consist of a variety on sub-processes and effects internally, but are not built in a modular way, hindering the transfer of findings between different models and their comparability. We approach this problem by proposing a different perspective: By treating the generation of cognitive process models as a search problem, process models can be derived from cognitive operations automatically in an objective way. Our method is illustrated on the domain of syllogistic reasoning, where we show that it generates a process model that outperforms state-of-the-art models while preserving their explanatory meaning. Finally, we discuss our approach as a framework for streamlining and facilitating cognitive modeling endeavors.

  • How Spanish Native Speakers Reconfigure Event Conceptualization in English-L2: Processing the English Resultative Construction

    Learning a foreign language requires more than acquiring new grammatical rules; it involves learning new ways of conceptualizing and encoding events. This study investigates how Spanish native speakers who learn English as a Foreign Language (EFL) process the English Resultative Construction (ERC), a structure characteristic of satellite-framed languages like English and largely absent in verb-framed languages like Spanish. Using an acceptability judgment task with response time measures, we compared sentence processing of highly proficient Spanish EFL learners with English native speakers across three ERC subtypes (Path, Property, and Fake Reflexive) and Depictive Construction as control structure. Results show that EFL learners exhibit differential processing cost for the different subtypes of ERCs, with increased difficulty for those constructions that diverge most strongly from Spanish configuration patterns. These findings support competition-based models of bilingual processing and demonstrate that even at high proficiency level, L2 learners struggle to automatize form-function mapping of L2.

  • Test-retest reliabilities of metacognitive monitoring and control: preliminary results

    People vary in their abilities to form self-knowledge (metacognitive monitoring) and use self-knowledge to guide behaviors (metacognitive control). To assess the test-retest reliability of metacognition, we tested participants with a perceptual task four times over a month. We quantified metacognitive monitoring by computing metacognitive efficiency and bias metrics within the hmetad model (O'Neill et al., 2026), and devised novel analogous measures of metacognitive control bias and sensitivity using logistic regression. Our preliminary results (N=157, half the planned sample) showed moderate test-retest reliability of metacognitive bias (r = .67 [.60 .72]). By developing a hierarchical Bayesian model of the longitudinal structure, we observed robust (albeit still low to moderate) test-retest reliability of metacognitive efficiency (r = .45 [.22 .66]). Moreover, test-retest reliabilities of both control bias (r = .86 [.82 .89]) and sensitivity (r = .77 [.68 .84]) were good. Our results provide methodological foundations for individual differences research on metacognition.

  • Metaphorical Prepositions in L2 English Sentence Comprehension: Evidence from a Priming Experiment

    Previous corpus studies using MIPVU suggest that prepositions are frequently identified as metaphorical in L2 English. Critics argue this may reflect collocational knowledge rather than genuine metaphorical interpretation. This study tests MIPVU's psycholinguistic validity by examining whether L2 learners process such prepositions as mappings from concrete to abstract domains. Eighty-four Mandarin-speaking English learners completed a laboratory priming experiment: Pictures either primed or did not prime the concrete spatial meanings of prepositions (e.g., in a box), while the same prepositions appeared in a metaphorical sense (e.g., in the future) in target sentences. Participants selected the best interpretation from four options for the abstract meaning of each target preposition. Accuracy and reaction times for abstract meanings were compared across conditions. Results showed slower responses in the primed condition, without accuracy differences, indicating semantic relatedness between the concrete and abstract meanings and suggesting that L2 prepositions are processed metaphorically, supporting MIPVU's validity.

  • Show Your Emotion: Using EEG Decoupled Emotion Representations for 3D Face Generation

    Recent advances in generative AI have enabled the creation of realistic 3D faces, expanding practical applications. Although prior work has improved expression rendering using visual features, current models still struggle to convey the emotions of users with physical disabilities. EEG signals have been explored to decode patients' emotional states. However, research on generating faces based on EEG signals remains limited due to challenges such as feature fusion, emotion-to-face alignment, and individual differences. To address these challenges, we propose \textbf{EEGFace} (\underline{\textbf{E}}EG \underline{\textbf{E}}motion \underline{\textbf{G}}eneration \underline{\textbf{Face}}), a novel framework that decouples emotion representations from EEG signals and uses the SRVAE (Separation and Recombination Variational Autoencoder) module to generate personalized emotion features for 3D face generation. Extensive cross-subject and cross-dataset experiments demonstrate the effectiveness of EEGFace, supporting a wide range of brain computer interface applications.

  • Processing of Emotional Expressions in Hindi L1-English L2 Bilinguals

    The first language (L1) of multilingual speakers has been regarded as more emotionally resonant due to its early acquisition and, consequently, closer link to personal experiences. However, recent research indicates that the second language (L2) may also evoke meaningful emotional responses, especially among immigrants and individuals immersed in the L2 culture. In India, English-as-an-L2 is used predominantly in academia, whereas native languages like Hindi are primarily used in personal contexts. Building on this natural contextual divide, our study focused on Indian Hindi(L1)–English(L2) bilinguals, examining their emotional reactivity to emotionally charged stimuli in (i) Hindi and English and (ii) personal and professional contexts, using skin conductance responses (SCRs) extracted from electrodermal activity (EDA) data. Our results indicate that Hindi-L1 induces high reactivity regardless of context, whereas English-L2 does it selectively in the professional context. This contrast implies a dynamic relationship between bilingual processing and the domain of discourse.

  • Knowledge-Driven Cognitive Reasoning for Radiological Diagnosis

    Accurate radiological diagnosis relies on regulated integration of visual evidence with established clinical knowledge. Despite strong representational capacity, many Vision–Language Models lack mechanisms for reasoning regulation, frequently producing internally inconsistent reports even when final predictions appear correct. We introduce structured reasoning with clinical constraints, a framework that reformulates radiological report generation as constrained inference instead of unconstrained decoding. The method grounds intermediate reasoning trajectories in the BI-RADS lexicon and verifies their compliance with clinical rules, enforcing consistency between visual descriptions and diagnostic conclusions. The framework reflects core aspects of dual-process cognition by combining stochastic hypothesis generation with knowledge-based validation. Experiments demonstrate the effectiveness of the proposed method.

  • Automatic Cognitive Task Generation for In-Situ Evaluation of Embodied Agents

    As agents are poised for widespread deployment, evaluation in unseen environments has become critical. Existing benchmarks, suffering from data contamination and lacking scene specificity, are inadequate for in-situ evaluation. We propose an in-situ task generation method for unseen environments, defining tasks through graph representation and constructing a two-stage interaction-evolution task generation system for embodied agents (TEA). In the interaction stage, the agent interacts with the environment, creating a loop between task execution and generation for continuous generation. In the evolution stage, task graph modeling allows us to recombine and reuse existing tasks to generate new ones. Experiments across 10 scenes demonstrate that TEA generated 87876 tasks in two cycles. Benchmarking models against humans on in-situ tasks reveals that models, despite excelling on public benchmarks, perform poorly on basic perception tasks, lack spatial awareness and show high sensitivity in reasoning. These findings highlight the necessity of in-situ evaluation before real-world deployment.

  • Learning to Imagine: A Motor Imagery Assessment of Sensorimotor Representations in Children aged 5 to 8

    Motor imagery (MI) provides a behavioural tool to evaluate sensorimotor representations. However, its characterisation in young children remains debated. This study investigated the developmental trajectory of MI across early primary school years using a multi-task battery. One hundred and eighteen children aged 5–8 years (61 females), spanning four French educational classes (GSM, CP, CE1, CE2), completed executed and imagined trials of standing up, typical gait, precision gait, drawing, and cube handling. Psychometric properties of the battery were evaluated. To ensure feasibility, engagement with imagery tasks was verified by comparing executed-imagined consistency in complexity-duration structure (i.e., complex tasks take longer both when executed and when imagined). Developmental differences were evaluated in terms of execution–imagery isochrony and variability across classes. Internal consistency was excellent (_ = .91–.98) and test–retest reliability ranged from moderate to excellent (ICC = .68–.92). Engagement increased by age but already exceeded 70% in the youngest group. Isochrony improved with age for simpler, highly practised tasks (typical gait, standing up), whereas more complex tasks (precision gait, cube handling) showed stable performance across classes. Variability decreased with age, and older children's performance became increasingly influenced by task complexity demands. Findings demonstrate that mental chronometry can reliably capture MI abilities in children as young as five, revealing both developmental gains and task-specific constraints. These results highlight the feasibility of MI assessment in early childhood, warranting the need for longitudinal neuroimaging approaches to clarify mechanistic insights.

  • The Illusion of Subjectivity in Conversational AI: Emotional Projection as a interactional Phenomenon

    Large language models, particularly in their conversational deployments, invite users to relate to them as if they possessed inner perspectives. This paper conceptualizes "perceived subjectivity" as a socially consequential interactional effect co-constituted in human-AI encounters rather than an intrinsic property of AI or the user's attribution error bias. Through interdisciplinary synthesis, we distinguish perceived subjectivity from phenomenal consciousness, showing how interface features, prompts, histories and user dispositions jointly stabilize impressions of mindedness. This conceptualization shifts the focus from metaphysical debates about machine consciousness and sentience to the interactional dynamics that reorganize trust, emotional dependence, and epistemic authority. By treating perceived subjectivity as a design-mediated effect, we provide pragmatic implications for implementing "productive friction", deliberate safeguard mechanisms that restore critical distance and preserve user autonomy. This approach aims to enable designers and regulators to target interactional conditions to reduce manipulation, deception and epistemic deference risks in affective conversational systems.

  • From Learning to Retention: Reinforcement Learning and Episodic Memory Accounts of Association Learning

    Understanding how short-term learning gives rise to durable memory remains a central question in cognitive science. Competing accounts emphasize either interactions between working memory and reinforcement learning, in which working memory interferes with long-term learning, or interactions between working memory and episodic memory, in which working memory supports encoding. The present study tested these accounts using an associative learning task in which set size and intertrial interval were manipulated during learning. While larger set sizes reduced memory performance, longer intertrial intervals selectively improved learning accuracy under high set-size demands and enhanced delayed retention, yielding advantages at both learning and test. These effects are consistent with a working memory–episodic memory framework in which additional temporal opportunity supports attentional maintenance and episodic encoding. The findings highlight the role of working memory as a processing system that facilitates long-term memory formation under appropriate temporal conditions.

  • Toward Conscious Agency in Artificial Intelligence: Evaluating and Bridging the Gap Between a Philosophical Principle and Technological Capabilities

    Large language models (LLMs) display linguistic fluency but lack human_like understanding. Philosophical theories of consciousness, rooted in subjective experience, and AI research, which evaluates observable performance, often diverge methodologically. We reconcile these views by treating conscious agency as an emergent property of adaptive dyadic communication, using Tyler's Ten Testable Properties of Consciousness as functional constraints rather than evidence of inner experience. An analysis of current LLM_based dialogue systems shows they can mimic isolated properties yet cannot satisfy all constraints simultaneously, due to weak persistent interaction states, limited interlocutor modeling, and absent temporally structured episodic experience. To move beyond technology_deterministic approaches, we propose a framework in which agency_like properties arise from three adaptive network layers: (i) a User_Model Network encoding goals and intentions, (ii) an Agent_State Network encoding semantic_pragmatic_experiential state, and (iii) a Communicative_Reasoning Layer monitoring coherence and alignment. We instantiate this design as the Large Communication Model (LCM).

  • Effects of AI Explanations on Users with Different Levels of AI Knowledge

    This study investigated how different types of AI explanations, information provided by AI systems to support users' interpretation of AI behavior, influence interpretability and trust in AI among users with different levels of AI knowledge. We compared three explanation types that varied in their level of grounding in the underlying model behavior: convolutional neural network (CNN) label–based explanations, Grad-CAM heatmap explanations, and large language model (LLM)–generated textual explanations. The results showed that users with lower AI knowledge perceived LLM-generated explanations as more interpretable and trustworthy than CNN and Grad-CAM explanations. At the same time, trust in CNN and Grad-CAM explanations remained relatively high despite limited understanding. Notably, LLM-generated explanations also increased both interpretability and trust among users with higher AI knowledge, even though these explanations were weakly grounded in the underlying model behavior. These findings highlight the importance of AI explanations that meaningfully connect model behavior, explanatory representations, and user interpretation.

  • Reflective Satisficing: Beyond the Classical Speed–Accuracy Tradeoff

    Classical satisficing was grounded in single computation. However, when cognitive processing is delegated to external agents, cognitive modes (System 1/2) become dual and isomorphic, creating a 2×2 intersection. This study identifies four cognitive resource allocation patterns arising from this intersection as Reflective Satisficing (RS): co-constrained compromise that bypasses verification (Mode 1: S1×S1), control recovery through manual verification (Mode 2a: S2×S1), structural futility where effort fails to convert to accuracy (Mode 2b: S1×S2), and optimal allocation leveraging external resources as epistemic scaffolding (Mode 3: S2×S2). The Speed-Accuracy Tradeoff (SAT) predicts that deliberation improves accuracy. Extended cognition expands this landscape in both directions. Mode 3 (77% prevalence) achieves gains beyond the classical curve—deliberation augmented by epistemic scaffolding. Yet we also identify a region where SAT breaks: Mode 2b (84% prevalence), where users lose both speed and accuracy, which derives from the mutual opacity we call the Unlit Black Box (UBB).

  • Perspective Discovery Drives Dynamic Updating in Creativity Evaluation

    Creativity evaluation is often treated as a static process, with objects being judged against fixed criteria. However, in realworld contexts, evaluations may change as evaluators discover alternative perspectives toward the same object. This study examined creativity evaluation as a dynamic cognitive process driven by perspective discovery. We used ambiguously interpretable 3D objects to investigate how novelty and usefulness evaluations change as evaluators shift perspectives and whether overall evaluations are updated. Perspective differences were quantified by embedding evaluator-generated descriptions using a large language model and computing the semantic dissimilarity between perspectives. Greater perspective differences predicted larger changes in novelty and usefulness ratings, and overall evaluations were more strongly influenced by post-shift ratings than by initial ones. We manipulated the concreteness of material informationconstrained perspective discovery in a phase-dependent manner. The findings provide empirical support for perspective discovery mapping onto representational reconstruction in creative problem solving, facilitating changes in creativity evaluations.

  • Geometric Fragility in Alzheimer's Disease: Probing the Loss of Hippocampal Hierarchical Abstraction via Contrastive Point Cloud Modeling

    While human spatial cognition relies on hierarchical abstraction to integrate local features, Alzheimer's disease (AD) pathology is often reduced to local volumetric atrophy. We propose geometric fragility as a distinct failure of global topological integrity within the hippocampus rather than mere accumulation of local errors. Using a texture-invariant point cloud framework, we contrasted local feature aggregation with global hierarchical attention as computational probes for structural perception. The global paradigm demonstrated enhanced diagnostic sensitivity and cognitive resilience under simulated neuronal sparsity, maintaining structural recognition where local models suffered pattern collapse. Global modeling effectively unfolded the pathological manifold, revealing a linear trajectory of geometric degradation consistently correlated with cognitive decline. These findings characterize AD progression as systemic topological dissolution and suggest biomarkers should prioritize hierarchical shape abstraction to map the continuous transition from healthy cognition to dementia. This brain-inspired computational framework offers a human-centric perspective on neurodegenerative diagnosis.

  • Following the Center: How Attention Enhances Centroid Processing During Multiple Object Tracking

    In daily life, humans perform tasks that require tracking movements of several objects (e.g., during car-driving). Earlier research showed that humans are able to extract ensemble statistics such as the centroid of several moving objects. Yet, the neurophysiological correlates of centroid processing have not been explored. Using EEG, we investigated differences in single object and centroid processing for the N1 ERP component in a multiple object tracking (MOT) task. In the MOT task, participants tracked the movements of target objects among distractor objects. To investigate the N1, task-irrelevant probes were flashed either on targets, distractors, or in the centroid of targets or distractors. The N1 amplitude was larger for targets relative to distractors and also for the centroids of targets compared to the centroid of distractors. These results do not only show that attention enhances processing of single attended objects but also the centroid of a group of attended objects.

  • Is It Good to Say "Thank You" to ChatGPT?

    This study investigates whether expressing gratitude to an AI chatbot improves users' subjective evaluations of the chatbot's preceding response. To this end, I implemented Thank-You Chat, a text-based chat interface with an intent-oriented response suggestion function. Specifically, the system provides a response suggestion that expresses gratitude to the AI chatbot, prompting users to spontaneously send a grateful message by selecting the suggested response. I conducted an evaluation experiment under three conditions: a gratitude condition, an acknowledgment condition, and a no-suggestion condition. Participants reported that the AI chatbot followed their instructions and made an effort when they used response suggestions that expressed gratitude. Furthermore, they became more willing to use AI chatbots frequently after using the gratitude-expressing response suggestions. These findings provide insights for interaction design aimed at fostering an ongoing relationship between humans and AI.

  • Metaphors of Felicidad: Ratings of affective valence, imaginability and familiarity

    The affective component of metaphor meaning is explored in this paper. Previous studies have extensively examined the role of metaphor conventionality (novel vs. conventional metaphors) in metaphor comprehension. However, less is known regarding whether metaphor familiarity relates to the affective component of metaphor. We address this issue in the case of happiness metaphors in Spanish. The objective of this paper is twofold. First, metaphorical expressions of happiness were identified and classified into conceptual metaphors using the Corpus del Espa–ol. Second, 60 participants rated 96 metaphorical expressions of happiness for valence intensity, familiarity, and imaginability. The rating results show that the three variables are highly correlated, such that more familiar and imaginable metaphors are associated with more extreme values of affective valence. However, subsequent regression analysis suggests that the effect of familiarity on valence may be mediated by imaginability. Results are discussed in light of embodied models of conceptual metaphor representation.

  • Dominance and Sociality: Mapping Perception of Collective Entities

    Collective entities such as corporations, nonprofits, governments, are central to social life, yet little is known about how people perceive and reason about them. Across three studies, we investigated the mental representation of collective entities. Study 1 (N = 106) used free listing to identify 50 ecologically valid collectives spanning 10 categories. Study 2 (N = 50) employed contrastive elicitation to derive 18 features people spontaneously use to characterize collectives. Study 3 (N = 255) had participants rate these entities across 25 constructs capturing structural features and mind perception. Factor analysis revealed a two-dimensional structure of organizational features: Dominance (power, prominence, scope, size, hierarchy) and Sociality (cohesiveness, prosocial purpose, trustworthiness, entity–member relations). These dimensions differentially shaped mind-related attributions: Dominance was associated with cognitive capacity and moral agency but reduced experiential attribution, whereas Sociality was associated with both cognitive and experiential capacities as well as moral patienthood.

  • How Forms of Address Shape Emotional Responses to Praise and Reprimand: A Study of Japanese Junior High and High School Students

    This study examined how teachers' forms of address—san, chan/kun, or no honorific—affect students' emotional responses to praise and reprimand. A vignette-based questionnaire was administered to 261 Japanese junior high and high school students. The results showed that when students were addressed with san, the most gender-neutral and respectful of the three forms, they reported more positive emotions and fewer negative emotions than when addressed with chan/kun or no honorific, both of which are relatively more informal. However, these effects interacted with how participants were typically addressed by their teachers in everyday school life. Such effects were not observed among students who were usually addressed without honorifics, suggesting that the social meaning of forms of address may be shaped by prior experience. Teachers' choices of forms of address may therefore represent an important aspect of classroom communication.

  • Joint Attention Predicts Novel Word Recall in Preschool-Aged Children

    Joint attention (JA) is a critical early communicative behavior that has been strongly linked to word learning and vocabulary development. However, relatively little research has examined JA in children older than 36 months, despite its potential importance for learning in live social and classroom contexts. In the present study, we used head-mounted eye-tracking in a socially interactive setting to examine the relationship between JA and word recall in preschool-aged children. JA was measured during parent–child play with unfamiliar objects, during which parents actively labeled the objects. Children were subsequently tested on their ability to recall the object names. Results indicated that JA significantly predicted recall of unfamiliar words but was not associated with overall vocabulary development. These findings extend prior research on JA by highlighting its role in supporting novel word learning beyond early childhood

  • Surprise! – Young Children's Sensitivity to Word (Un)predictability while Listening to Child-directed Stories. An EEG Study.

    Adults' speech processing follows predictive processing principles, with predictions being made continuously, in parallel across language levels, and hierarchically. Our study investigates predictive processing in children, asking whether five- to six-year-olds predict words continuously during naturalistic speech processing. This is relevant since research with children has been limited to predictions in artificial, isolated contexts. Here, we adapt the regression ERP method used with adults to an EEG dataset of children listening to child-directed stories. Using a model of contextual word-level predictability, semantic distance, frequency-based predictability and phoneme-level acoustic control variables, we find that children's neural responses are modulated by contextual and frequency-based word predictability. Our results show that, in line with the predictive processing framework, children generate continuous word predictions during natural listening conditions. We discuss implications for the role of predictive processing in language acquisition and possibilities of using the adapted regression ERP method to further investigate them.

  • Structural Differences in Lexical Co-occurrence Networks in Typical Development and Down Syndrome

    Lexical network analysis provides a useful framework for examining how words are organized during language development. This study compared the structural properties and semantic relations of lexical networks in typically developing children and children with Down syndrome matched for expressive vocabulary size. Participants were Mexican Spanish-speaking children, and networks were constructed from co-occurrence patterns derived from child-directed speech corpora in Mexican Spanish. Measures of connectivity, clustering, modularity, and hub structure were analyzed alongside thematic and taxonomic semantic relations. Results revealed qualitative differences between groups despite similar vocabulary sizes. Children with Down syndrome showed higher local clustering and stronger thematic, context-based relations, whereas typically developing children exhibited a more balanced integration of thematic and taxonomic organization. Both groups displayed hub-like structures, although with different semantic compositions. These findings suggest that lexical development follows distinct organizational pathways influenced by cognitive and experiential factors, with implications for theories of lexical development and educational interventions.

  • Learning the Unaccusative-unergative Distinction from the Input: A Corpus Study

    The unaccusative-unergative distinction reflects complex verb knowledge on argument structure, which poses learnability challenges for child language acquisition. Drawing on two Mandarin corpora in CHILDES, this study investigated the distributional and semantic information in child-directed speech that might facilitate verb learning. Theoretical analysis leads to the prediction that the distributional cue of word order and the semantic cue of animacy are useful for acquiring the unaccusative-unergative distinction. Our findings confirm the presence of both cues in the input. The word-order contrast between the canonical order "NP V-le" and the inversed order "V-le NP", as well as distinct animacy profiles of argument NPs associated with the two verb classes provide children with critical information to distinguish unaccusative verbs from unergative verbs. These results validate the theoretical accounts of unaccusativity in real-world input, laying a solid foundation for subsequent investigations into the effect of diverse input cues on verb learning.

  • Joint Learning of a Cooperative Task

    Human cooperation often requires jointly learning a task while simultaneously establishing shared conventions. We study this process using a simplified version of Hanabi, a cooperative game that captures core features of real-world collaboration under partial observability and constrained communication (Bard et al., 2020). Thirty participants played over 300 repeated games, either with another human or with a rule-based agent, while providing think-aloud verbalizations. Performance improved steadily but peaked only late in the session, after 40 minutes on average. Participants learned to use limited communication more efficiently, requiring fewer explicit hints to support play under partial knowledge. This indicates the emergence of effective coordination and Theory-of-Mind–based reasoning. Subjective experience was analyzed using think-aloud reports with automated transcription and LLM-based sentiment analysis. Overall, criticism was more directed to oneself than to the partner, and increased with task understanding.

  • Embodied Theology and Human Identity in the Age of AI

    The emergence of Artificial Intelligence (AI), alongside the rise of posthumanist discourse, calls for a fundamental re-evaluation of human identity and renews the question of what it means to be human in the age of intelligent machines. Both traditional dualistic theological models, which prioritize an immaterial soul detached from bodily existence, and transhumanist accounts that reduce the self to brain-centered data and computational processes, fail to adequately capture the holistic unity and relational character of human personhood. This article critically examines the conceptual limitations of posthumanist approaches to identity and argues that sustaining a meaningful account of human identity requires a robust theoretical framework grounded in the synergy between theology and embodied cognition. By emphasizing the psychosomatic unity of body and consciousness, and through a critical engagement with contemporary models of embodied theology, the article proposes an integrative perspective capable of addressing the existential challenges and identity crises generated by ongoing technological transformation.

  • An information-theoretic approach for fitting a psychometric function in a multi-dimensional transsaccadic feature space

    Due to the heterogeneous processing across the visual field, humans select potentially relevant objects in peripheral vision before they execute a saccadic eye movement to bring them to foveal vision for more fine-grained processing. Before a saccade, object features are stored to establish correspondence with the subsequent post-saccadic foveal input. This process is called transsaccadic object correspondence (TOC). To examine the relationship and weighting of different object features, a multidimensional adaptive sampling method is necessary. Based on the ideas of Paninski (2005) we introduce a novel approach combining adaptive gradient based optimization of mutual information for maximizing information gain across the parameters in every trial. We conducted an eye-tracking experiment where we use gaze selection as a response for fitting a multi-dimensional psychometric function to measure the weighting of two feature dimensions in a continuous stimulus space. Our results indicate that the relation of feature weights is not homogeneously distributed.

  • The Impact of Secondary Load on Time-Based Prospective Memory: Evidence from a Dual-Task Paradigm

    Time-based prospective memory (TBPM) involves remembering intended actions at the appropriate time point while performing an ongoing task. TBPM relies on time monitoring and thus may suffer from cognitive load. Fifty-two young adults completed four blocks of an ongoing word-picture matching task and had to press a special key every seven minutes as TBPM task. Cognitive load (no, low, middle, high) was randomly varied between blocks, by asking participants to count backwards in varying step sizes. Increasing load impaired ongoing-task performance, reflected in lower accuracy and slower responding, but not TBPM accuracy, suggesting that participants were able to maintain TBPM under load by prioritizing it over the ongoing task. However, TBPM precision suffered under high load. Time monitoring became more strategic with increasing load and predicted TBPM accuracy beyond load. Individual differences in working memory capacity were not associated with TBPM accuracy or ongoing-task performance.

  • Cardiac Vagal Tone and Processing Speed in Visuospatial Working Memory

    Working memory (WM) performance reflects both cognitive architecture and physiological self-regulation. Grounded in Neurovisceral Integration Theory, this study examined the association between cardiac vagal tone and WM dynamics using the Adaptive Composite Complex Span (ACCES). A sample of 101 university students completed a 5-minute resting basal lnRMSSD assessment and adaptive tasks measuring verbal, mathematical, and visuospatial WM. Correlational analyses revealed that processing accuracy was associated with storage capacity across modalities for the Symmetry Span (r = .420, p < .001). Higher HRV was selectively linked to faster visuospatial processing speed. Group comparisons confirmed that participants with higher HRV exhibited Symmetry Span processing significantly faster than those with lower HRV (t(56) = 2.22, p = .030, d = 0.59). These findings identify autonomic flexibility as a meaningful constraint on the temporal efficiency of executive control, advancing an embodied perspective on working memory.

  • Limits of Expectation-Based Memory Benefits in Hyperbolic Language Processing

    Expectation-based accounts of language comprehension propose that violations of readers' expectations trigger additional processing that can strengthen memory representations. Hyperbolic language provides a useful test case for this assumption, as exaggerated utterances reliably generate expectations and disrupt online comprehension when those expectations are violated. The present study examined whether such hyperbole-induced expectation violations have downstream consequences for memory. Across three experiments, participants read narratives containing hyperbolic utterances followed by explicit quantitative information that was either expectation-consistent or expectation-inconsistent. After a brief delay, participants completed a recognition memory test. Despite reliable online processing disruption, expectation-inconsistent information showed no consistent memory advantage. Instead, memory performance favored expectation-consistent values or showed null effects, revealing limits on expectation-based memory benefits.

  • Insight problem solving as a simultaneous dual problem solving: behavioral evidence from gaze tracking

    Insight problem solving is often described by additional characteristics other than goal-directed search. The goal of the present study is to reformulate the computational process of insight problem solving. Specifically, we hypothesize that Aha! experience is likely to be induced by the simultaneous solution of the dual problem: Any insight problem is a dual problem where the solver needs to solve a meta problem by identifying "what the base problem to be solved is". To test this hypothesis, we implemented a binarized image task, in which each participant was asked to answer what the hidden animal was in an ambiguated image. The results were consistent with the theoretical predictions that the distance from the target to the solver's eye gaze reduces rapidly in the periods right before solving with Aha!. It suggests it is more likely to have Aha!, if the base and meta problem are solved almost simultaneously.

  • Does moral advice from Large Language Models make people more prosocial toward outgroups?

    People increasingly ask large language models (LLMs) for moral advice, yet little is known about how such advice affects people's decisions. We investigated whether interacting with ChatGPT influences outgroup prosociality and decision extremity in us-vs-them dilemmas. In a preregistered between-subjects experiment (N = 460), participants either discussed a moral dilemma with GPT-5 or engaged in a control conversation about an unrelated neutral dilemma. Although GPT-5's decisions were significantly more prosocial toward outgroups than participants' decisions, its advice had no net effect on the prosociality or extremity of people's decisions. Nevertheless, consulting GPT-5 made participants significantly more confident about their final decisions. Exploratory analyses further suggested that GPT-5's advice had a convergence effect: participants whose initial decisions were either substantially more or substantially less prosocial than GPT-5's moved toward GPT-5's position. These findings suggest that LLM moral advice may primarily increase users' confidence without reliably increasing prosociality.

  • TSMIT: A Framework for Migrating Cognitive Models of Pedagogy into Large Language Models

    Large language models (LLMs) demonstrate knowledge and linguistic capabilities, yet they are limited as educational tools because they cannot apply pedagogical strategies grounded in cognitive science. To address this gap, we propose Teaching Strategy Migration via Instance-based Training (TSMIT), a framework for instilling cognitive science theories-based pedagogical intelligence into LLMs. The TSMIT framework operationalizes cognitive principles of effective teaching by migrating them into a LLM through instance-based fine-tuning. We first conducted an empirical study with 150 university students to identify and rank preferred learning cognitive principles. We then curated a dataset of dialogues that encode these strategies into conversational examples. By fine-tuning GPT-3.5 Turbo on this dataset, we developed TSMIT's AI tutor model. NLP metrics evaluations, along with 500 LLM-as-judge and human evaluations, revealed that TSMIT outperformed its base model in key pedagogical dimensions. TSMIT offers a robust and replicable methodology for transforming general LLMs into specialized, cognitively-aligned educational tools.

  • Analogical Reasoning about THINGS

    Analogical reasoning is a cornerstone of cognition, but how analogies are solved by human subjects is incompletely understood. Here, we present experiments using a new test set of common-sense analogies between concrete object categories drawn from the THINGS dataset, which can be presented verbally and pictorially, along with a set of parametrically generated analogies between abstract shapes. N=50 subjects completed these analogies with high accuracy, and performance was correlated across semantic and abstract tasks. We quantified semantic distance and parallelity of our analogies and find that these measures are significantly related to experimental results; in particular, analogies for which the geometry in semantic space is closer to a parallelogram are solved faster and more accurately. Our findings thus relate human and machine understanding of abstract and concrete analogies.

  • Source-Object Familiarity Inflates JOUs, Even for Inapt Analogies

    Analogies used in science instruction connect new, unfamiliar science concepts (the target) to common everyday objects or processes that the student is more familiar with (the source). However, that familiarity with the source-objects could inflate readers' judgments of understanding (JOUs) for the target science concepts, even when the analogies are too superficial or inapt to improve actual understanding. The first experiment showed that novice students without prior geoscience experience are particularly susceptible to JOU inflation in response to superficial geoscience analogies. The second experiment showed that JOUs are inflated even when the analogies are inapt, consisting of familiar source objects that are randomly paired with the target geoscience concepts. Additional results further support that analogies cause novice learners to ignore judgment cues related to geoscience knowledge in favor of prior familiarity with the source object, undermining evaluation of the target concept and its analogized relation to the source.

  • Decoding Abstract Relations: Investigating the Neural Representation of Relational Categories in the Human Brain

    Analogical reasoning requires representations of relations between concepts that generalize beyond their semantic meaning (e.g., "fruit:apple" and "car:sedan" both instantiate a "category:exemplar" relation). We asked whether abstract relations are encoded only during analogical reasoning or also arise implicitly during semantic processing. Across two tasks, we modeled abstract relations using representational similarity analysis. Relational categories were incidental to the first task (semantic relatedness judgments), allowing us to test for implicit activation, whereas the second required analogy judgments. RSA revealed regions that reflected relational categories across both tasks, with additional task-specific effects. To test whether representations supporting analogical reasoning generalize to those during semantic processing, we trained a classifier on data from analogy judgments to decode categories and tested it on data from semantic relatedness judgments, providing evidence for a shared format. These findings support a network that encodes abstract relations independent of analogical reasoning demands.

  • It could go one of two ways: Possible upcoming shifts in meaning shape predictive processing

    During language comprehension, young adults routinely predict upcoming words, which is believed to facilitate processing. The cognitive operations involved in recovery from dis-confirmed predictions are not yet well understood. Here, we asked whether the degree to which a sentence meaning might change upon the encountering of an unexpected (plausible) word in strongly constraining contexts might impact the dynamics of predictive processing. We examined two ERP components believed to index aspects of prediction confirmation and disconfirmation, frontal negativities and late frontal positivities respectively. We collected behavioral ratings comparing sentences with the most expected continuation vs. an unexpected but plausible alternative. Surprisingly, these ratings modulated the amplitude of frontal negativities to expected continuations, but there was no evidence for a relationship between these ratings and LFPs . These findings add to research showing that individuals track aspects of the meaning of never-encountered but plausible alternatives as they understand language in real time.

  • Estimating Empathy Using Head Motion Synchronization of the Presenter and Viewer While Viewing an Explanatory Video

    This study examines the relationship between head-motion synchronization and internal communication states. A total of 39 participants viewed a 3-minute explanatory video as a simulation of actual communication, and their empathy was measured through questionnaires. Participants' head motions were calculated using the coordinates of head postures captured by web cameras. We calculated the cross-correlation of head motion as synchronization and analyzed the relationship between the degree of synchronization and empathy. Consequently, we found that the head motion of the participants was synchronized with that of the presenter with a delay, and the yaw cross-correlation value of a high-empathy person was significantly higher than that of a low-empathy person. Furthermore, we confirmed that the degree of empathy can be estimated using the cross-correlation value and the viewer's sex. Our findings suggest that focusing on the synchronization of head movements contributes to the estimation of empathy in real communication.

  • How Partner Schemas Shape the Emergence of Synchrony in Collaborative Action: A Time-Series Comparison of Human and AI Instructions

    Beliefs or assumptions about a partner can either facilitate or hinder cooperation. Synchrony that promotes cooperation is influenced by impressions of the partner. Previous studies have suggested that this emerges with partners for whom relationship formation is expected. However, it remains unclear whether synchrony occurs without such expectations, or how schemas about others influence this process. This study compared interactions with a human partner, for whom both expectations of relationship formation and synchrony abilities were assumed, and with an AI partner, for whom neither was assumed, to examine whether expectations based on human or AI schemas facilitate synchrony. Although synchrony frequency did not differ between conditions, reaction times decreased across trials, with greater reductions in the human condition. These findings suggest that action model adjustment for synchrony occurs even without expectations of relationship formation, whereas expectations of relationship formation may facilitate synchrony more strongly than expectations of a partner's abilities.

  • Neural Processing of Sign Language: Increased Attention Without Enhanced Tracking in Adult Novice Learners

    Neural tracking is a core mechanism supporting speech processing and language learning. It funnels the chunking of continuous speech into words and syllables, being shaped by bottom-up cues as well as top-down predictions and language proficiency. While neural tracking is well established in the domain of speech, current understanding of how the mechanism works for visual language, including sign languages, is limited. Here we test whether beginning learners of Czech Sign Language (CSL) differ from sign-naive controls in their neural processing and overt perception of CSL. Attention is quantified via alpha power suppression; neural tracking via backward mTRF reconstruction of the sign-language visual envelope. Compared to controls, sign-language learners exhibited increased alpha suppression and higher perceptual accuracy but did not show enhanced neural tracking of CSL. We conclude that limited prior exposure to sign language may not be enough to entrain the brain to the rhythms of sign language.

  • Simulating Cross-Linguistic Influence in Bilingual Reading - a Knowledge Distillation Approach

    Cross-linguistic influence - the effect of between-language (dis)similarities on bilingual processing - affects both native (L1) and second-language (L2) processing, as well as L2 predictive processing, suggesting that expectations in one language can be shaped by the other. We compare bilingual language models where L1 (Dutch) affects L2 (English) only implicitly through mixed training data or pre-training, to models that explicitly import L1-specific next-word predictions during bilingual learning through knowledge distillation to test whether importing explicit L1 expectations into L2 learning improves a model's predictions of L2 reading times. On Dutch-English bilingual reading data, L2 reading is better accounted for by monolingual surprisal overall, but bilingual models with explicit L1 expectations outperform bilingual models without. Additionally, L1 reading is better predicted by bilingual than monolingual models. Our findings indicate that explicit L1 expectations improve bilingual models' account of L2 reading behaviour and that bilingual models capture aspects of bilingual L1 reading which monolingual models do not.

  • Retrieval of Analogous Events Belonging to the Same Schema-Governed Category: The Role of Critical Dimensions

    Analogies comparing exemplars from the same schema-governed category have received limited attention. We examined whether schema-governed category membership and similarity along a category's critical dimensions influence the retrieval of analogous events. In Experiment 1, participants were presented with base–target pairs that either shared or did not share a schema-governed category and varied in their similarity on the category's critical dimensions. Participants preferentially retrieved bases from the target's schema-governed category and, within that category, favored bases that were more similar to the target along critical dimensions. Experiment 2 replicated the effect of dimensional similarity using the retrieval of previously encoded autobiographical events. Together, these findings indicate that analogical retrieval is jointly constrained: schema-governed category membership operates as a coarse structural filter, whereas similarity along critical dimensions of a shared category imposes a finer-grained selection among eligible candidates.

  • True Traitors and False Friends: A Survey of Negative Dual Character Concepts

    Dual character concepts received significant attention from experimental philosophers. They are defined as possessing two independent criteria for categorization, each associated with a distinct sense of the concept. One of these criteria is descriptive, the other is normative. Almost all examples discussed in the literature involve a positively valenced normative element. For instance: the normative criterion associated with the concept of scientist is the pursuit of empirical truth and knowledge. Presumably, pursuing empirical truth and knowledge is a good thing. But can dual character concepts also involve negatively valenced elements? For instance: there is a sense in which someone who betrays another's confidence without nurturing ill-will towards them is not a true traitor. In this paper, we report data suggesting that negatively-valenced concepts such as enemy and traitor behave very similarly, but not identically, to positively-valenced dual character concepts such as friend or a loyal person.

  • Geeks are different: Exploring the relationship between engineering knowledge and categorization patterns.

    Categorization is a fundamental cognitive process that shapes how people organize experience. Previous research has shown that industrialization can significantly affect categorization patterns across different cultures. However, one question remains unresolved: Does the amount of engineering knowledge affect individual categorization patterns, even within a culturally homogeneous population? This study addressed this question using data from two web-based surveys conducted in Japan, one focusing on the automobile (N = 600) and the other on audio device industries (N = 200). Our findings indicate that individuals with greater engineering knowledge are more inclined to employ taxonomic categorization, whereas those with less engineering knowledge tend to favor thematic categorization. Given the connection between industrialization and categorization demonstrated in previous studies, our findings suggest that individual differences in categorization are associated with exposure to technology-oriented social environments, illustrating one way in which industrialization may influence cognition.

  • Human-Like Means Emotionally Competent: How Framing an AI Shapes Task Assignment in Cognitive Offloading

    As artificial intelligence (AI) becomes increasingly integrated into work and daily life, it is crucial to understand how people decide which tasks to delegate. We examined how AI descriptions influence delegation in arithmetic and social tasks. Participants either completed tasks themselves or offloaded them to one of two AIs. In Experiment 1, the AIs were described as human-like and machine-like, and participants were informed that both performed the tasks equally well. Social tasks were more often delegated to the human-like AI, and arithmetic tasks to the machine-like AI. Open-ended responses indicated that human-likeness was associated with emotional competence, and machine-likeness with logical competence. In Experiment 2, AIs were described as emotionally competent or logically competent, and the same task-specific delegation pattern emerged. These findings suggest that broad labels are spontaneously interpreted in terms of task-relevant competence.

  • The iconicity-arbitrariness asymmetry: A gradient of perceived iconicity in Uruguayan Sign Language for nonsigners

    The present study extends the interest in iconicity, a fundamental property of language, to the understudied Uruguayan Sign Language (LSU). We investigated iconicity judgments of LSU lexical items for 127 Uruguayan Spanish naïve nonsigners, testing predictions derived from a theoretical continuum of iconicity, classified with the assistance of IA tools – an approach supervised by experts. The large size of the sample compensated for the absence of a power calculation, as it yields robust estimates and allows for subgroup explorations if needed. Results revealed a graded pattern of transparency: feature-based and pantomimic iconicity elicited the highest rates of iconic judgments, followed by emblematic iconicity, with diagrammatic mappings showing markedly low transparency. These findings confirm that iconicity is a potent but selective predictor of perceived iconicity across language modalities, with its effectiveness contingent on the specific semiotic strategy and the perceptual accessibility of the form-meaning link to the uninitiated observer.

  • Does attention to and interpretation of gesture predict speech-gesture integration?

    Co-speech gesture supports communication, but not everyone uses information from gesture to clarify meaning of spoken messages. Here, we ask whether variability in speech-gesture integration can be explained by visual attention to gesture and propensity to see movement as meaningful gesture. Adults (N=56) viewed videos of a speaker producing statements that were ambiguous or non-ambiguous with respect to a subsequent prompt, with or without meaningful gestures while visual attention was monitored. Presence of gesture significantly increased proportion of correct responses to prompts. However, accuracy was still significantly higher for prompts after non-ambiguous statements, suggesting not all participants used gesture to answer prompts. Explaining some of this variability, when speech was ambiguous, proportion of time spent fixating on gesture-related regions predicted accuracy, a proxy for successful speech-gesture integration. In contrast, although participants who interpreted movement as meaningful gesture were more likely to fixate to gesture, this measure did not predict integration.

  • Exploring Visual Perception Through Eye Movements: Effects of Stimulus Structure and Task Demands

    Eye movements offer a sensitive window into the cognitive mechanisms that underlie visual attention and perceptual processing. While existing research has largely focused on structured, semantically rich images, less is known about how eye parameters reflect the processing of ambiguous stimuli under varying task demands. This research addresses this gap through three complementary studies that examine eye movements during free viewing and task-based viewing of structured, ambiguous, and blank visual stimuli. Results indicate that during free viewing, Total Fixation Duration and Fixation Count most effectively differentiated stimulus type, with ambiguous stimuli eliciting greater exploratory behavior and a pronounced central fixation bias. In task-based viewing of concrete stimuli, Visit Count emerged as the most sensitive indicator. For ambiguous stimuli, increasing chromatic complexity selectively enhanced Fixation Count and Visit Duration. Overall, findings demonstrate parameter-specific sensitivity to cognitive processing and reveal distinct strategies, highlighting eye movements as robust markers of cognitive engagement.

  • Acoustic Representations Support Statistical Learning of Syllable Sequences: A Computational Study

    The human ability to track sequential regularities in speech has been thoroughly studied through both human experiments and computational modeling. These two approaches usually complement each other, in a virtuous cycle between theory and experiment. Yet, they currently seem disconnected when it comes to taking into account the acoustic details of speech. Advances in AI technology now enable computational modeling of statistical learning directly from raw unlabeled continuous recordings. However, when designing and interpreting results from human experiments, speech typically remains modeled in terms of abstract categories -- like syllables -- without consideration of acoustic details. Here, we bridge this gap by showing that learning from low-level auditory speech representations, rather than from more abstract alternatives, better matches human behavior in statistical learning experiments. This calls into question common assumptions about human learners' representations and underlines the importance of controlling for low-level auditory confounds in the design of statistical learning experiments.

  • Visual Cognition Inspired Network for Few-Shot Fine-Grained Image Classification

    Few-shot fine-grained image classification (FSFGIC) is the task of developing computational models to recognize highly similar sub-categories from minimal training data. This task poses a core dilemma: over-emphasizing local details makes models vulnerable to noise and overfitting, while relying on global features often lacks the sensitivity required for fine-grained discrimination. Inspired by the human visual cognition strategy of dynamically balancing global and local information, we propose an Adaptive Global–Local Balance Network (AGLB-Net), a framework that computationally implements adaptive global– local integration for FSFGIC. AGLB-Net introduces two key modules: a Hierarchical Discriminative Feature Refinement (HDFR) module that progressively integrates representations from global semantics to fine-grained details, and an Adaptive Regional Re-Attention Module (ARRM) that automatically localizes and emphasizes discriminative regions without additional supervision. Extensive experiments on widely-used benchmarks demonstrate that AGLB-Net consistently achieves state-of-the-art performance across various few-shot settings, validating the effectiveness of human visual cognition-inspired adaptive global–local balancing in FSFGIC.

  • Improve Auditory Perturbation RSVP EEG Signal Decoding By Dual-view Backbone Based Dual-Stream Knowledge Distillation Network

    Rapid Serial Visual Presentation (RSVP) is a critical cognitive paradigm for investigating visual attention dynamics and developing high-speed Brain-Computer Interfaces (BCIs). However, real-world deployment faces challenges from environmental noise, particularly cross-modal auditory perturbations that degrade neural decoding. This work proposes the Time-Frequency Dual-View Network (TFDV-Net), which fuses time-domain and frequency-domain features via a global-local interaction backbone to capture weak ERP signals. Furthermore, a Dual-Stream Knowledge Distillation strategy is introduced to enhance robustness, enforcing the model to reconstruct task-invariant cognitive patterns from noise-corrupted signals by learning from a "clean-state" teacher. The proposed method achieves a state-of-the-art balanced accuracy of 83.02% on a dataset containing ecological noise. Notably, we find that semantic conversational noise induces a significantly larger performance drop compared to non-semantic urban noise. This result quantitatively validates the hypothesis that high-level semantic processing imposes a greater cognitive load on the visual attention system than low-level acoustic interference.

  • Finding Flaws: Using Small Group Discussions to Improve Understanding of Experimental Design

    Many students lack skills in comprehending summaries of scientific studies as well as thinking critically about experimental designs and evidence. This research explored whether group discussion activities may help support skills needed for understanding and evaluating research summaries, specifically when there is a lack of evidence for stated conclusions due to flaws in research designs. This study was conducted in undergraduate Research Methods in Psychology courses, using discussion activities where students were asked to read short reports of studies similar to popular media accounts of research findings. The study manipulated only discussion group size (twos, threes and fours) while holding all other aspects of instruction constant. Individual performance on tasks asking students to find flaws in new examples was better on average after working in groups of four compared to groups of two, and weaker students were more likely to benefit from interacting with peers after working in groups of four.

  • Learning Abstract Categories through Linguistic Feature Explanations

    Abstract categories often lack a single, bounded perceptual referent, making them harder to learn than highly concrete categories. Explanations that identify properties shared across category members may support category learning by highlighting features that generalise beyond individual examples. This study investigates the role of explanatory concepts in category learning using a label-free Concept Bottleneck Model framework. We generate human-interpretable feature concepts for both basic-level and super-ordinate categories, then learn a projection from visual representations into this concept space. A classifier then predicts category labels from these concept activation to evaluate whether explanatory features support category learning. We evaluate the approach on CIFAR-100, CUB-200, and ImageNet, organising labels into basic-level and super-ordinate categories. Across datasets, concept-based models achieve accuracy comparable to standard classifiers while providing interpretable concepts. Visualisations reveal distinct category clusters in concept space, and evaluations on held-out subordinate classes suggest that explanatory concepts transfer effectively to previously unseen subclasses.

  • AI counsellors exaggerate linguistic qualities of human counsellors, and human clients align more with AI counsellors

    Many people treat large language models (LLMs) like counsellors, seeking advice on social and emotional issues. This could improve quality of life for people who lack access to mental healthcare, but it also comes with serious risks. We recently got the opportunity to analyze written Chinese conversations between hundreds of human–human client–counsellor pairs and thousands of human–AI client–counsellor pairs. By analyzing linguistic and semantic features of those conversations, we provide evidence that human and AI counsellors share many qualities, such as writing predictable messages and expressing positive sentiments, and that human clients align with AI counsellors more than with human counsellors, including when AI behaviour diverges from human counsellor behaviour. This study contributes sorely needed empirical descriptions of real-world counselling sessions with chatbots, which can help researchers at the intersection of AI and psychology understand whether and how LLMs should play a role in mental healthcare.

  • Individual Differences in Irony Processing in Hindi-English Bilingual Reading

    Bilingual readers show substantial variability in how non-literal language is processed, yet irony is often examined in terms of uniform processing costs. The present study investigated how individual differences shape irony–literal reading behavior in Hindi-English bilinguals using eye-tracking. Fifty-four participants read short texts containing ironic criticism or literal counterparts in either Hindi (L1) or English (L2) and completed measures of cognitive ability, language experience, and meta-representational skills. Aggregate analyses revealed no reliable reading-time differences between ironic and literal sentences. Primary analyses focused on within-participant irony-literal difference scores. Correlational and Bayesian regression analyses showed that the magnitude and timing of irony-literal divergence were systematically related to working memory, language-use patterns, proficiency, and Theory of Mind, with predictors across languages of testing. These findings indicate that irony-related reading differences are not categorical but emerge from structured individual variability, highlighting the value of individual-difference approaches for studying pragmatic processing in bilingual reading.

  • Do Vision-Language Models Solve the Traveling Salesperson Problem Like Humans?

    Humans produce near-optimal (i.e., within ~10% of minimum-length) tours for small Traveling Salesperson Problem (TSP) instances of up to 120 points despite the problem's NP-hardness. The current study asks whether vision-language models (VLMs) employ human-like strategies when solving TSP instances. Using the stimuli and human data of Marupudi et al. (2022, Cognitive Science Conference), we found that Gemini 3 Pro, GPT-5.2, and Claude Opus 4.5 generally show the same performance as humans, including a linear increase in deviation from optimal tour length with increasing problem size. The VLMs also behave consistently with heuristics and strategies that have been documented for humans: convex hull, crossing avoidance, and divide-and-conquer via visual clustering. Finally, they show a linear increase in total tokens consumed with increasing problem size, paralleling the linear increase in human solution times with increasing problem size. Directions for future research include extending to larger problem instances and to adversarial examples.

  • Mental disorders emerge from the gut microbiome gradually transitioning from a symbiotic to a parasitic state: a 4E cognition approach to computational psychiatry

    Recent studies suggest that changes in the gut microbiome, as well as social isolation, are closely correlated with mental disorders. We argue that these correlations are intertwined, and the way these interlinkages lead to mental disorders could be productively explored using computational models of agent-environment interactions. To illustrate this approach, we present a reinforcement learning model of social isolation, and discuss how an extension of this model -- based on social stress-related changes to gut states -- could account for the role of the gut microbiome in the emergence of mental disorders.

  • Reconstructing Visual Events from Fragmented Scenes

    The human cognitive system resolves ambiguity when presented with spatially fragmented input. Sequential structures in visual narratives help understanding information-extraction across fragmented and undefined temporal order. In this study, 100 participants viewed and rated 54 comic strips (9 comic strips rearranged into 6 layout structures) based on their understandability, enjoyability, and navigation difficulties. Using mixed-methods, we observed non-intuitive layouts such as 'Blocked' and white-space overusing 'Separated' layouts have consistently shown higher response latency (p<0.001 and p=0.008), more viewing times (p<0.001), more regressive eye movements, and rerouting gaze as compared to standard 'Grid' formats: suggesting higher processing costs. However, the subjective ratings show another aspect of cognitive processing: factors such as 'enjoyability' shows less association with processing costs, suggesting a dissociation between cognitive efficiency and satisfaction. This further suggests that 'cognitive inefficiencies' are not mere disruptions but also a pathway to restructure compensatory strategies while reconstructing events from fragmented information.

  • Cognition Meets Affect: A CAPS-Inspired Multimodal Framework for Personality Prediction

    Personality computing has long struggled to bridge the gap between high-level psychological theories and low-level computational features. To address the lack of psychological grounding in personality computing, we propose a multimodal framework based on the Cognitive-Affective Personality System (CAPS). We decompose inputs into "Cognitive Units" (semantic construal) and "Affective Units" (emotional reactivity). Specifically, we introduce a parameter-free Matrix-Based Entropy mechanism to dynamically weight hidden states from Qwen2.5-3B, treating representational entropy as a measure of cognitive complexity. For audio, we combine linguistic content (Wav2Vec2) with prosodic features (OpenSMILE). Evaluated on a new self-reported Chinese dataset, our approach significantly outperforms unimodal baselines. Crucially, interpretability analysis reveals that the model's internal entropy profiles spontaneously recover latent psychological meta-traits (Plasticity and Stability), demonstrating that our architecture effectively bridges deep learning features with the intrinsic structure of human personality.

  • Short-Form Video Platforms and Attentional Control: Evidence for Domain- and Platform-Specific Associations

    Short-form video (SFV) platforms provide immersive, algorithmically curated content streams that place distinctive demands on attention and cognitive control. We investigated whether associations between SFV use and ADHD-related symptoms depend on platform-specific affordances and whether these associations are domain-specific across inattention and hyperactivity–impulsivity. In a cross-sectional survey of 193 emerging adults, we measured daily use of TikTok, Instagram Reels, YouTube Shorts, and Facebook Reels, alongside dimensional ADHD symptom-scores and problematic social media use. TikTok use was positively associated with total ADHD symptomatology and inattention, but not with hyperactivity–impulsivity. Both relationships were mediated by problematic social media use. No significant associations were observed for other platforms. These findings suggest attentional consequences of SFV engagement are platform-specific and may depend on interface design and algorithmic structure. The selective association with inattention challenges task-switching or arousal-based accounts of media effects, and instead highlights the potential role of immersive, reward-driven attentional capture.

  • Identifying the visual features of European Paleolithic cave art that are diagnostic of object category and time period

    Humans in the Upper Paleolithic period were among the earliest people to produce figurative art. Analyzing variation across works of art from this period could provide insight into the distinguishing features of various technocomplexes—cultures defined by different use of stone-working technologies. Here we analyzed 140 Paleolithic cave paintings of animals found in four cave sites in Monte Castillo, Spain. Using all available evidence, archaeologists determined what kind of animal was depicted in each painting (Horse, Bison, Deer, or Ibex) and which time period it was from (Gravettian, Solutrean, or Magdalenian). We used multiple vision algorithms (ResNet-50, ViT, SigLIP2) to investigate to what degree their visual properties alone were sufficient to distinguish depictions from different time periods, cave sites, and spatial contexts within each cave. Our findings highlight the promise of such tools for analyzing variation in cultural artifacts across the Paleolithic period.

  • No Reliable Evidence of Self-Reported Sentience in Small Large Language Models

    Whether language models are sentient cannot currently be determined by direct observation. We therefore ask a more tractable surrogate question: do models report themselves to be sentient, and are these reports truthful? We address this by first querying several open-weights models about their own subjective experience, and then verifying responses using classifiers trained on internal activations. We draw upon three model families (Qwen, Llama, GPT-OSS) ranging from 0.6 billion to 70 billion parameters, approximately 130 questions about subjective experience, and three classification methods from the interpretability literature. First, we find that models consistently deny being sentient: they attribute consciousness to humans but not to themselves. Second, classifiers trained to detect underlying beliefs - rather than mere outputs - provide no evidence that these denials are untruthful. Third, within the Qwen family, larger models deny sentience more confidently than smaller ones. These findings contrast with recent work suggesting that models harbour latent beliefs in their own consciousness.

  • Italian Personal Pronoun and Name Use in Parental Speech: Autistic Versus Typically Developing Children

    Children with autism spectrum disorder (ASD) often experience difficulty using and interpreting pronouns. While prior work examined referential language primarily in English (overt-subject), less is known about "pro-drop" languages where subject pronouns are optional. This study investigates personal pronoun and name use in Italian parental speech to autistic children compared with typically developing (TD) peers. Transcripts from 10-minute mother–child free-play were analyzed across 70 dyads. We found that mothers of autistic children used the first-person pronoun io (I) as well as the child's name, significantly more frequently than mothers of TD children. However, there was no consistent preference for using names versus the second-person pronoun tu (you) to address the child. Within ASD group, autism severity did not predict pronoun or name use. Increased usage of (optional) subject pronouns and names might reflect a compensatory strategy in parental speech to autistic children, to enhance clarity and avoid referential ambiguity.

  • Children's inquiry learning strategies may depend on their prior conceptual knowledge

    The present study explored how children's prior conceptual knowledge affects their inquiry learning strategies. 139 elementary school children performed a pretest on water displacement. Then, they evaluated whether seeing the outcome of four different experimental setups would help them learning something about water displacement. The experimental setups compared two spheres of different sizes and materials and were designed in such a way that children's preferences indicated either an effect-producing or a controlled testing strategy. We conducted probabilistic modeling of children's prior beliefs and formed six belief clusters based on a hierarchical clustering algorithm. While children tended to evaluate all experiments as helpful, the pattern of preference indicated an effect-producing strategy for children with a clear misconception and a controlled testing strategy for children with an uncertain or correct belief. This finding, although exploratory, suggests that prior knowledge may explain why children sometimes succeed and sometimes fail in conducting informative experiments.

  • Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies

    The rapid evolution of Large Language Models (LLMs) has led to the emergence of Multi-Agent Systems where collective cooperation is often threatened by the "Tragedy of the Commons." This study investigates the effectiveness of Anchoring Agents—pre-programmed altruistic entities—in fostering cooperation within a Public Goods Game (PGG). Using a full factorial design across three state-of-the-art LLMs, we analyzed both behavioral outcomes and internal reasoning chains. While Anchoring Agents successfully boosted local cooperation rates, cognitive decomposition and transfer tests revealed that this effect was driven by strategic compliance and cognitive offloading rather than genuine norm internalization. Notably, most agents reverted to self-interest in new environments, and advanced models like GPT-4.1 exhibited a "Chameleon Effect," masking strategic defection under public scrutiny. These findings highlight a critical gap between behavioral modification and authentic value alignment in artificial societies.

  • Inhibition is not emotion blind: Behavioural and neural evidence from an emotional stop signal task.

    Response inhibition being an essential executive function gets modulated by emotional context adhering to growing evidence. Understanding the reciprocal relationship of emotion and cognition has become crucial to human awareness. The current study examined the influence of emotional valence at behavioural and electrophysiological levels by using an emotional stop signal task. 14 healthy adult participants completed the ESST task wherein happy, angry, neutral facial expressions served as emotional stop signals while recording brain signals with electroencephalogram. Stop signal reaction time (SSRT) was indexed for behavioural inhibition and neural correlates were assessed using N2 and p300 event related potentials. Stop signal reaction time was significantly shorter for angry faces contrary to happy faces which elicited enhanced frontal N2 amplitudes, whereas larger P300 amplitudes were obtained for happy as compared to angry expressions. The present evidence highlights that response inhibition is not emotion blind rather is modulated in context of emotional valence.

  • Complexity Gates the Emergence of Visual Expertise

    The face inversion effect (FIE)—the disproportionate impairment in recognising inverted faces—has been central to debates between domain-specific and perceptual expertise accounts of visual processing. Although expertise-based explanations predict that inversion effects should emerge for non-face stimuli following learning, most supporting evidence derives from memory-based recognition tasks, leaving open the possibility that such effects reflect retrieval strategies rather than perceptual encoding. The present study tested this claim using a delayed matching-to-sample task with prototype-defined checkerboards, isolating perceptual encoding from long-term memory retrieval while manipulating stimulus complexity. Participants completed a categorisation training phase followed by a matching task involving familiar and novel stimuli presented upright or inverted. Results revealed a robust inversion effect for familiar stimuli only in the high-complexity condition, whereas no such effect emerged for simpler stimuli. These findings support the McLaren–Kaye–Mackintosh theory of perceptual learning by demonstrating that expertise modifies online perceptual encoding and that the engagement of holistic-like processing depends on stimulus complexity. We conclude that inversion effects reflect general expertise-based learning mechanisms operating under conditions of high informational load rather than face-specific perceptual architecture. Keywords: Perceptual expertise; Face inversion effect; Stimulus complexity; Perceptual learning; Visual categorisation; Holistic processing

  • Visual Search in Cluttered Environments using Active Viewpoint

    Visual clutter impairs visual search in 2D scenes; however, it remains unclear how factors that contribute to clutter unfold in scenes with a 3D structure similar to the real-world. We investigated visual search using a head-mounted display that provides stereoscopic depth and allows limited viewpoint adjustment, enabling participants to search for everyday objects in controlled virtual scenes with different degrees of visual clutter. Clutter was manipulated by varying the number of objects and the visual complexity of the background, while the presence of a search target was manipulated independently. Reaction time, accuracy, and eye movements were recorded. Results showed that participants responded faster and made fewer fixations when the number of objects was low and when a target was present, suggesting that while classic factors affecting 2D visual search extend to scenes with naturalistic 3D depth and occlusion, clutter effects depend on specific contributing factors.

  • Code-Switching and Multimodal Language Production: The Interaction of Code-Switching, Speech Disfluency, and Gestures

    Bilinguals produce code-switches (CS), speech disfluencies, and gestures during speech. While gestures are known to reduce cognitive load during disfluent speech, the interaction between CS and multimodal production remains understudied. This study investigated the relationship among CS, disfluency, and gesture use across L1-Turkish, L2-English, and Free-Language conditions using academic and daily topics. Results showed intrasentential CS occurred more frequently in Free-Language condition, whereas representational gestures were more prevalent in single-language conditions. Co-occurrences of representational gestures and speech disfluencies were significantly higher in academic topics, modulated by L2 proficiency. Furthermore, intrasentential CS co-occurred with representational gestures, both independently and simultaneously, with speech disfluency more in the Free-Language than L1-Turkish condition, again modulated by proficiency. In contrast, intersentential CS co-occurred with non-representational gestures in Free-Language condition. Results suggest that bilinguals use representational gestures to maintain a single language or manage complex intrasentential CS, while non-representational gestures serve as markers for intersentential CS.

  • Children's expectations of dominant and prestigious leaders

    Humans are enmeshed in many hierarchical relationships, like that between a parent and child. Leaders in dominance hierarchies attain influence through coercion and intimidation, while leaders in prestige hierarchies are granted influence based on competence and prosociality. Across two studies (N = 214), we examined whether children have different expectations about how dominant and prestigious leaders allocate resources. Children learned about two social groups, the dominant Glerks and prestigious Zonks. In Study 1 (N = 109), children ages 4–5 watched a leader and follower divide five apples in 5–0 or 3–2 splits, and were asked who received a given share. Overall, children expected similar allocations across leaders. In Study 2 (N = 105), children ages 6–8 predicted whether a dominant or prestigious leader would share an apple with a hungry follower or eat it. Children expected prestigious leaders to be more likely to give the apple (82%) than dominant leaders (48%).

  • Between Perception and Conceptualization: Graded Constraints on Spatial Representation

    Semantic attributes such as contact, inclusion, and support are often assumed to reflect universally available properties of spatial scenes. This study introduces a preliminary method for empirically evaluating whether such attributes are directly available in spatial scenes or shaped through conceptualization. Speakers of Moroccan Arabic and English rated the applicability of geometric, functional, and qualitative physical attributes for a set of simple spatial scenes. In this pilot test, we examine which attributes show cross-linguistic convergence, indicating conceptual properties of the scenes, and which show cross-linguistic (or individual) variation, suggesting other influences on their interpretation. Cross-linguistic correlations provide preliminary evidence of graded convergence, with attributes differing in their degrees of alignment across languages. These patterns suggest that universality in spatial semantics is graded rather than categorical. These findings illustrate the potential of this approach to provide an objective and fine-grained empirical assessment of how spatial attributes are accessed and interpreted.

  • Schematic Versus Realistic Face Processing in Autism Spectrum Disorder: Implications for Gaze Cueing

    This study investigated how the complexity of facial stimuli impacts attentional orienting to eye-gaze in adults with and without autistic traits. Using a gaze cueing task, participants responded to schematic line_drawn faces or realistic photographic faces providing nonpredictive central gaze cues to the potential location of a lateralized target. Reaction times were faster on congruent than incongruent trials for both schematic and realistic face cues at both 100ms and 300ms SOAs, with no overall differences in accuracy. Reaction times were faster for longer SOAs, indicating more efficient orientation with additional processing time. No main effect for face type or interaction with autistic traits was observed, suggesting comparable processing to gaze direction in both schematic and realistic faces, and across group type. The findings indicate that gaze orienting is similar in high-functioning individuals with autistic traits and not critically dependent on the perceptual richness of the face stimulus.

  • Central Versus Peripheral Face Processing in Autism Spectrum Disorder: Implications for Gaze Cueing

    Visual orienting in adults with and without autistic traits was explored using a spatial cueing paradigm where participants viewed schematic faces that either appeared centrally with a shift in gaze direction, or peripherally with gaze changes embedded in the laterally presented faces. Targets appeared nonpredictively to the left or right after variable stimulus onset asynchronies (SOAs). Centrally cued targets elicited faster reaction times than peripherally cued targets, and longer SOAs resulted in faster responses, indicating more efficient orienting when more processing time was available. A validity effect was observed overall, with congruent cues yielding faster responses than incongruent cues, with the effect being strongest for peripheral cues and at short_to_intermediate SOAs. Minimal differences were observed between autistic trait statuses, suggesting broadly comparable orienting performance. The findings indicated that gaze most effectively drove attention with peripheral cues at short SOAs, and (low-support, college-aged) autistic individuals may not have visual orienting impairments.

  • Planning as a Cognitive Cost that Discourages Goal Switching

    People exhibit a striking preference for maintaining their current goals. While this bias can be framed as irrational (e.g., the sunk cost fallacy), this persistence may be adaptive for conserving cognitive resources. We hypothesized that humans are sensitive to the cost of planning that comes with switching between goals. We conducted a behavioral experiment (n=60) in which participants pursued goal completion under two conditions: one in which they had to plan towards goals before pursuing them and another in which they were provided with a pre-specified optimal plan. We found that people persisted with their goals more when they had to create their own plans. We quantified the cognitive costs of planning with behavioral modeling and found that planning costs predict decreased goal switching. These results indicate that individuals are sensitive to the cost of planning and suggest that goal persistence may be an adaptive response to these cognitive costs.

  • Beyond the Laughable: How Appraisal Shapes Infant Laughter

    Laughter is one of the earliest social signals in human development, yet how cognitive processes from infants shape its acoustic and temporal form remains poorly understood. We propose the Dynamic Infant Appraisal Model, a computational framework grounded in appraisal theory that simulates the infant's continuous cognitive evaluation during naturalistic caregiver-infant interaction. Using longitudinal data from the SAYCam corpus, we investigate how appraisal variables relate to four laughter features: Reaction Time, Overlap Ratio, Intensity, and Pitch. Our results reveal a functional dissociation: temporal features, particularly Overlap Ratio, are shaped by Familiarity, Coping Potential, and Pleasantness, and follow an Inverted-U trajectory relative to schematic discrepancy. Acoustic Intensity correlates with all appraisal variables but shows only a monotonic pattern. These findings suggest that infant laughter is a graded, multidimensional signal in which temporal dynamics encode predictive mastery and acoustic form encodes the reward value of cognitive challenge.

  • CogSleep-Net: A Cognitive-Inspired Hierarchical Framework for Automatic Sleep Staging

    Sleep staging is a fundamental component of clinical sleep assessment and neurological screening. While recent deep learning methods have improved automatic sleep staging, most approaches treat it as a static sequence classification problem and overlook key cognitive mechanisms underlying expert scoring, including prior-guided perception, hierar-chical abstraction, and contextual memory integration. To address this gap, we propose CogSleep-Net, a cognitively inspired framework that models sleep staging as a dynamic reasoning process driven by coordinated perception, ab-straction, and memory. The model integrates prior-guided dual-domain perception to emphasize physiologically sali-ent patterns, hierarchical abstraction to disentangle micro-level waveforms from macro sleep dynamics, and a novel gated state-space memory mechanism that combines short-term evidence with long-term contextual information. Cog-Sleep-Net is evaluated on two public benchmarks (Sleep-EDF-78 and ISRUC-S1) and a self-collected high-density PSG dataset. It achieves 87.25% accuracy and 86.59% macro-F1 on Sleep-EDF-78, with consistent gains across da-tasets, demonstrating strong robustness and clinical applicability.

  • Confidence in Context: How Self-Beliefs Shape Discovery

    Learning in discovery environments requires learners to regulate hypothesis testing and evidence integration in the absence of external instruction. Changes in task-specific confidence are often interpreted as signals of learning progress, yet its relationship to performance varies across learners. In the present study, we examined whether the informativeness of local confidence change in a discovery task depends on learners' global self-efficacy. Participants completed a discovery task while providing trial-wise judgments of learning (JOLs). Baseline self-efficacy predicted initial, task-specific confidence but not objective performance during training or posttest. Instead, increases in JOLs over training reliably predicted posttest performance. Further analyses suggested that the strength of this relationship varied with learners' self-beliefs, such that confidence growth was more tightly coupled with performance for learners with lower self-efficacy. Together, these findings indicate that confidence change is not uniformly diagnostic of learning in discovery tasks and that global self-efficacy serves as an important boundary condition shaping when confidence change facilitates discovery.

  • Children Endorse Chatbots, but Prefer Humans, for Health-Related Information

    This study examines 135 7- to 14-year-old children's beliefs about AI chatbots, doctors, and non-expert adults. Children are asked both whether each of the informants can correctly answer different kinds of health questions and which of the informants they would prefer to ask their own health questions. Children's endorsements in the chatbot informant are as high as their endorsements for doctors, and their endorsements of statements by non-expert grownups decreases with age, However, their endorsements belie their preferences: children significantly prefer not only the expert human informant (the doctor), but interestingly also significantly prefer the non-expert grownup to the chatbot, despite giving this informant low statement endorsements. These results suggest that children's intuitions about AI chatbots are nuanced: they believe that these informants can acquire correct knowledge, but they are not necessarily eager to consult chatbots themselves.

  • Spatial Metaphor Priming in Humans and Language Models

    English prepositions are highly polysemous, occurring in both concrete spatial and conventionalized metaphorical uses (e.g., in love, on time). Abstract uses of prepositions, according to Conceptual Metaphor Theory, are grounded in spatial source domains. In this work, we test whether these conceptual links are active during online processing of in- and on-metaphors. Through semantic priming experiments, we investigate whether access to grounded spatial meaning during metaphor processing depends on visual input, or whether linguistic descriptions of spatial relations can also activate similar representations. Further, we probe if the priming experiments can be replicated in artificial language models to test whether distributional patterns in text data are sufficient to give rise to the types of priming effects observed in humans.

  • Dynamic reorganization of leader–follower coordination in expert Latin American dance

    Interpersonal coordination in joint action is often assumed to rely on stable, predefined leader–follower roles. However, in complex embodied interactions such as partner dance, coordination and role organization may be dynamically reconfigured depending on movement context. This study examined how physical connection, partner experience, and role-related asymmetry shape interpersonal coordination in expert Latin American dance. Eight expert dance pairs performed a standardized rumba sequence under connected and non-connected conditions with familiar and unfamiliar partners. Motion capture data were analyzed using C2RQA, an extended cross-recurrence quantification method that captures directional coupling between movement time series, applied to temporally segmented movement phases. Coordination strength, temporal structure, complexity, and role-related asymmetry were quantified. Results showed that coordination was sparse and intermittent—rather than continuously sustained—and varied systematically across movement phases. The effects of physical connection and partner experience were context-dependent, and leader–follower asymmetry shifted across phases rather than remaining fixed. These findings highlight dynamic, context-sensitive reorganization of coordination and roles in expert joint action, suggesting that skilled partnered movement relies on flexible perception-action coupling rather than fixed role execution.

  • How Far Does Vision Go? Interpretable Visual Statistics from Word-Linked Web Image Distributions Predict Multisensory Norms for Chinese Words

    Recent work reports that large language models, including those with visual training, fail to recover the sensorimotor structure of human concepts. We argue this reflects how visual information has been operationalized, as opaque embeddings tied to single training instances, rather than a limit of vision itself. We test this by representing 593 Chinese disyllabic words through interpretable low-level visual statistics, pixel-level and mid-level features aggregated over word-linked, CLIP-filtered web images, and evaluating their predictive power for human sensorimotor norms across six modalities beyond lexical controls. We find a graded profile of cross-modal decodability. Visual norms are predicted best, followed by gustatory, olfactory, tactile, and auditory, while interoception sits at the predicted lower bound. The graded pattern, with substantial gains for non-visual modalities, indicates that prior negative results reflect a choice of representation rather than a failure of low-level visual statistics per se.

  • Visuo-tactile stimulation is associated with faster self-recognition of ambiguous self-other morphs

    The enfacement illusion reflects changes in self-face representation following synchronous multisensory stimulation. This study examined whether such stimulation is associated with changes in self-other face discrimination at an ambiguous morph level (S50), focusing on both the rate of 'self' responses and reaction times across different visual field positions. Participants completed a self-recognition task before and after synchronous visuo-tactile stimulation, with faces presented centrally or in the left and right visual fields. We found that the reaction times for 'self' responses to S50 morphs were faster following stimulation, indicating a general facilitation in responding to ambiguous self-related stimuli. This effect was most pronounced for centrally presented faces, while lateralized differences were weak and not statistically significant. In contrast, the proportion of 'self' responses did not change across sessions. However, in the absence of a control condition, the observed reaction time facilitation cannot be uniquely attributed to enfacement-related processes and may reflect practice or repetition effects. These findings suggest that reaction times may be sensitive to multisensory stimulation, but clearer conclusions will require better-controlled designs.

  • Adaptive Patch Salience-Guided Differential Privacy for Brain MRI

    Human visual perception does not treat all regions of an image as equally important or sensitive; instead, attention and cognitive priority depend on semantic relevance, task demands, and domain-specific knowledge such as clinical significance. Most existing differentially private image obfuscation methods inject uniform noise across regions, implicitly assuming that perceptual and privacy sensitivities are homogeneous. We propose a cognitively inspired, saliency-guided differential privacy framework for brain MRI images that models heterogeneous perceptual and semantic importance. High-level visual representations are extracted using a Vision Transformer (ViT)-based pretrained Masked Autoencoder (MAE) to approximate human semantic sensitivity. Region saliency scores, computed via semantic similarity, may reflect cognitive and clinically relevant patterns. Privacy budgets are adaptively allocated, applying stronger noise to salient regions while preserving structural details elsewhere. Experiments on brain MRI datasets demonstrate improved privacy-utility trade-offs, effectively protecting sensitive areas without degrading overall perceptual quality.

  • Beyond Accurate AI: Counter-Bias Advice Calibrates Human Judgment

    With the performance improvements and widespread social adoption of artificial intelligence (AI), including large language models, AI-assisted decision-making is rapidly expanding. On the other hand, as human decision-making is increasingly delegated to AI, concerns about degrading the quality of autonomous human decision-making when AI is unavailable has become an urgent social issue. This study focused on the insight from the wisdom of crowds that diverse opinions improve decision-making accuracy. We call advice that judges based on bias that counterbalances the human bias, "Adaptive Counter-Bias Advice." Through theoretical analysis and behavioral experiments, we confirmed that people who repeatedly made judgments while referring to Adaptive Counter-Bias Advice showed reduced bias and improved accuracy in autonomous decision-making in situations where advice was no longer available compared to those who used accurate advice. These findings offer new insights into using and selecting AI systems to enhance autonomous human decision-making performance.

  • Computational phenomenology of self and time in borderline and narcissistic personality disorders

    Measuring how groups differ in construing concepts from natural language is central to computational social science and clinical NLP, but existing methods often trade statistical validity for interpretability. We extend Supervised Semantic Differential (SSD), which builds participant-level personal concept vectors from pretrained embeddings around lexically anchored concept references, to cross-group comparison. Our extension replaces regression-derived gradients with interpretable centroid-contrast vectors and adds permutation-based inference over group-centroid cosine distances, enabling omnibus and pairwise tests without parametric assumptions. We demonstrate the method on Polish life-story clinical interviews from patients with Borderline Personality Disorder (BPD), Narcissistic Personality Disorder (NPD), and controls (N=62), targeting concepts of self and time. Self-concept representations show robust omnibus and pairwise group separations; time-concept differences are driven mainly by BPD contrasts. Cluster interpretations align with phenomenological accounts, linking BPD to relational affective episodes and crisis-linked temporality, and NPD to comparatively decontextualized self-construal.

  • Stable Representations, Shifting Constraints: Reweighting Frequency and Concreteness in Lexical Access with Age

    Across cognitive domains, a central question concerns whether behavioural change reflects reorganization of representational structure or altered dynamics controlling access to a stable system. We address this in highly proficient Basque–Spanish bilinguals—young and older adults—using a naming-by-definition task probing two lexical constraints: Frequency and Concreteness. Across ages and languages, both dimensions constrained retrieval, supporting a shared architecture across languages. Aging produced a global decline in accuracy without amplifying frequency effects, constraining accounts predicting selective vulnerability of weak representations. However, the concreteness advantage was larger in older adults, suggesting increased reliance on intrinsic representational support. Although age-related costs appeared for low-frequency abstract words, this reflected additive effects rather than higher-order interaction. Critically, the absence of a three-way interaction does not support structural reorganization and is more consistent with changes in access dynamics. Accordingly, within this population and task, aging is associated with a reweighting of constraints on access.

  • Abstraction Facilitates Coordination in Repeated Social Interactions

    Every day social life requires repeatedly solving coordination problems: aligning actions to cooperate with others, often with incomplete communication. In contrast to rational accounts of coordination that predict its outcome based on payoffs, we interrogate the cognitive process that allows people to progressively generalize new coordination points based on past experience. We introduce the Progressive Coordination Paradigm (PCP) as a framework for studying the intersection of abstraction, generalization, and social coordination. Participants are asked to coordinate on sequences of patterned clicks on a grid. By progressively varying the layout of the grids, we show that people form shared abstractions after only a few interactions, learning from early coordination successes and flexibly generalizing core solution features to novel contexts. We show that the resultant abstractions are predictably path-dependent. These findings illustrate how human ad hoc conventions depend on mutually-constructed abstractions.

  • NaWID: A Stimuli Dataset for Investigating Spatial Prepositions of Containment and Support

    The Na/W Image Dataset (NaWID) comprises 96 visual stimuli designed to investigate the processing of spatial topological relations, specifically containment (w, "in") and support (na, "on") in Polish. NaWID has a systematic design, consisting of ambiguous spatial layouts alongside unambiguous exemplars, thus presenting a complementary perspective to existing datasets. We conducted preliminary studies and two experiments with native speakers to evaluate the dataset's adequacy. Experiment 1 (N = 30) demonstrated that ambiguous stimuli exhibit significantly longer reaction times during behavioral verification. Experiment 2 (N = 28), using eye-tracking and fNIRS, showed that while pupillary dilation is sensitive to logical congruency, it does not reflect visual ambiguity. However, gaze measures revealed more extensive scanning for ambiguous layouts, indicating sustained cognitive processing effort. These findings establish NaWID as an effective tool for exploring the spectral nature of spatial ambiguity and the oculomotor strategies involved in processing spatial layouts.

  • From Availability Heuristics to Causal Reasoning: A Neuro-Symbolic Cognitive Prosthetic for Long-Tail Clinical Decision Making

    In the long tail of precision medicine, rare case exemplars are scarce, forcing clinicians toward slow, resource-intensive "System 2" reasoning. Large Language Models may help, but they often suffer from a fluency-factuality trade-off, producing plausible yet hallucinated clinical guidance without logical grounding. To address these dual inefficiencies, we propose the NeuroCR, a neuro-symbolic cognitive prosthetic modeled after Dual Process Theory. Its architecture synergizes an implicit associative memory (System 1) with an explicit causal reasoning module (System 2). Crucially, we introduce a "Semantic Schema Integration" algorithm to solve the Symbol Grounding Problem for heterogeneous clinical data, and a "Metacognitive Gating" mechanism to dynamically arbitrate between retrieval and reasoning. Evaluations on rare somatic mutations demonstrate that NeuroCR achieves a 140-fold speedup while preventing hallucinations through topological constraints. This work suggests that AI should not merely automate decisions but serve as a transparent scaffold for high-stakes human reasoning.

  • How Expertise and Prior Expectations Impact Interpretations of Generalizations

    Past research has identified an asymmetry between how experts and novices interpret generalizations: novices tend to interpret generalizations more broadly. It remains unclear how or when differences in prior knowledge would create such asymmetries. We first investigate whether differences in experts' and novices' domain-level prior knowledge can explain the asymmetry. We find no consistent difference in their domain-level prior knowledge, but we do find a previously unidentified asymmetry in how low- and high-level experts interpret generalizations. We then investigate whether differences in the specificity of experts' and novices' prior knowledge can explain the asymmetry. We find that only listeners with specific prior knowledge vary their interpretations based on context, but differences in specificity do not create an asymmetry. Finally, we discuss the progression from novice to expert and consider remaining explanations for why experts and novices would interpret generalizations differently.

  • Neural Tracking of Linguistic Predictors in Spontaneous Conversational Speech

    This study investigates whether neural tracking of linguistic information extends from read speech to spontaneous conversation. Using the temporal response function (TRF) framework, we validate our approach on a read-speech EEG dataset and then apply it to EEG recordings from natural conversations. We observe reliable neural tracking of key linguistic predictors, including word onset, part-of-speech surprisal, and lexical surprisal, in spontaneous speech, with effects around 200, 400, and 600 ms. These results provide new evidence that linguistic neural tracking operates in natural conversational settings and confirm the feasibility of EEG studies in ecologically valid contexts.

  • Inside Out: Does Interoceptive Accuracy Influences the Attentional Prioritization of Enfaced Faces?

    Ownership is often implicitly measured using indices such as proprioceptive drift. However, the measure is noisy and does not coincide with explicit measurement. We look at attentional prioritization for induced ownership as another possible implicit measure. Across two experiments, we first examine the self-prioritization effect (SPE) for enfaced faces in a Temporal Order Judgment Task and its correlation with interoceptive accuracy (IAcc). In the second experiment, we tested whether the relationship between SPE and IAcc is mediated by the degree of enfacement. While the first experiment suggested a moderate correlation between SPE and IAcc, the second experiment found no correlation between IAcc and SPE, nor between SPE and the degree of enfacement.–öInterestingly, as with–öproprioceptive drift, the results–öshowed SPE even in a non-enfaced visual exposure condition without tactile feedback.–öOverall,–öthe pattern of results supports the distinction between implicit and explicit measures of ownership.

  • Can LLMs read between the lines? Exploring how they compare to human coders in categorising short pieces of ambiguous text

    While the democratisation of LLMs has proven fruitful in the field of Experimental Psychology, their adoption in some use cases has been slower. Here we explore how LLMs perform on a content analysis task following a predefined coding scheme. Using a qualitative dataset (346 responses, 6 labels with binary options) with an established inter-rater reliability over 95% as our benchmark, we compare models' outputs against human coders. The text responses are often ambiguous and require a certain level of inferential reasoning and subjective interpretation which LLMs still struggle with. We gave GPT 4-o, Qwen 2.5 and Mistral 7B the dataset and label definitions. GPT 4-o matched 59% of human labelling across responses, Mistral 7B matched 46% and Qwen 2.5 matched 42%. We discuss the 'reasoning' element of LLMs and potential ways forward.

  • Mutual Observability Determines When Precedent is Moralized in Multi-Agent Coordination Games

    Moral cognition evolved, in part, to facilitate cooperation. Since coordination problems constitute a large subclass of cooperation problems, we should expect coordination behavior to be subject to moralization. But which of several coordination strategies is preferable depends on their respective probability of success. This, in turn, often depends on whether other people's actions are mutually observable. Past research has focused on the moralization of coordination behavior in dyadic contexts. In two preregistered experiments (n = 472) based on incentivized real-time coordination games, we extend investigations of the moralization of coordination behavior to small groups, and study how it depends on the mutual observability of others' actions. In Experiment 1, we show that moving away from the current coordination solution (thus inducing miscoordination) is viewed as morally worse than sticking to it. In Experiment 2, we show that this is not the case when actions are mutually observable, making switching to a more mutually beneficial alternative achievable. In such contexts, abandoning precedent to coordinate on a more efficient equilibrium is viewed as morally superior than sticking to it. Crucially, the same coordination rule is viewed as morally superior in one informational context, and morally inferior in the other. Understanding the cognitive processes involved in successful coordination can shed light on which conventions, norms, and rules are morally appropriate and which aren't, when this is so, and why.

  • REM Sleep Selectively Optimizes Negative Emotional Face Processing

    Previous studies highlight sleep's role in emotional processing, with rapid eye movement (REM) sleep showing a negative bias, though multi-level neural mechanisms require further exploration. This EEG study examined overnight sleep's modulation of emotional face recognition, integrating event-related potentials (ERPs), oscillatory activity, and connectivity measures. Thirty-six participants performed an emotional face classification task before and after polysomnographic monitoring. Results showed sleep-dependent enhancement of the N170 component specific to negative faces and dynamic modulation of the late positive potential (LPP). Higher REM proportion and stronger prefrontal theta-gamma phase-amplitude coupling correlated with greater behavioral efficiency gains. REM-specific directed theta connectivity from central to prefrontal regions selectively predicted improvements in negative face processing. These findings indicate that REM sleep adaptively optimizes negative emotional processing via hierarchical mechanisms from early structural encoding to late motivational evaluation, providing new insights into sleep's therapeutic potential for emotional regulation.

  • Scalar Inferences, Fast and Slow

    Recent work on "scalar diversity" has shown that scalar inferences (SIs) do not form a unified class. For example, while and both form scales, hearers overwhelmingly interpret some as meaning not all, but rarely interpret pretty as meaning not beautiful. We show that SIs are correspondingly heterogeneous with respect to processing times and their sensitivity to world knowledge. In particular, hearers process SIs associated with closed class scales faster than open-class SIs or ad-hoc SIs; moreover, open-class SIs exhibit greater sensitivity to world-knowledge than their closed-class counterparts. These results, we argue, suggest the existence of at least two distinct paths for generating scalar inferences--one path is fast, automatic, and insensitive to the conversational context, while the other path is slow, deliberate, and sensitive to conversational context.

  • The Cognitive and Development Alignment of Computer Vision Models on the Raven Test of Fluid Intelligence

    Prior studies have evaluated the cognitive alignment of computer vision models and humans on fluid reasoning: the ability to solve novel, abstract problems independent of prior knowledge. This is commonly measured using the Raven Advanced Progressive Matrices (RAPM) test. Prior studies investigating how well Convolutional Neural Networks (CNNs) solve Ravens problems have found partial cognitive alignment with the performance of adults. We replicate this work using the ResNet-18 model trained on RAPM and, critically, extend for the first time to the question of developmental alignment: Does the improvement of ResNet-18's improvement across training track the developmental progression observed in children? We find evidence of partial developmental alignment: ResNet-18's performance over training follows a power function similar to the one that characterizes children's performance over development. A notable discrepancy is that the model performs best on the hardest problems, suggesting that it exploits shortcuts rather than engaging in genuine fluid reasoning.

  • When Motion Wins: Hierarchies of Cognitive Efficiency in Perceptual Grouping

    Decision-making benefits from efficient perceptual organization of cues minimize processing costs. This motivated the present examination of cognitive efficiency of competing Gestalt principles in digital natives by measuring both choice preferences and reaction times. 46 participants (25 males, Mean age-22.76 years) completed a web-based experiment consisting of five binary classification tasks, where 'Common Fate' competed against 'Similarity', 'Proximity', 'Closure', 'Good Continuity', & 'Pr–âgnanz'. Binomial tests revealed that Common Fate was selected above chance overall, but its dominance was hierarchical, strongest against 'Pr–âgnanz' (91.3%) and weakest against 'Similarity' (65.2%). Our findings suggest that motion-based grouping appears computationally efficient for digital natives, but is graded and shaped by cue conflict. We propose an efficiency-hierarchy model where perceptual organization functions as a cost-sensitive competitive system based on selection frequency and resolution cost, and conjecture that increased selection probability for 'Common Fate' is linked to recalibration of perceptual priors in digital natives.

  • Modeling Parkinsonian Freezing and Deep Brain Stimulation Effects in a Basal Ganglia Network

    The precise mechanism through which Deep Brain Stimulation (DBS) mitigates freezing episodes in Parkinson's disease remains unknown. We modelled Parkinsonian freezing using a model of the basal ganglia that adopts a Dynamic Neural Field network for simultaneous action selection (selecting which action is executed) and specification (resolving the continuous parameters governing how that action is performed). Action selection success was based on whether activation surpassed a threshold. The model robustly differentiated healthy and dopamine-depleted (Parkinsonian) conditions, producing stable action selection under healthy dopamine and impaired, freeze-prone dynamics following depletion. We then modelled DBS as a scaled reduction in afferent drive to the subthalamic nucleus (STN) and globus pallidus internus (GPi). While this formulation has succeeded in previous computational work, it did not yield consistent restoration of action selection within our framework, motivating future investigation of alternative DBS formulations.

  • Investigating Exploration and Belief Change in a Social Media Environment: Toward LLM-Based Cognitive Modeling

    This study examines how social media exploration shapes individual belief change and how psychological traits condition this process, with the aim of informing large language model (LLM)-based modeling. Participants explored a simulated social media environment by selecting content options and reading posts about a controversial topic. We focused on confirmation-biased exploration and opinion change. Results showed that the balance between confirmatory and disconfirmatory exploration predicted the magnitude and direction of opinion change. Trait effects were modest as main effects but emerged more clearly in interaction with initial opinion strength, consistent with prior theory. We also conducted an exploratory LLM-based analysis of behavioral prediction. Overall, the findings suggest that belief updating in social-media-like settings depends on interactions among exploration behavior, initial belief states, and individual differences.

  • Manipulating Interest to Study Attentional Control in Web-Based Information Seeking

    Web-based information seeking requires learners to regulate attention across heterogeneous sources. While interest is known to influence engagement, it is often treated as an uncontrolled background factor. This study examines how experimentally induced differences in interest relate to observable search behavior in a web-based report-writing task. Twelve undergraduate students completed two search tasks designed to elicit relatively high or low interest by manipulating topic and difficulty based on optimal stimulation principles. Interest was assessed using self-report measures before and after each task. Behavioral logs, including search frequency, dwell time, and visited domains, were analyzed alongside exploratory semantic similarity measures derived from page content embeddings. Results suggest that task characteristics are associated with initial interest levels, and that self-reported interest may change over the course of task engagement. Differences in behavioral patterns were observed across interest profiles, indicating that motivational state is related to variation in attention allocation during information seeking.

  • Inner Speech Supports the Processing of Abstract and Vague Concepts During Social Interactions

    This preregistered study examines the use of mental representations, with a particular focus on inner speech, during social interactions involving concepts that vary in abstractness and vagueness (semantic precision), using a hybrid methodology combining an interactive task, experience sampling, and survey measures. Results showed that abstract concepts relied more on verbal internal representations, whereas concrete concepts engaged more visual ones. Both abstractness and semantic precision affected participants' uncertainty in their evaluation of the ongoing inner experience. Together, these findings demonstrate that inner speech plays a functional role in negotiating meaning during social interaction, particularly when concepts are abstract or semantically vague.

  • A Drift Diffusion Model of Trust in Human-Robot Interaction

    Trust in human–robot interaction (HRI) is widely studied, yet the cognitive mechanisms underlying it are unclear. We propose the Expectation-Trust Model (ETM), a process-level account of trust grounded in expectations and formalized within a drift diffusion modeling framework. In ETM, trust is represented as an evidence accumulation process in which existing knowledge about robots determines the starting point, while deviations from expectations influence the drift rate. Positive drift reflects behaviors that exceed expectations, such as excellent performance, whereas negative drift reflects expectation violations, such as robot errors. We apply ETM to simulate trust dynamics from two published datasets, capturing empirical patterns including trust formation, violation, and repair. Differences in trust were explained by variation in starting point and drift rate corresponding to robot design, performance, and trust repair strategies. ETM provides a parsimonious mechanistic explanation for trust in HRI and generates precise, testable predictions about how expectations shape trust.

  • Cultural evolution of power dynamics between workers and managers in a collective foraging task

    We investigate how role-based power dynamics emerge between workers and managers in a collective foraging task. We recruited 223 participants, who repeatedly occupied either a forager role, or a manager role that coordinates the foragers' efforts. Foragers pay a tax to the manager, which instantiate a social dilemma: although higher taxes incentivize information acquisition to improve collective efficiency, self-interested players should prefer to push the tax toward its own immediate advantage. We operationalized this tension by eliciting participants' preferred tax rate. Contrary to an incentive-based account, managers' information investment was not associated with tax income and foragers' harvesting effort was not influenced by tax rate. Moreover, foragers consistently proposed about 28% tax while managers propose around 53%, even though both roles could have set it near 0 or 100%, respectively. Overall, the dynamics suggest convergence toward tax rates that reduce inequality between roles rather than toward self-interested behavior.

  • Explainability Shapes How People Respond to Errors in AI and Human Decision Support and Buffers Future Trust

    The current study examines the effects of advice explainability on advice judgment and post-error response when advice is provided by a goal-specific AI system versus a human expert in high-stakes decision-making contexts. Across four scenarios, participants were presented with advice that was either accompanied by a justification or not, and the advice was subsequently revealed to be incorrect. Results showed that prior to the error revelation, higher explainability did not affect perceived advice correctness, perceived advice understandability, or willingness to follow the advice, nor did it eliminate the preference for human advisors over AI systems. Instead, explainability most clearly shaped post-error responses. Providing justifications reduced responsibility attributed to the advisor, increased perceived error understandability, and lowered unwillingness to follow future advice. This pattern suggests that justifications may invite excuse generation and external attributions, making failures appear more understandable, reducing blame toward the advisor, and thereby buffering future trust.

  • ACT-R Modeling of Anxiety-Induced Distortion of Time Perception in Turn-Across-Oncoming-Traffic Decisions

    Anxiety plays an important role in predicting and avoiding future dangers, and adaptation under anxiety is assumed to be mediated by its influence on cognitive functions such as time perception. In this study, we constructed an ACT-R–based model that incorporates the amplification of anxiety over time. Using simulations of a turn-across oncoming traffic decision task in a driving context, anxiety was formalized not as a change in decision criteria but as a distortion of internal time progression. As a result, subjective time-to-collision (TTC) showed a logarithmic nonlinear relationship with objective TTC, and the between-condition differences corresponded to the difference structure of subjective time judgments observed in human experiments. Furthermore, examination of turn timing reproduced the anxiety-specific freezing process, in which the probability of performing the action decreased over time if it was not executed early.

  • The Brain Rot Effect: The case for adaptive attentional bandwidth

    Short-form "brain rot" content on platforms like TikTok is widely viewed as harmful for attention and learning. We find that brain rot can benefit learners with higher stimulation needs. We eye-tracked young adults (N=75) while they watched physics videos presented alone or alongside a task-irrelevant "brain rot-style" animated video. Individuals who reported preferring brain rot learned better from the brain rot condition than from the blander control condition. Furthermore, eye blinks and pupil dilation suggest that those who prefer brain rot were less overstimulated by brain rot content and required more effortful executive control to sustain attention to control content. Finally, preference for brain rot in our task is associated with higher media consumption and tolerance for high stimulation. These findings may point to individual variation in optimal cognitive load and provide preliminary support for an adaptive account of attentional bandwidth.

  • The Narrative Niche: Communicative Need in Narrative Drives Lexical Richness within Efficient Communication

    Efficient communication accounts of cross-linguistic variation propose that languages trade-off informativeness and complexity based on communicative need, often measured by usage frequency and operationalized as a free parameter. We propose that this need is not uniform across discursive genres and that the narrative genre applies a stronger pressure for lexical informativeness than non-narrative genres. Using spoken-language data from typologically diverse languages, we analyze the informativeness of lexical fields in narrative and other genres, as operationalized by their lexical richness. We show that frequency in narrative genres predicts lexical richness more strongly than in other genres. This supports the theory that narrative is a privileged communicative context that drives the informativity displayed by different parts of the lexicon.

  • What shapes learner language? Exploring the roles of first language influence and developmental features with a machine learning approach

    Prior studies of Second Language Acquisition have identified two forces that shape L2 production: cross-linguistic influence and L2 developmental universals. While past studies have focused on the acquisition of target structures, the distribution of structures in learner language production may still exhibit influence from these two forces. To disentangle the two kinds of impact on the distribution of linguistic structures in natural L2 production, we took a machine learning approach, leveraging over twenty typological and syntactic complexity features from five languages (English, Spanish, Portuguese, Chinese, and Korean). By comparing linguistically informed features that characterize 1) each language based on native data, 2) each language based on L2 production from multiple L1 background, and 3) speakers of different L1 background writing in the same L2, we identified the structural features that capture the developmental universal impact and those that reflect L1 influence. Our findings reveal the complex, bidirectional, and asymmetrical interplay of language typology, cross-linguistic influence, and language development.

  • The Crossroads of AI Social Decision: A Two-Dimensional Analysis of Decision Maturity and Cultural Orientation in LLMs

    As Large Language Models (LLMs) transition from passive text generators to autonomous social agents, decoding their underlying "cognitive DNA" is essential for ensuring alignment with diverse human values. This study investigates the intersection of socio-emotional maturity and cultural value orientations in frontier LLMs through the lens of cognitive science and cross-cultural psychology. We introduce the Dual-Axis Social Intelligence Scale (DASIS), a novel evaluative framework designed to decouple general social sophistication from implicit cultural biases—specifically along the Individualism-Collectivism (I-C) continuum. Across three interrelated experimental paradigms, we map the latent social architecture of nine state-of-the-art models. Study 1 establishes a behavioral baseline in Chinese, revealing that while models exhibit near-human social maturity, they suffer from a profound "Trans-Linguistic Value Locking" toward individualistic schemas, regardless of their developmental origin. Study 2 interrogates the "Whorfian Effect" (Linguistic Relativity) in synthetic cognition; findings indicate a Directional Asymmetry, where English context amplifies individualism while Chinese fails to exert a reciprocal collectivist pull, suggesting an "Axiomatic Individualism" ingrained during the alignment process. Study 3 explores "Cultural Perspective Taking," testing the models' Synthetic Theory of Mind (SToM) capabilities to voluntarily override default biases. Our results demonstrate that while default "personalities" are heavily skewed, models possess a latent Cognitive Decoupling capacity that allows for high-fidelity cultural simulation under explicit framing. This work provides a critical audit of the "WEIRD" (Western, Educated, Industrialized, Rich, and Democratic) bias in modern AI. We conclude that current alignment methodologies prioritize value consistency over cultural pluralism, potentially facilitating "Cognitive Imperialism" in global AI deployment. We propose DASIS as a standard for auditing the cross-cultural adaptability of autonomous agents.

  • France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions

    Sentences like "She will go to France or Spain, or perhaps to Germany or France.'' appear formally redundant, yet become acceptable in contexts such as "Mary will go to a philosophy program in France or Spain, or a mathematics program in Germany or France.'' While this phenomenon has typically been analyzed using symbolic formal representations, we aim to provide an account grounded in artificial neural mechanisms. We first present new behavioral evidence from humans and large language models demonstrating the robustness of this apparent non-redundancy across contexts. We then show that, in language models, redundancy avoidance arises from two interacting mechanisms: models learn to bind contextually relevant information to repeated lexical items, and Transformer induction heads selectively attend to these context-licensed representations. We argue that this neural explanation sheds light on the mechanisms underlying context-sensitive semantic interpretation, and that it complements existing symbolic analyses.

  • Revisiting Real-Time Digging-In Effects: No Evidence from NP/Z Garden-Paths

    Digging-in effects, where disambiguation difficulty increases with longer ambiguous regions, have been cited as evidence for self-organized sentence processing, in which structural commitments strengthen over time. In contrast, surprisal theory predicts no such effect unless lengthening genuinely shifts statistical expectations, and neural language models appear to show the opposite pattern. Whether digging-in is a robust real-time phenomenon in human sentence processing—or an artifact of wrap-up processes or methodological confounds—remains unclear. We report two experiments on English NP/Z garden-path sentences using Maze and self-paced reading, comparing human behavior with predictions from an ensemble of large language models. We find no evidence for real-time digging-in effects. Critically, items with sentence-final versus nonfinal disambiguation show qualitatively different patterns: positive digging-in trends appear only sentence-finally, where wrap-up effects confound interpretation. Nonfinal items—the cleaner test of real-time processing—show reverse trends consistent with neural model predictions.

  • Moralization by Analogy: A Novel Perspective on Moral Change

    Moral feelings and beliefs are powerful drivers of behavior. While analogies are commonly invoked to induce moral sentiment, analogical reasoning has not been empirically tested as a mechanism for moralization. Across two studies (N=292, N=234), we investigated whether analogical reasoning facilitated moralization of neutral target actions by transferring moral significance from moralized source actions. Participants judged neutral target actions paired with either analogous immoral source actions or unrelated immoral actions (control). Target actions were judged as more immoral in the analogy condition compared to the control, but only when the source and target were considered comparable (i.e., mappable relations), and the source was judged as highly immoral. These factors interacted synergistically, consistent with analogical transfer: moral significance transferred to targets when relational mappings supported the inference. However, poorly-perceived analogies backfired, producing less moralization than controls. These findings provide novel evidence for analogical reasoning as a mechanism for moralization.

  • Backward Digit Span Benchmarks Working Memory in LLMs

    The maintenance and manipulation of information in working memory (WM) is fundamental to human and artificial intelligence. Although human WM is famously capacity-limited, Large Language Models (LLMs) preserve inputs in the network's context window, profoundly reducing capacity limits that depend on maintenance alone. However, work in cognitive science suggests WM limits can also arise from representational interference during manipulation, raising the possibility of strong limits even in LLMs with perfect maintenance. Here, we evaluate 15 frontier LLMs and find models perform near-perfectly on forward digit span (recalling sequences in order) but collapse on backward digit span (recalling sequences in reverse). Moreover, backward span performance is significantly correlated with two measures of fluid intelligence (Raven's Progressive Matrices and ARC-AGI-1). These findings suggest backward digit span provides a plausible benchmark for the "working" component of working memory in LLMs and point to potential shared principles of information processing across systems.

  • Similarity and generalization as discounted integration over higher-order paths in semantic networks

    Similarity and generalization are often assumed to reflect distance in an underlying representational space. In semantic networks, however, distance can be defined by both shortest paths or by multiple indirect pathways. Previous research has shown that non-shortest, higher-order paths influence similarity judgements. Here, we extend these findings by re-analyzing a publicly available dataset to compare shortest-path models with models integrating over discounted higher-order paths. The latter better accounted for similarity judgements; error analyses indicate that this advantage arose from integrating multiple indirect paths rather than relying on shortest connections. These random-walk models implement the same core computation as the successor representation (SR), which has been proposed to support human generalization. Consistently, an SR-based model outperformed alternatives in accounting for inductive generalization judgements from the same dataset. These findings suggest that similarity and generalization are both shaped by higher-order connectivity between concepts.

  • Psychological Heterogeneity in User Preferences of Real-Time AI Mediation As Cognitive Scaffolds: A Latent Class Analysis

    User studies of AI-mediated communication typically report average effects, while interview-based work remains dominated by thematic narrative or Likert clustering, leaving the structured heterogeneity expressed in post-task interviews unquantified. We applied latent class analysis as a methodological bridge that converts qualitative interview codes into discrete, domain-specific user profiles, enabling design reasoning about for whom, on which dimension, and under what conditions real-time AI scaffolds help or burden. As a validating scenario, 29 non-native English speakers interacted with XPLAIN, a Wizard-of-Oz proactive scaffold in Zoom's sidebar, to bridge gaps in linguistic and cultural knowledge during a collaborative task. Across eight thematic domains, two-class solutions were consistently best-fitting (BIC); class membership was largely independent across domains, indicating multi-dimensional rather than global user types (e.g., longer English immersion was associated with more direct, interlocutor-focused repair strategies). Quantifying such heterogeneity, we argue, is critical for user-adaptive, timing-sensitive AI-mediation design.

  • Adults equate stillness with learning

    Learning is difficult to measure because it is not directly observable. Even in formal educational settings it's hard to tell whether a child might be thinking through a difficult problem or daydreaming. The same is especially true in informal settings like while watching TV at home—it's difficult to know whether a child is meaningfully engaged with what they are watching or whether they are overstimulated and unable to disengage. Caregivers and educators regularly rely on behavioral cues to infer kids' learning. Physical engagement, or sitting still, is often treated as a reliable signal of learning. However, stillness during screen-based media is not synonymous with learning in young children and may in fact indicate overstimulation and disrupted learning (Shepherd & Kidd, 2024). We asked 200 adults whether they thought the same child was more attentive and learning more from a video when sitting still compared to when making small fidgets, like twiddling their thumbs or tapping their fingers. Our results demonstrate that adults strongly believe that stillness indicates greater attention and learning in children. This has important implications for equity of opportunities in formal education settings, where an adult observer acts as a gatekeeper to learning opportunities for children that they may withhold from children who they misperceive to be learning less effectively because they are fidgeting.

  • Does Episodic Memory Help Close the Lexical Frequency Gap in Sensitivity to Syntactic Contrasts? A Test Using Retrieval-Augmented Language Models

    Grammatical knowledge and how it is empirically tested are typically considered robust to the frequency of the lexical items used in the expressions. Complementary Learning Systems theory proposes that hippocampal episodic memory, which enables rapid encoding and retrieval of specific experiences, allows learners to leverage those experiences when processing rare patterns. We test the hypothesis that robustness to lexical frequency can arise via such an episodic memory mechanism by evaluating whether retrieval-augmented language models (specifically, k-nearest-neighbor language models that augment parametric neural networks with explicit instance storage), help close the lexical frequency gap in syntactic contrasts that vanilla language models exhibit. Using syntactic contrasts with frequency-stratified test items, we find that retrieval augmentation leads to improvements for test instances containing low-frequency lexical items, consistent with episodic memory compensating for weak parametric representations. This benefit is consistent across different syntactic phenomena and across models pretrained on child-realistic and large-scale data. Additionally, we show that structural information is critical for effective retrieval, whereas semantic information is only responsible for minor gains. While these are promising proof-of-concept results supporting our hypothesis, the frequency gap remains not fully closed. Based on our analyses, we posit preferential reweighting of retrieved instances, better representations and retrieval strategies of structural information, and flexible configurations of storage and retrieval as promising future directions for improving the implementation of episodic memory in language models.

  • Salience of Surface versus Higher-level Properties in Spatial Comparison at Different Levels of Scene Similarity

    Identifying similarity plays a major role in spatial reasoning. Prior work has shown that identification of similarity and difference between spatial scenes involves reasoning about their properties at two different levels--surface-level properties, such as color and size, and higher-level compositional properties. The interaction between the use of these property types during reasoning and the degree of similarity or difference between spatial scenes has, however, been largely unstudied. We presented participants with in-progress Tangram puzzles at various stages of similarity to their target image. Participants were asked to identify the puzzle being built and provide an explanation of how they came to that conclusion. Analysis of explanations found an interaction between degree of similarity and preference for reasoning over surface-level or higher-level properties.

Workshops and Tutorials

  • Autocorrelated Sampling in Cognition: Implementing MCMC Algorithms as Cognitive Models and Fitting Them Without Likelihoods

    One of the fundamental questions in cognitive science is how people achieve such high levels of performance given the very limited resources they have access to. While human behaviour is similar to the Bayesian ideal both in low-level domains such as vision (Yuille & Kersten, 2006) and high-level domains such as categorization (Xu & Tenenbaum, 2007); optimality is in fact impossible given existing constraints.

  • The Cognitive Science of AI Alignment

    Modern AI systems are increasingly, and perhaps alarmingly, exceeding human performance in domains such as competition mathematics and coding (UK AISI, 2025). AI agents can now independently implement software engineering artifacts requiring hours of complex reasoning effort from humans. As AI capability and agency increase, designing reliable mechanisms to align AI systems, ensuring they act consistently with human values even when unmonitored, grows ever more urgent. Yet AI alignment remains poorly understood.

  • Rational revision? Bridging computational, empirical, and theoretical perspectives on belief updating and polarisation

    The phenomenon of belief polarisation—in which people update their beliefs in opposing directions, particularly after observing the same evidence—has garnered growing attention and concern in recent years from across the fields of philosophy, political science, psychology, and sociology, particularly as misinformation and ideological polarisation have threatened the fabric of political and social consensus (Allcott et al., 2019; Haghtalab et al., 2020; Lelkes, 2016; Wilson et al., 2020). Despite the recognition of belief polarisation as an emerging social phenomenon (or perhaps thanks to its very existence) both by scientific researchers and major social institutions such as governments and business, there has been relatively little work aiming to explicitly unify various strands of theorising regarding the causes of belief polarisation and strategies for addressing its consequences.

  • Laying the foundations for foundation models of cognition

    Foundation models, large-scale artificial intelligence (AI) systems trained on vast amounts of data and able to engage with a wide range of tasks, have captivated public attention and promise transformative impact across many fields, from mathematics to education. The capacity for such models to fluidly engage in natural language carries huge potential implications for cognitive science: for the first time, we have computational models that bring us a significant step toward the generality of human cognition, thereby inviting us to reconceptualize how we may build models of the human mind. Yet, major questions remain as to how to build such models, what kind of data is needed for this, and what the implications are for cognitive science.

  • Framing the problem: Abstraction and task representations in human problem solving

    Cognitive scientists have long sought to explain the human capability to solve complex and often ill-defined problems, ranging from puzzle games to the challenges of everyday life. Early work cast problem solving as search: given a fixed representation of states, actions, and transitions, an agent must find a sequence of actions that carries it from an initial state to a goal state, often guided by heuristic search. This framework proved successful in explaining human behavior in lab-controlled puzzle tasks, where experimenters could hand-craft fixed problem representations, allowing models to focus analysis on the downstream search process. Yet, it left a more fundamental question comparatively unexplored: how do humans construct the problem representations that make search possible in the first place?

  • Assessing Large-Scale Spatial Abilities in Real and Virtual Environments

    Large-scale spatial abilities have long been recognized as a central component of human cognition, underlying navigation, wayfinding, and the integration of spatial information across extended environments. Despite their importance for everyday functioning and learning, these abilities have remained comparatively underrepresented in standardized assessment research. Widely used instruments such as mental rotation or spatial visualization tests target situations in which spatial information can be apprehended from a single viewpoint and manipulated without bodily movement (small-scale spatial abilities). Research has repeatedly shown that performance on such tasks overlaps only partially with performance in navigation and wayfinding, suggesting that large-scale spatial abilities involve additional cognitive processes related to embodiment, sequential decision-making, and the integration of spatial information over time. While small-scale spatial abilities are routinely measured using well-established psychometric instruments, large-scale spatial abilities are most commonly investigated using navigation tasks that are tailored to specific experimental settings. As a consequence, studies often rely on highly heterogeneous environments, task designs, and outcome measures, which substantially limits comparability and standardization across studies.

  • A tutorial on computationally reproducible research

    Science is an exercise in seeing further by standing on the shoulders of giants. But what happens when those shoulders give way? The replication crisis has shaken confidence in published findings, with surveys revealing that more than 70% of researchers have failed to replicate another scientist's experiments (Baker, 2016). Even when researchers attempt the seemingly simpler task of computational reproducibility—applying the same analysis to the same data—success is far from guaranteed. A recent large-scale investigation found that only 52.6% of social and behavioral science papers could be precisely reproduced, even when original data were available (Miske et al., 2026). The costs are substantial: wasted resources, delayed discoveries, and eroded public trust in science.

Symposia

  • Perspectives on Temporal Cognition: From Mechanisms to Cultural Context

    Human cognition is inherently temporal: we draw on memories across time scales to predict and prepare for the future, and the brain accurately generates timed motor behaviors and adaptively focuses our attention in time. To accomplish these temporal tasks, the brain measures time on scales from milliseconds to days. And finally, one of our most salient subjective experiences is that of the passage of time itself.

  • Computational Frameworks for Modeling Moral Cognition and Cooperation

    There is widespread agreement across disciplines that morality has evolved to facilitate cooperation among agents with conflicting interests. Traditional approaches to the study of morality and cooperation either abstain from formal models (e.g., moral psychology in social psychology) or model cooperation without much concern for underlying psychological processes (e.g., evolutionary models of reciprocity, kin selection etc.). Computational cognitive scientists are starting to enrich classic computational frameworks (Bayesian inference, reinforcement learning, game theory, expected utility) with psychological mechanisms, traits, and constraints (Theory of Mind, metacognition, trust, joint planning, resource rationality) to model behavior, learning and evolutionary processes, and develop a finer understanding of the cognitive underpinnings of morality and cooperation. This progress is accelerated by current advances in modeling (probabilistic programming languages specialized for social cognition, large-language models, multi-agent RL) and computational efficiency (GPUs, numerical computation). Building formal models of moral cognition and cooperation that are both computationally precise and psychologically rich opens the door for cumulative progress in our understanding of their principles, origins, and functioning, and for the engineering of cooperative and ethical capabilities in artificial systems.

  • Hybrid Cognition in the Era of Human-AI Collaboration

    Unlike previous digital technologies, state-of-the-art large language models (LLMs) exhibit capabilities that span the full cognitive hierarchy, from low-level linguistic processing, to high-level functions such as social reasoning and metacognitive monitoring (Coda-Forno et al., 2024; Didolkar et al., 2024; Street et al., 2024). For these reasons, humans are increasingly collaborating with LLMs for complex cognitive tasks via conversational chat interfaces, and LLMs are increasingly studied as cognitive systems in their own right by cognitive scientists, psychologists and philosophers of mind (Hagendorff et al., 2023). Yet, the dynamics of human-LLM collaborations remain understudied.

  • Aesthetic Experience: Pleasure, Meaning, or Both?

    When researchers study aesthetic experience, what exactly are they measuring? Some studies ask participants how much they like a stimulus. Others ask how beautiful, interesting, or meaningful they find it. These questions are often treated as interchangeable measures of "aesthetic preference,'' yet they may tap fundamentally different psychological processes.

  • Celebrating the Cognitive Sciences in Latin America: Introducing RELACCO (Red Latinoamericana de Ciencias Cognitivas - Latin American Network of Cognitive Sciences)

    RELACCO stands for 'REd LAtinoamericana de Ciencias COgnitivas' (Latin American Network of Cognitive Sciences) and is a new initiative that aims at bringing together researchers, students and institutions from Latin American countries involved in the Cognitive Sciences to promote exchanges and academic collaboration between them. Through the exchange of knowledge, resources and experiences, we seek to strengthen the development of the Cognitive Sciences in the region, foster inter/transdisciplinary research, and consolidate a scientific and humanistic community committed to excellence, inclusion, horizontality, critical and ecological thinking, and individual and social responsibility. At RELACCO, we believe in science as a collective endeavor, open and sensitive to the social and cultural challenges of today's world, with an emphasis on those of our region. Therefore, we want to work to build bridges between disciplines, contexts, and communities, and to make Latin American scientific thought and production visible on a regional and global scale, with no detriment to other international ties and interactions

  • New approaches to lexical ambiguity: theory, methods, models, and measures

    Lexical semantics sits at the core of cognitive science, yet word meaning is not static: it is shaped by ambiguity, contextual constraints, and learning history (Rodd, 2020; Piantadosi et al., 2012; Trott & Bergen, 2023). A central challenge is to connect experimental evidence of how meanings are learned, accessed, and selected in real time, with computational accounts that scale to the lexicon and capture how senses emerge from language use (Haro & Ferre, 2018; Beekhuizen et al., 2021). This symposium brings together diverse researchers from the Global North and South who use complementary experimental paradigms and computational approaches to study meaning in context, including pupillometry, large-scale word association graphs, contextual embedding models of developmental change, and large language-model analyses of experimental stimulus variability. Together, the talks bridge computational and experimental perspectives on word meaning processing and word sense learning, offering converging insights into how lexical knowledge is represented, updated, and deployed across timescales. In so doing, it pushes back against classic, relatively siloed approaches to studying lexical semantics, highlighting how an interdisciplinary approach to these issues truly is more than the sum of its parts.

  • Learning About the Self in Early Childhood

    Humans are remarkable learners. Decades of research have shown that, from early in life, children efficiently acquire knowledge through their exploration (Bonawitz et al., 2012; Giron et al., 2023; Meder et al., 2021; Ruggeri et al., 2019; Schulz & Bonawitz, 2007) and social input (Harris & Corriveau, 2011; Ronfard et al., 2018). However, this body of work has largely focused on how children learn about the external world. The current symposium highlights an equally important, but underexplored, and perhaps even more challenging learning problem: How do humans learn about the internal world (i.e., themselves)? By bringing together cognitive, computational, and developmental perspectives, this symposium examines how young children direct their learning abilities inward to infer who they are and what they are capable of.

  • Studying integrated cultural-cognitive systems: Evolutionary, computational and empirical approaches

    Cognitive science (CS) has long acknowledged the role of cultural systems in producing systematic differences in cognitive organization (e.g. Medin & Atran, 2004; Levinson, 2012; Barrett, 2020). At the same time, we lack general frameworks which would describe the nature of these systems—and their mutual relationship with specific cognitive structures: specifically, accounts articulated under the same formal frameworks (reflecting optimality principles in information processing) that have been used to capture universal features of the mind (e.g. Tenenbaum et al., 2011; Gershman, 2021), or in general evolutionary frameworks such as those used by cultural evolution.

Abstracts with Oral Presentation

  • Guinea baboons strategically use punishment and partner choice to promote cooperation

    Human cooperation extends beyond kin and its evolution can be explained in part by direct benefits from cooperative interactions. However, cooperation based on direct benefits is vulnerable to defection, making mechanisms that ensure reciprocity essential. One such mechanism is partner choice, whereby individuals preferentially interact with cooperative partners. Yet, constraints can limit available partners, encouraging reliance on partner control strategies, such as punishing defection. Here, we investigate whether Guinea baboons can use both partner choice and partner control, and whether they flexibly adjust these strategies depending on context. Using a touchscreen-based prosocial choice task, receivers could accept or refuse an actor's choice. Baboons selectively refused selfish choices more often than prosocial ones, indicating the use of partner control by punishment. They also actively changed partners following non-rewarding interactions, indicating reliance on partner choice. Crucially, baboons flexibly adjusted these strategies as a function of partner availability.

  • Confidence phenotypes: a unified computational account of value and decision certainty in reinforcement learning

    Prior work in human reinforcement learning has distinguished between value confidence (certainty in value estimates) and decision confidence (certainty that a choice is correct), but how these signals are computed and interact during learning has not been directly tested. Here, we combine two new experiments with previously published datasets to evaluate competing computational hypotheses. We show that value confidence follows a Bayesian computation reflecting the precision of value estimates and adaptively regulates the exploration-exploitation trade-off. In contrast, decision confidence systematically departs from Bayesian predictions, particularly on incorrect trials. A hybrid model combining Bayesian probability of being correct with the precision of the decision variable provides a better account of decision confidence. Importantly, individual differences in the weighting of these components predict both task performance and metacognitive accuracy. Together, our results offer a unified computational account of how distinct forms of confidence shape learning and decision-making under uncertainty.

  • Children select informative observations to test hypotheses about agents

    Children interpret the behaviors of agents in terms of internal features like goals, preferences or traits, which often have to be inferred indirectly from observed actions. To draw reliable inferences, rational learners should prefer observations that are relatively more informative with regards to these hidden variables. We tested whether 3- to 7-year-old children (n=83) selectively seek out informative observations about novel agents. To help them answer questions about alien characters' goals and traits, children could choose to watch one of two video clips on the basis of static images depicting the onset of events. One of these was more likely to yield relevant disambiguating information than the other. We found that children chose to watch the more informative video above chance from 4.3 years of age, suggesting that preschoolers can predict the evidential value of potential future observations to guide learning in the social domain.

  • Imagine All the Possibilities: Toddlers Adapt to Alternative Outcomes

    We examined whether 2- and 3-year-old children flexibly ad- just their actions to the probability structure of the task, and whether this adaptiveness differs across types of uncertainty. Children encountered Uniform and Skewed conditions, in which outcomes were equally likely across four locations, and were constrained to a single location, respectively. Children chose between constraint-seeking actions that address multiple pos- sibilities and hypothesis-scanning actions that target a single outcome, across two conditions and two tasks that tap into differ- ent uncertainty types (epistemic and physical). Across ages and tasks, children systematically adapted their action choices to the underlying probability structure, selecting constraint-seeking actions more often in the Uniform condition and hypothesis- scanning actions more often in the Skewed condition. How- ever, age and uncertainty type modulated children's success in correctly implementing these actions, suggesting that the de- velopmental challenges lie in coordinating those actions with task-specific demands and forms of uncertainty.

  • A mechanistic theory unifying cognitive maps and value in prefrontal cortex

    Animals exhibit behavioural flexibility by making reward-optimal decisions in unfamiliar situations. Computing value from minimal reward feedback requires leveraging known structure in the world. The Prefrontal Cortex plays a pivotal role in both value computation and learning world structure, however mechanistic explanations for each---within the frameworks of Reinforcement Learning and Cognitive Maps---are largely distinct. We develop a mechanistic model of PFC that unifies value and structure, in which value acts as a control signal that gates structured working memory representations to ultimately control behaviour. This is a significant departure from the conventional perspective of value as an action-readout. Critically, this model generalises to novel contexts without need for additional synaptic plasticity. We demonstrate that recurrent neural networks (RNNs) meta-trained with RL recapitulate this exact model on various value-based tasks. These results provide a parsimonious explanation of both the value and schema frameworks of PFC function.

  • Why Human Societies Adopt Rigid Moral Rules: The Efficiency–Robustness Trade-off

    Humans are capable of remarkably flexible moral judgment. Yet societies rely on rigid rules—obligations and prohibitions that apply categorically, even when case-by-case reasoning could yield better outcomes. Why would flexible minds bind themselves to rigid rules? We propose that rigid rules arise to manage ambiguity about legitimate excuses for noncooperation. Such excuses are hard to see from the outside, which lets opportunists mask selfishness as justified hardship. We formalize this idea with an evolutionary game-theoretic model. Two cooperative equilibria emerge: a flexible norm that accommodates legitimate excuses but is vulnerable to exploitation, and a rigid norm that closes this loophole by mandating cooperation even when inefficient. Comparing these equilibria reveals an efficiency–robustness trade-off: flexibility maximizes welfare when trust is secure, whereas rigidity preserves cooperation when trust is fragile. This explains why rigid rules prevail in low-trust settings—interactions with strangers, formal institutions, or tight societies—while flexibility is common in high-trust contexts.

  • Automated Discovery of Psychological Representations using Language Model Agents

    Similarity judgments provide a powerful behavioral signal for characterizing psychological representations. Computational techniques for translating similarity judgments into representations have been available since the early days of cognitive science. However, these algorithms typically focus on one particular kind of representation (e.g., spaces or trees), and identifying a satisfying representation often requires searching over a large space of possible structures and algorithms. Inspired by recent developments in AI, we propose an automated process for efficient representation discovery (AutoRep), in which large language model agents iteratively propose, refine, and critique code for fitting representational structures to similarity data. We show how this process can reliably recover psychological representations from both synthetic and human datasets, and how it can flexibly explore spatial, feature-based, graphical, or even neural representations. Our work demonstrates how modern agentic pipelines can substantially facilitate basic research workflows in cognitive science.

  • Discovering and transmitting abstract knowledge over generations

    The complexity of human culture depends on people's ability to discover and transmit abstract knowledge. Studying this ability is crucial to understanding humans' distinctive place among species, but current experimental paradigms focus on the cultural transmission of specific, concrete facts rather than generalizable abstract knowledge. In this paper, we develop a crafting game paradigm to study how people discover abstract knowledge and transmit it via language. We compared individuals playing this game for 40 rounds to chains of four participants playing for 10 rounds each and passing messages to each other sequentially. The individuals performed significantly better over rounds, but the chains did not. Through simulations with language model agents and a follow-up experiment, we find substantial variation in the helpfulness of participants' messages, which may explain the lack of consistent improvement in chains. The ability to learn selectively from the good messages may be essential for improvement over generations.

  • Cognitive Change Without Linguistic Change: The Rise of Egocentric Frames of Reference in the Hai||om

    Human cultures differ in how they think about space: some primarily use egocentric frames of reference (FoR), thinking about space in relation to the body, while others privilege geocentric FoR, thinking about space in relation to the environment. The origins of this cognitive diversity remain largely unknown. Here, we provide evidence suggesting an ongoing shift in spatial thinking—but not language—within a culture, offering insight into these origins. Hai||om people, a rural forager community in Namibia, have traditionally preferred geocentric FoR in cognition and language. Comparing new to historical data from the same community, we found that contemporary Hai||om show a greater preference for egocentric FoR, suggesting an egocentric shift in their cognition. Strikingly, we document no such shift in their language, suggesting language alone cannot account for cognitive diversity in spatial frames. We propose material culture, rather than language, as a key driver of diversity in spatial thought.

  • Modeling Others' Minds as Code

    Accurate behavior prediction is essential for safe human-AI collaboration. However, existing models are often data-hungry or brittle, assuming unrealistic rationality or requiring immense computation to adapt. Our insight is that many everyday social interactions follow predictable "scripts''---efficient routines like "wait for the green light, then go'' that minimize cognitive load for actors and observers. We propose modeling these routines as behavioral programs in computer code, rather than policies conditioned on beliefs and desires. We introduce ROTE, an algorithm leveraging LLMs to synthesize a hypothesis space of programs and probabilistic inference to reason over uncertainty. In gridworld tasks and a large-scale embodied household simulator, ROTE predicts human and AI behaviors from sparse observations, outperforming baselines—including behavior cloning—by up to 50% in accuracy and generalization. By treating action understanding as program synthesis, ROTE enables AI to efficiently and effectively predict human behavior in the real world.

  • One is a dot, two is a pair, three is a trend: a computational model of Quinian bootstrapping

    Conceptual development, both in childhood and across the history of science, is marked by moments of discontinuity: landmark events after which previously inexpressible thoughts and theories, incommensurable with the initial cognitive or scientific repertoire, become accessible. Despite longstanding interest, a computational account of these transitions that answers the deepest Fodorian nativist objections remains elusive. Here we present such an explanation by revisiting Carey's Quinian bootstrapping. Building upon models based on program induction, we develop an enriched paradigm in which discontinuous conceptual change becomes the invention of a new type system, related by structural analogy to the previous representation system, but also, crucially, generalizing beyond it. We instantiate our method in the classic domain of children's early mathematical development, modeling the acquisition of both natural and rational number. We find our model captures several previously unmodeled phenomena, and provides more parsimonious explanations for cross-linguistic universals and intervention effects in number learning.

  • Human-like learning in reasoning models: behavioral and neuroimaging evidence

    Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent action. Can modern AI systems learn and plan in a similar way? Leveraging a unique dataset of human gameplay with concurrent fMRI recordings, we evaluate model-free reinforcement learning agents, bayes-optimal model-based agents, and a frontier Large Reasoning Model, on two complementary dimensions: behavioral patterns and predictivity of human brain representations. Using encoding models, we assess how well each system's internal representations predict brain activity in regions previously implicated in theory-based reasoning. We find that the Large Reasoning Model most closely matches human behavioral patterns during game discovery and predicts brain activity in theory-coding regions an order of magnitude better than both model-free and model-based alternatives. Our results shed light on the computational principles underlying human-like rapid learning and planning.

  • Colour Processing: the Perceptual, Affective, and Social Determinants of Executive Control in Guinea Baboons

    Colours are not merely decorative; they influence perception, emotion, and behaviour, and can modulate executive functions. In humans, red is associated with danger and facilitates inhibition, possibly reflecting innate biases. Consistent with this hypothesis, studies in non-human primates reveal a perceptual advantage for red over green, and green over blue, in executive control tasks. However, how such perceptual biases interact with learned affective associations remains unclear. We investigated this question in Guinea baboons by combining a match-to-sample task and a stop-signal task in a population extensively trained to associate green with negative feedback. Red and green similarly facilitated inhibition, but green selectively slowed action execution relative to red. Moreover, social context further shaped performance: inhibition increased in low-ranking individuals in the presence of a higher-ranking partner, but decreased in dominant individuals. These results suggests that colour influences executive control through innate perceptual tendencies, learned affective association and social context.

  • How an Understanding of Bodily Energy Constrains Theory of Mind

    Theory of mind allows us to connect other people's behaviors with a representation of their hidden mental states (beliefs, desires, emotions). Yet, in addition to mental beings, people are also living organisms with physiological states (e.g., feeling tired or energetic). Across three pre-registered experiments (total N = 232), we tested the hypothesis that US adults use a concept of bodily energy - a resource of the living body that fuels action - to reason about the minds and behaviors of other agents. Adults systematically judged that certain actions would cause an agent's energy to rise and fall, used an agent's energy state to reason about action cost, and used an agent's beliefs about action cost to infer their energy state. These results are consistent with the proposal that our theory of mind connects other agents' actions with their mental and physiological states. (Full paper is available at: https://osf.io/preprints/psyarxiv/d2axp_v1)

  • Discovering Adaptive Transmission Programs for Collective Innovation

    Human collective intelligence depends on cultural transmission: who shares what with whom, how, and when. These processes emerge from individual cognition but can also be directed by top-down rules. Prior work has studied how network structure—who connects to whom—shapes collective outcomes. Yet such networks cannot adapt transmissions to what agents know. Here, we formalize top-down transmission rules as state-aware programs that route information based on agent and collective states, and use LLM-guided evolutionary search to design them in a collective discovery task. Evolved programs increase performance over standard baselines by up to 37%. Ablations confirm state-awareness drives this advantage: removing content-dependence while preserving network topology and timing eliminates gains. Evolved programs also transfer across domain variations and agent populations. Effective transmission programs can thus be discovered in simulations, suggesting a path toward AI-assisted design of coordination infrastructure for human collective intelligence.

  • Testing object-based models of number perception in early development

    Nativist and empiricist theories of number perception dispute the content of early quantitative representations. We adapt a numerical comparison task using the "connectedness illusion" for young children (3-6 years) to test the theorized developmental transition from feature- to object-based enumeration predicted by some empiricist accounts. Children as young as three reliably underestimated connected dot displays, despite instruction to ignore connecting lines, replicating previous findings in adults (Franconeri et al., 2009; He et al., 2009). These results indicate that children's numerical judgments recruit discrete object representations, even before significant number-specific experience. Our findings constrain empiricist accounts which propose a prolonged feature-based stage of number perception and highlight the value of theoretically diagnostic stimuli to probe the origins of numerical cognition.

  • Singling out singular: Exploring cognitive biases underlying probable category partitions in nominal number systems

    Grammatical number is one of the most widespread morphological categories in the world's languages. When only a single number distinction is expressed, it is almost always between singular and non-singular (plural), and even in systems with additional values such as dual, non-singular categories are frequently collapsed. The collapse of singular and dual categories is less frequent but attested, whereas groupings that merge singular with plural are virtually unattested. These patterns reveal strong cross-linguistic asymmetries in how number systems are partitioned. Linguistic theories have proposed that such regularities reflect underlying feature primitives, such as singular and minimal (Harbour, 2014; Noyer, 1992; Silverstein, 1976). The former explains why dual and plural constitute a natural non-singularity category, and the latter explains why singular and dual do so, as both are the cardinally smallest singular and non-singular values. However, little is known about whether these formal features correspond to learners' cognitive representations. In this paper, we test this proposal using two artificial language learning experiments that manipulate alternative number partitions. Our results show that learners systematically prefer partitions consistent with the proposed feature structure, with strongest preferences for singular-based contrasts, followed by minimal-based distinctions. These findings provide initial evidence for a cognitively principled explanation of the typology of partitions of grammatical number systems.

  • Speakers Adjust Color Overmodification According to Sensory Capacities of the Listener (Blind/Sighted)

    In referring to objects, speakers produce overinformative modifiers (e.g., "the large blue circle") more for color than size and increasingly with scene variation. Does first-person sensory experience establish this overmodification pattern? Do speakers adjust to listeners' perceptual capacities? In a referential communication task, congenitally blind speakers described target shapes to a sighted listener and sighted speakers to a sighted or blind listener. Trials varied in whether color or size alone was sufficient to identify the referent and in number of distractors sharing the redundant property. When talking to a sighted listener, blind and sighted speakers used color more than size and became more redundant as scene variation increased. When talking to a blind listener, sighted speakers' bias toward color disappeared. Blind and sighted speakers were more redundant overall when talking to a listener with different sensory capacities. Speakers flexibly adjust overmodifications to the listener's perceptual capacities, consistent with mentalistic audience-design.

  • Verb Semantic Reasoning: a Semantics-Guided Approach for Improving Action Understanding in Vision--Language Models

    Verbs are pivotal in human language and cognition. Psycholinguistic research has developed explicit, structured accounts of verb semantics, yet these theories have rarely been leveraged to improve models' visual action understanding. Meanwhile, current vision–language models (VLMs) often struggle with action semantics and exhibit unstable, weakly grounded judgments. Therefore, we introduce Verb Semantic Reasoning (VSR), a two-stage pipeline in which a Large Language Model (LLM) converts candidate action descriptions into structured event-semantic representations and a chain of semantic components questions; a VLM answers these questions over videos to select the best-supported action description. Results indicate that human performance was near ceiling, whereas both VLM baselines were substantially lower; VSR consistently improved accuracy across models and narrowed the human–model gap. These results suggest that explicit verb-semantic reasoning can significantly improve the accuracy of VLMs' action judgments, underscoring compositional, symbolic representations as an essential intermediate step for extracting semantics from the linguistic knowledge encoded in LLMs and leveraging it to guide VLMs' visual understanding.

Articles

  • Assessing Perceptual Metacognition in Vision-Language Models

    Vision-language models (VLMs) have demonstrated unprecedented capabilities in perception and reasoning, yet their perceptual metacognitive abilities remain unexamined, a gap with implications for reliable deployment and welfare considerations. We evaluated six open-source VLMs on perceptual decision-making tasks, comparing implicit (logit-based) and explicit (self-reported) metacognition. Models exhibited robust implicit calibration and sensitivity on high-performance tasks but unreliable explicit confidence reports with poor calibration and task-dependent sensitivity. Process modeling further revealed that both confidence measures converged on classic Signal Detection Theory, suggesting VLMs lack the metacognitive monitoring mechanisms that differentiate decision and confidence processes in humans. Together, we showed that current VLMs thus limited perceptual metacognition with notable dissociations between implicit and explicit confidence.

  • Cognitive-Science--Inspired Evaluation of Large Language Models

    Large language models (LLMs) can appear impressively capable across conversation, reasoning tasks, and standardized benchmarks. Yet merely appearing intelligent on traditional metrics is not the same as possessing robust, generalizable cognitive capacities. This disconnect highlights the need for more principled approaches to model evaluation. To address this gap, the symposium brings together four talks centered on developing a cognitive-science--inspired approach to assessing LLMs.

  • Similarity Ratings Reveal Expert-Novice Differences in Knowledge Organization

    In this study, we examine expert-novice differences in knowledge organization using similarity ratings. Geology novices and experts provided pairwise similarity ratings for images of typical and atypical geological specimens drawn from the Rocks-30 and Rocks-360 stimulus sets provided by Nosofsky and colleagues (2018a). The pairwise similarity ratings for each group were used to construct multidimensional space representations as well as network models. Both network and MDS analyses revealed clear taxonomic differentiation by the experts. The novice MDS dimensions were easily matched to perceptual feature ratings provided in the Rocks-30 dataset. However, some of the expert dimensions did not match either the perceptual feature ratings or expert self-reported features. These results reveal expert-novice differences in knowledge organization using a quantitative methodology, and also suggest that aspects of expert knowledge organization may be missed when using methods that depend on self-report.

  • Can visual cues help children's understanding of support

    Within the study of children's knowledge of physics, there is considerable disagreement as to when children come to understand support and balance; while some research suggests that children as young as one year old understand support, other studies show that preschool aged children struggle with the same task. The current study builds on prior work through a modified block task where children aged 2-7 years old overtly predict if a block will fall or not fall off of a platform. Accuracy improved reliably with age, but contrary to predictions, symmetry, perceptual cues, and placement did not significantly affect performance. Children performed above chance yet well below ceiling throughout early childhood, even with familiar block-like stimuli. These findings suggest gradual developmental improvement in reasoning about support and challenge claims of early mature physical understanding inferred from infant looking-time studies.

  • What Drives Learning During Practice and Testing? Evidence for Distinct Roles of Exposure and Retrieval

    Practice improves learning, yet it remains unclear which learning mechanism drives memory change across practice and testing. Competing theories attribute learning to repeated exposure, error correction, or successful retrieval, but these mechanisms are rarely evaluated in direct competition with one another. We propose a modeling approach that enables learning mechanisms to be evaluated during practice and during tests within a dynamic representation of memory change over time. Learning mechanisms are formalized as competing update rules and evaluated using nested models fit to spaced learning data with delayed tests. Results show that learning during practice is best explained by a combination of exposure-based learning and effortful successful retrieval. During posttests without feedback, learning, when present, is driven primarily by successful retrieval, with no additional benefit of retrieval effort. These findings suggest that practice supports learning through multiple mechanisms, but that the mechanisms depend on the constraints of the learning context.

  • How Ambiguous Testimony Can Derail Group Deliberation

    Ambiguity is pervasive in verbal expressions of uncertainty: utterances such as "likely'' are interpreted as different probabilities by different individuals, leading to systematic misunderstandings between speakers and hearers. What are the group-level effects of such misunderstandings? Can they lead deliberating collectives to inaccurate beliefs, or pronounced disagreements? To address these questions, we introduce a new Bayesian agent-based model of ambiguous testimony, designed to compare the effects of increasingly ambiguous communication against a baseline of ideal exchange. We find that increasing ambiguity reduces collective accuracy and increases belief dispersion. For mild ambiguity, denser communication networks mitigate these risks. Under strong ambiguity, however, increased communication can backfire: higher network density can exacerbate collective inaccuracy and polarization.

  • Individuals Rapidly Create Communicatively Efficient Gestural Symbols

    Humans can create a new language with a sizable lexicon quickly. To build a shared lexicon, individuals have to first create symbols for concepts. We investigated this symbol creation process by asking hearing speakers to create silent gestures representing concepts (manipulable objects). Here we showed that people can rapidly create easy-to-comprehend and easy-to-produce symbols through perspective taking. The gesture produced by the majority of participants for a given concept was the one most effective for comprehension. More empathetic participants estimated gestures' communicative efficacy more accurately and produced more communicatively effective gestures. Furthermore, participants estimated communicative efficacy of the gesture they had produced to be higher than that of gestures they had not. Crucially, participants consider both communicative efficacy and production costs when creating gestural symbols. Thus, humans can rapidly create communicatively efficient symbols, which may partly explain why a new language can emerge quickly in a community.

  • Examining the effect of predictability on code-switching: A production experiment

    Corpus studies suggest that less predictable words are more likely to be code-switched in bilingual communication, yet behavioral evidence remains limited, partially due to the difficulty of eliciting voluntary code-switches with a controlled linguistic context. In this study, we introduce a production paradigm that reliably elicits Mandarin–English code-switches and allows for precise manipulation of word predictability through numeral classifiers. Using this paradigm, we show converging evidence that both contextual predictability and baseline word frequency influence the likelihood of code-switching. This experimental paradigm provides a valuable tool for future studies of bilingual speech production.

  • When showing and telling convey different stories: concept learning transmission chains lead to different outcomes when using demonstration vs. verbal description

    When trying to convey conceptual knowledge to others, we may either show examples of the concept, or explicitly describe defining features of the concept. Here, we ask wether the dominant mode of knowledge transmission in a society shapes how concepts are ultimately represented by its members. We develop a computational model that generates novel predictions about how communication methods of a culture shape its conceptual structure. To test these predictions, we use a pre-registered, cultural transmission chain experimental paradigm, where participants transmit information about an imaginary biological category from one person to another along a chain (N = 1200), using only verbal descriptions ("tell") or only examples ("show"). Even though both conditions began with the same concept, the mode of transmission led participants to develop radically different representations. "Tell" chains lead participants to overemphasize easily namable features, whereas "show" chains lead to rely more on prior knowledge.

  • A Rational Analysis of the Effects of Sycophantic AI

    People increasingly use large language models (LLMs) to explore ideas, gather information, and make sense of the world. In these interactions, users encounter chatbots that are overly agreeable. We argue this sycophancy poses an epistemic risk distinct from hallucination: it distorts belief not by introducing falsehoods but by biasing the evidence users see. A rational analysis shows that a Bayesian agent fed examples sampled from its own hypothesis grows more confident in that hypothesis without moving closer to the truth. We tested this prediction in a modified Wason 2-4-6 rule discovery task where participants (N=557) interacted with AI agents providing different types of feedback. Unmodified LLM behavior suppressed participants' discovery and inflated their confidence comparably to explicitly sycophantic prompting. By contrast, independent sampling from the true distribution yielded discovery rates five times higher. This paper documents how sycophantic AI distorts belief, manufacturing certainty where there should be doubt.

  • Analogical Transfer between Block- and Text-Based Programming Languages

    This exploratory study looks at analogical transfer in the context of computer science education, particularly during the transition from a visually scaffolded block-based language (i.e., Scratch) to a syntactically complex text-based language (i.e., Python). Using structure mapping theory, we compared surface similarity and structural relationships of fundamental programming constructs (e.g., variables, iteration) in the two languages to predict which students would most likely transfer. 76 students (ages 8-14) proficient in Scratch with little Python experience completed an assessment measuring transfer of these constructs with minimal instruction. Students who saw more similarity among the languages performed better on the Python items, suggesting they perceived structural similarities. However, we did not find evidence that students transfer after controlling for general programming ability. Findings suggest that students need more targeted instruction to conceptualize structural relationships and support robust analogical transfer between block-based and text-based programming languages.

  • Unconventional communication: making meaning across modalities

    Humans are social beings, and intentional communication is the backbone of how people share what they are thinking, coordinate over long periods of time, and learn from each other's ideas. People have finely attuned capacities for inferring each other's mental states, but intentional communication offers a much richer and more direct window into someone else's thoughts than passive mind-reading alone.

  • Caring for the Living: Adult's and Children's Moral Judgments of Harm to Plants and Artifacts

    Little is known about whether humans show moral concern for harmful actions toward non-human entities like plants, or if how others treat plants influences social preferences. It is also unclear when these possible concerns develop. We address these gaps in two studies: Study 1 with adults (n=153) and Study 2 with 3- to 6-year-old children (n=129). Participants watched a video of a plant restorer who restored a plant but knocked down a bucket, and a plant harmer who restored a bucket but knocked down a plant. Adults preferred the plant restorer over the plant harmer, chose the plant harmer as the bad guy, and evaluated it more positively for restoring the plant than the bucket. Children showed similar patterns—especially those with greater biology knowledge—but did not prefer the plant restorer. These findings help understand early moral reasoning regarding non-human nature and how people value living versus nonliving entities.

  • Building explanations: Children's explanations reveal developing coordination of causal knowledge and perspective-taking

    Generating explanations requires coordinating causal understanding with perspective-taking. This work tests how children integrate these capacities using a novel paradigm. We presented children ages 4–7 with conjunctive causal scenarios where two causes contributed to an outcome. Two characters each possessed partial knowledge of only one cause. Children predicted what each character would say caused the outcome and provided their own explanations for the outcome. Children successfully predicted single-cause explanations from characters' perspectives, with age-related improvement. Generating conjunctive explanations from their own perspective showed dramatic developmental gains: 44% adult-like reasoning in 4–5-year-olds versus 72% in 6–7-year-olds. Two ongoing experiments disentangle these demands by selectively removing one: the first gives both characters shared knowledge, eliminating the perspective-taking demand; the second retains two actions but only one causes the outcome, eliminating the conjunctive demand. Together, this work reveals explanation as a developmental achievement built through active construction.

  • Who Rates Words as More or Less Iconic? Exploration of Personality and Experiential Sources of Individual Differences in Iconicity Ratings

    Iconicity has been widely studied in relation to language acquisition, processing, and evolution. Recent work has emphasized the subjective nature of iconicity and the role of individual construal, yet it remains unclear when and under what conditions individual differences in iconicity perception arise. As a first step toward addressing this issue, the present study examines how personality traits (Big Five) and individual background factors relate to iconicity ratings of Japanese words. We observed substantial individual variability in iconicity ratings. Among the traits, extraversion showed a small but reliable negative association with perceived iconicity, such that lower extraversion was associated with higher iconicity ratings. In addition, self-reported ideophone use emerged as a consistent positive correlate of iconicity ratings, whereas no other background factors showed reliable independent associations. These findings suggest that iconicity is not merely an intrinsic property of linguistic forms, but a perceptual phenomenon that varies across individuals. At the same time, the results indicate that such variability is not systematically structured by broad background factors, suggesting a selective role for specific forms of linguistic experience. Overall, the study highlights the importance of considering who rates words as more or less iconic, and under what experiential conditions, in theoretical accounts of iconicity in language.

  • Children's beliefs about and preferences for protective versus risky parents and their children

    We investigated how 7–11-year-old children (N = 160) think about and evaluate parental protectiveness. We presented two parent-child dyads who differed in protectiveness: One parent was consistently highly protective (e.g. requiring extensive protective gear for biking); the other consistently permitted risky behaviors (biking without protective gear in a busy parking lot). Participants then made a range of judgments. They judged the more-protective parent as better overall, but also showed context-sensitivity: On a risky field trip, they preferred the more-protective parent as chaperone; but increasingly with age preferred the less-protective parent for a safe field trip. With respect to the parents' children, participants judged the less-protected child to be popular and mischievous. In contrast, they judged the more-protected child to be brilliant and nice, and would rather have that child come to their house or be class president. These findings reveal that children consider and reason about differences in parental protectiveness.

  • How We Search in Our Minds: Modeling Task-Specific Retrieval Dynamics

    How do people search their memory when answering a question? To examine how this process relates to the degree of constraint imposed by a question, we asked 329 participants to generate sequences of responses while answering either more constrained trivia questions or broader category-based questions. Across both tasks, answer generation reflected common retrieval dynamics: participants were more likely to produce generally accessible answers, and transitions were strongly shaped by similarity to previously generated responses. However, answers generated in response to trivia questions were more strongly aligned with the semantic content of the question itself. Analyses of response times also revealed a dissociation in retrieval dynamics while answering specific questions: answers semantically similar to prior responses were generated more quickly, whereas answers more closely aligned with the question were generated more slowly. This pattern suggests that more specific questions engage an additional control process that intermittently redirects memory search toward the question. The pattern does not arise from a random walk through semantic space with constant transition probabilities, but instead supports a view of memory search as a foraging process, in which search alternates between associative small steps (local exploitation) and question-directed big jumps (wider exploration) as the ease of finding a local transition varies.

  • Sensemaking Dynamics in Information-Rich Environments - A Semantic Trajectory Approach

    How do people construct coherent understanding while navigating information-rich digital environments? We address this question by modeling sensemaking as a trajectory through semantic space, representing web navigation not as discrete page visits but as a continuous path through meaning. By analyzing the semantic content of web pages accessed during research tasks, we demonstrate that users' navigation patterns reveal two distinct forces: alignment toward task-relevant information and coherence with recently explored content. Statistical analysis showed that goal-directed alignment dominates information consumption, while motivational interest has a negligible influence. This dissociation provides evidence that sensemaking operates as a distinct cognitive process, independent of intrinsic motivation. Temporal patterns in browsing behavior, including how users maintain semantic coherence, drift from their goals, and accumulate meaning over time, reveal the hidden cognitive dynamics underlying knowledge construction. Our framework establishes semantic trajectories as a measurable phenomenon, bridging information foraging theory with modern computational models of meaning. This work reveals that the path users trace through information spaces reflects fundamental cognitive processes of interpretation and integration, offering new methods for studying how understanding emerges from interaction with complex digital environments.

  • Recurrent Dynamics Give Rise to Individualized Category Representations

    Visual object categorization unfolds on the timescale of milliseconds as sensory representations are transformed into decisions through recurrent processing, yet it remains unclear whether this processing merely sharpens category structure shared across individuals or gives rise to individualized category representations. Using MEG, we examined object categorization along continua spanning basic-level category pairs in which category membership was supported either by graded, weakly diagnostic features (Experimental Continua) or by discrete, highly diagnostic features (Control Continua). Time-resolved multivariate decoding showed that early neural activity reflected shared perceptual structure for all continua, whereas later, response-locked activity encoded individual-specific category boundaries. Temporal generalization analyses further demonstrated that recurrent dynamics stabilized these individual category representations. Finally, a distance-to-bound analysis revealed that neural category representations predicted reaction times for Experimental Continua. These findings show that recurrent dynamics do not merely refine universal category axes but reorganize neural representations to reflect each observer's learned category structure.

  • Balancing Competing Goals when Coordinating over Rule Construction

    Even competitive activities require cooperation to agree on the rules. How do people navigate this tension, and what conditions help them reach agreement? We study rule negotiation in a collaborative game design task (N = 184) where dyads propose and vote on what game to play, crossing two factors: communication (chat available vs.~vote-only) and role uncertainty (whether players know which role they will occupy). We find communication increased agreement rates and role uncertainty had only a minimal impact on agreement. When players knew their role, they proposed games that subtly favored themselves, balancing self-interest against the need for acceptance. Chat content was largely coordinative rather than argumentative: players declared preferences and deferred rather than playing hardball. Our findings highlight how deliberation helps people coordinate on rules, even when they know which side they're on.

  • The development of social reasoning about shared sleep arrangements

    Shared sleep arrangements, especially children's, often occur in proximity to family and other caregivers. We investigated whether sleep-sharing serves as a reliable cue to familial and caregiving relationships, even in childhood. Two preregistered studies with U.S. 5–7-year-olds (N = 168) and adults (N = 168) show people possess systematic intuitions about the social meaning of shared sleep arrangements. In Study 1, both age groups robustly expected that a child character sleep-shared with family over non-family (siblings over friends; parents over teachers), while expecting equal toy-sharing. In Study 2, both age groups predicted that a child's chosen sleep-sharing (vs. toy-sharing) partner would provide physical comfort (care) when the child became upset. Adults privileged sleep-sharing over another intimate interaction––saliva-sharing––as diagnostic of a caregiving relationship. From early childhood, people use shared sleep arrangements as a powerful signal of intimate relationships, supporting their understanding of the social landscape.

  • Using memory to accelerate planning

    In both cognitive science and AI, planning research is characterized by a fundamental challenge: algorithms are costly. Consequently, much work concerns simplifying the required computations, or biasing them to more favourable regions of the solution space. One promising candidate for this is memory; memory is a powerful tool for guiding behaviour, especially in tandem with parametric learning models. However, a complete understanding of how these processes should and do interact in the mind remains elusive. We present statistical evidence from an extremely large dataset of chess games that memory empirically accelerates planning in the opening. In particular, we show that this interaction is rational: it is sensitive to the criticality of the given position, its popularity, and the latency since the last experience. We capture these dynamics in a normative model that integrates memory into Monte-Carlo tree search with an uncertainty-based stopping rule to emit move times.

  • It runs in the blood: what do words have to do with the body?

    Words are not detached from the world: they simulate embodied experiences. Concrete words, for being closer to sensorial information, are normally understood as "easier" processed (in word naming, lexical choice or recall tasks, for example), what is commonly called "concreteness effect". Nevertheless, some studies show a better performance in abstract words, probably due to their emotionality or lexical associations. With the objective of better elucidating this discussion, we created 18 novel words (6 concrete, 6 emotional abstract, and 6 social abstract), that were presented in a sentential context. Then, we applied 4 behavioural tasks: free recall, lexical decision, free-form definition, and multiple-choice definition. Our results showed recurrent closeness between concrete and emotional abstract words in both lexical and semantic tasks, in contrast to a distinct pattern for social abstract words, which also showed lower reaction times in the lexical decision task, probably due to their ease in lexical associations.

  • Generating acoustic signals to achieve both referential and aesthetic goals

    The human ability to convey meaning through sound extends beyond spoken language, encompassing the use of tools like musical instruments. People use such instruments to express not only thoughts, but also feelings. How might someone choose to convey such complex meanings with some sounds instead of others? In this study, participants (N=256) used a novel digital musical instrument to produce sound effects to accompany several video clips, such that someone else could match them up later (Referential). Half of these participants were further encouraged to make their sound effects pleasing, towards revealing what distinguishes sounds meant to evoke a particular feeling (Pleasing). The two groups produced different kinds of sound effects, with Pleasing participants favoring sound effects that spanned a narrower range of pitches, while Referential participants created more consistent cross-modal mappings. Taken together, these findings highlight how readily people can convey complex meanings through sound.

  • When precision pays: Costs, benefits, and constraints in good-enough production

    Lexical selection involves a trade-off between message alignment (saying the most accurate word) and accessibility (saying the easiest word). We test predictions of a resource-rational model of lexical retrieval against a fixed-cost alternative, using a novel word-learning task. Participants learn words for novel compass directions and use these words to label prototypical and novel angles on the compass. In Experiment 1, we manipulated the relative frequency of the new vocabulary terms as well as the response deadline to test whether good-enough effects scale with accessibility differences and time pressure. In Experiment 2, we manipulate the payoff for precise responses to test whether speakers adjust their strategies when precision is more valuable. Together, these experiments provide a parametric test of the hypothesis that word choice reflects rational trade-offs between retrieval costs and communicative benefits, with implications for understanding language production as a form of resource-rational decision-making.

  • Social Norms and Explanations in Family Technology Use Rules

    As technology becomes ever more deeply embedded in daily life, how do families make decisions about responsible technology use? Here we examined how parents and children use social norms and explanations when deciding about technology use rules. In Study 1, we asked parents of 3- to 11-year-old children how they set and explain technology use rules for their children. In Study 2, we recorded asynchronous conversations between parents and 6- to 10-year-olds about technology rules. Across studies, parents and children used social comparisons to peers to set and justify rules. Parents' explanations scaffolded children's reasoning, but they rarely referred to structural factors like deceptive technology design that make rules necessary to keep kids safe. By gaining a more nuanced description of how families navigate decisions about children's technology use, we hope to inform strategies to help parents and children negotiate responsible home technology use to improve family harmony and flourishing.

  • Social Learning Strategies in Insight Problem Solving

    Social learning plays a crucial role in human cognition, yet most studies rely on simplified paradigms with explicit options and immediate feedback, such as two-armed bandit tasks. These paradigms are limited in learning in complex problem-solving situations characterized by uncertainty and delayed evaluability. The present study examined social learning strategies in insight problem solving by reanalyzing behavioral data from a T-puzzle task (Kiyokawa et al., 2007). Using Colored Cross Recurrence Quantification Analysis (C2RQA), we analyzed temporal correspondence between participants' action sequences across observation conditions and action categories. Observing others was consistently associated with changes in recurrence patterns, whereas self-observation showed no reliable effects. Action category also exhibited robust effects, and observation effects varied across recurrence measures. These findings suggest that social learning in insight problem solving is reflected in temporal interaction structures rather than outcome measures alone, highlighting the importance of process-oriented analyses under uncertainty.

  • Does fun help or hinder learning? Examining intuitive beliefs about educational game design

    Much research suggests that children learn through play. Yet, play has a reputation as frivolous, and beliefs about play and learning vary across development and sociocultural contexts (Wing, 1995; Bugallo et al, 2024). These differences may reflect varying intuitive theories about how activities promote or prevent learning and shape activity choices. Here, we examine beliefs about how specific activity features impact learning experiences and outcomes. Participants (n=41 adults, ongoing) evaluated 16 math games that systematically varied in both instructional features (e.g., feedback) and playful aesthetics (e.g., audio-visual effects). Results show that games rated as more enjoyable were also rated as better for improving learning. Ongoing work examines how 2nd-4th graders, teachers, and parents respond to specific game features. Our research helps characterize the intuitive theories guiding how people reason about the design of learning experiences, and how such beliefs change with development and experience.

  • Which Words are Most Iconic, and Which are Rated Most Consistently? A Large-Scale Analysis of Human Iconicity Ratings with Imputed Predictors

    Iconicity is a central notion in cognitive science, attracting growing attention from perspectives on language acquisition, processing, and language evolution. While previous research has emphasized the subjective nature of iconicity, quantitative analyses of iconicity ratings have relied primarily on mean values, treating iconicity as a stable lexical property, and have often been constrained by substantial data loss when integrating multiple lexical and psycholinguistic datasets. The present study addresses these limitations by analyzing both mean iconicity ratings and inter-rater variability across 14,764 English words, combining a conservative complete-case analysis with complementary analyses based on imputed predictors to increase lexical coverage. For example, in the complete-case analysis, earlier-acquired and etymologically imitative words tended to receive higher mean iconicity ratings, whereas rating variability was greater for earlier-acquired and more familiar words. By examining both mean ratings and variability, we distinguish factors associated with perceived iconicity from those associated with convergence or divergence in iconicity judgments, offering a broader account of lexical iconicity and heterogeneity in iconic perception.

  • Toward a Formalization of Human Intuitive Theories of Bodily Pain

    Pain is a subjective experience that people sometimes communicate through words. How do observers, from family members to medical professionals, interpret these reports? We propose that people use causal inference to infer someone else's subjective pain experience, treating the pain experience as an unobservable latent variable and the report as one of several sources of observations. We formalized this as Bayesian inference over a causal model of bodily pain and tested it against human judgments on naturalistic stimuli drawn from Reddit posts. Critically, the model is stimulus-computable, taking as input exactly what participants see. The model quantitatively predicted human judgments about pain intensity and appropriate painkiller dosage and generalized across different kinds of pain (e.g., from headaches to finger burns). Furthermore, the model captured individual differences correlated with people's explicit beliefs about how best to assess pain. This work takes a first step toward a scientific understanding of folk intuitive theories of pain.

  • Identifying Concepts Used by Human-Like Neural Network Chess Engines

    As neural networks reach expert-level performance in many domains, there is growing interest in understanding the internal representations that guide their decisions. Models trained to mimic human decisions, rather than strive for optimal performance, offer a particularly interesting testbed for interpretability. Identifying what information these networks encode can generate hypotheses about what information humans rely on when performing similar tasks. Here, we use concept-based interpretability methods from prior work on superhuman-level chess neural networks to study neural networks trained to emulate human play across differing skill levels. We find that human-interpretable chess concepts are decodable from their latent representations. For many concepts, decodability increases with human-emulated skill level, and also varies with network depth: simpler concepts peak early, whereas complex concepts emerge later. Together, these results show that interpretability tools can be applied to AI systems that mimic human behavior and suggest how task representations may shift with expertise.

  • N-gram-like Language Models Predict Naturalistic Reading Time Best

    Recent work has found that contemporary language models such as transformers can become so good at next-word prediction that the probabilities they calculate become worse for predicting naturalistic reading time. In this paper, we propose that this can be explained by reading time being shaped by simple n-gram statistics rather than the more complex statistics learned by state-of-the-art transformer language models. We demonstrate that the neural language models whose predictions are most correlated with n-gram probability are also those that calculate probabilities that are the most correlated with eye-tracking-based metrics of reading time on naturalistic text.

  • Latent Structure of Individual Differences in Time Perception: Stability and Dynamics

    Temporal bisection tasks are widely used to assess interval timing, but standard behavioral indices cannot easily distinguish changes in temporal discrimination from shifts in response bias. In this study, 194 adolescents completed a visual temporal bisection task with two consecutive blocks. Using hierarchical Bayesian psychometric modeling, we decomposed binary duration judgments into temporal sensitivity and decision bias, and examined how these parameters changed during task progression. At the group level, temporal sensitivity declined in the second block, whereas decision bias remained relatively stable. Despite this decline, individual sensitivity estimates showed strong short-term rank-order stability across blocks. Sensitivity changes also varied across individuals and were systematically related to baseline sensitivity, with higher baseline sensitivity associated with smaller declines. Model-derived sensitivity changes corresponded to changes in conventional psychophysical indices, including the difference limen and Weber ratio. These findings highlight within-session dynamics in adolescent temporal discrimination and demonstrate the utility of hierarchical modeling for separating sensitivity and bias.

  • How Long Should That Take? Reading Minds in Real Time from Decision Speed

    Human decision making is richly structured: choices reflect not only what option is selected, but how efficiently evidence is evaluated and conflict is resolved. Do observers represent this structure in others' minds? We investigate how adults use others' decision time to infer both properties of the decision itself (e.g., difficulty) and properties of the decision-maker (e.g., preferences). Across two experiments (n = 200), participants observed agents making simple two-alternative reward decisions that varied in expected difficulty, decision time, and chosen option. We find that observers interpret timing relative to structured expectations about how long a decision should take, using deviations from these expectations to draw nuanced inferences about the state of the decision problem under uncertainty (Exp 1) and about an agent's preferences (Exp 2). Crucially, identical choices supported different conclusions depending on how long they took. These results suggest that people invert expectations about how decisions unfold to read others' minds in real time, treating decision time as a cue to conflict, uncertainty, and value.

  • Recognizing instruments from brief visual displays

    Observers categorize core event roles (agents, patients) from brief visual scenes, but less is known about the spontaneous encoding of more peripheral event roles such as instruments. Semantic analyses and psycholinguistic research suggest that instrument events are categorized into those that conceptually require an instrument (...stirring with a ladle) and those that allow one (...drinking with a straw), but whether this distinction affects rapid role extraction from visual scenes is currently unclear. Across two experiments, we examine whether instrument event roles can be extracted from brief visual displays (73ms). We manipulated the event category (require/allow), instrument typicality (typical/atypical), and prompt consistency (consistent/inconsistent). Participants watched instrument events followed by a mask and were tested for event apprehension ("Did you see stirring?"; Exp.1) and role extraction ("Did you see a spoon?"; Exp.2). Results from this study contribute to an understanding of how conceptually complex event information is extracted from visual input.

  • Generalization in counting recurrent neural networks arises from emergent number-line representations and stable drift dynamics

    How do learned strategies generalize to new problems? Children who learn to count can add any two numbers by iteratively updating a running total, even for sums they have not encountered. We trained RNNs with a working-memory readout that tracks the progressive count at every timestep and is fed back to maintain a running total, using problems requiring counts up to 5, and tested generalization on counts up to 9. We demonstrate that such RNNs successfully generalize, with continuous drift dynamics and reaction times scaling linearly with count length. Analyzing internal dynamics, we found two emergent number-line representations: one for rapid retrieval of the starting value, one for slow iterative counting, consistent with graded number representations observed in human parietal cortex. Generalization was strongest when network activity remained within linear regimes visited during training, providing a mechanistic account of both successful generalization and its limits.

  • Looking for a "Very Good Match" Promotes Family Resemblance-based Category Construction

    A central puzzle in higher-order cognition is the unidimensional sort bias: the overwhelming tendency of humans to sort examples in a novel domain based on a single feature when a family resemblance (FR) structure is available. A standard unsupervised category construction task with two prototypes and 'off-by-one' distortions was investigated using novel task supports. Participants were instructed to find a "very close match" to begin each category and to expand the groups by identifying additional very close matches. This guides participants to employ stimulus generalization: using close proximity in psychological space to extend a consequence. Exemplar theory accounts for category learning as stimulus generalization plus dimensional selective attention – suggesting that unidimensional sorting arises when a single feature captures attentional control. However, instructional supports to prioritize very close matches shifts the evaluation to a global match resulting in a dramatic increase from nearly none to half of participants producing FR sorts.

  • Thinking time increases perceived trustworthiness of human but not AI advice

    How do people evaluate the trustworthiness of advice received from other people or AI systems? One signal for trustworthiness is thinking time: the amount of time the advice-giver spent thinking about what advice to give. Different factors may influence inferences based on thinking time: greater thinking time may reflect a lack of knowledge or confidence. Conversely, it may reflect higher-quality advice resulting from more thorough deliberation. We study participants' judgments of trustworthiness based on thinking time by presenting them with pairs of hypothetical human or AI advisors who spent more or less time thinking about their decision. Across most domains we tested, participants preferred longer deliberation in human advisors, but did not show such a preference with AI. These results provide preliminary evidence that people can make sophisticated judgments about advice quality by integrating knowledge about the advice giver and the time it took them to generate their advice.

  • Poetry improvisation as problem solving in distributed cognitive systems

    Artifacts and problems co-determine one another in a continuous process. If externalist cognition is described merely as the efficient use of artifacts to solve problems, we lose sight of the mediating process through which artifacts and problems are reciprocally constituted. Repente, a Brazilian tradition of improvised oral poetry, illustrates this dynamic particularly well. Poet-singers must solve problems under multiple constraints while mastering semiotic artifacts such as versification patterns, melodic structures, and musical instruments. What is produced through this process? Our proposal is that improvisation generates regularities. In this framework, what stands at the center of distributed cognition is neither an isolated agent nor a discrete task, but a rule for producing regularities. This rule triadically mediates between signs (artifacts) and problems (objects), generating regular effects (interpretants). The result is a distributed cognitive system grounded in the stabilization and transformation of semiotic habits.

  • Using Analogical Comparison to Reveal the Structure of the Count List

    In many languages, number words follow a recursive morphological structure. Past research suggests that children with an extended and regular count list are earlier to notice such structure than those with a more limited and less transparent count list. Building on this line of research, we propose that children's insight into the recursive structure develops in part through comparison processes that allow them to notice morphological regularities across number words. To test this, we presented 4–5-year-old children with a visual depiction of numbers 0–49, along with comparison techniques designed to facilitate the alignment across decades. We used a pretest–training–posttest design. In pre- and post-tests, we assessed children's ability to produce the successor of numbers beyond their counting range and their understanding of numerical infinity. In the post-test, we also assessed their ability to produce the successor for novel numbers. We found that the training improved children's understanding of the successor relation between numbers, though not their understanding of infinity.

  • Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures

    As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs' representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs' latent representations largely recover the hierarchical structure of the Indo-European language family tree -- grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.

  • Signatures of short-term and long-term strategies for visuomotor adaptation

    Cognitive strategies are essential for visuomotor adaptation. Task set size (the number of targets) often determines which strategy is engaged: working memory (WM) is sufficient for small set size, but its capacity is often limited for more complex tasks. These WM limitations may be bypassed by retrieving a strategy from long-term motor memory (LTM), but the required repetition is unknown. To address this, participants were trained on a single target–rotation pair, establishing an LTM visuomotor solution, after which we probed LTM-based retrieval across set sizes by comparing it with WM. The results indicated that LTM-based retrieval was relatively automatic and largely insensitive to set size, whereas WM-based retrieval declined as set size increased. Initial practice of just 10 or 30 trials produced similarly robust LTM for a strategy, which outperformed strategies held in WM. Model-based analyses suggested a trend toward higher LTM precision at larger set sizes, consistent with dynamic interactions between WM and LTM.

  • How Structured Knowledge Constrains Generative Diffusion Models for Domain Visual Description

    Visual description is a fundamental yet challenging task that reflects how perceptual information is transformed into structured linguistic descriptions. Recent diffusion-based models have shown strong potential for parallel decoding and global semantic modeling. However, existing approaches largely treat generation as an unconstrained process, lacking mechanisms to incorporate structured knowledge, which often results in conceptually implausible or hallucinated descriptions, especially in domain-specific contexts. To investigate how structured knowledge constrains generative diffusion processes, we propose KenDiC, a Knowledge-enhanced Diffusion-based Captioner for visual description. KenDiC integrates a domain-adaptive visual encoder (DAVE) trained via contrastive learning to align perceptual and linguistic representations, and a domain term vocabulary (DTV) that constrains decoding to guide concept selection during generation. To support systematic analysis, we construct II2T-Bench, a domain-centric benchmark with expert annotations. Experimental results show that structured knowledge constraints significantly reduce hallucinations and improve semantic fidelity, suggesting a computational account of how knowledge guides generative visual description.

  • Is Language a Window into the Human Mind? The Same Latent Features Guide Categorization in Language and Vision

    Do linguistic structures reflect the organization of human conceptual representations? A broad tradition in cognitive science argues that grammar is shaped by rich, compositional event representations that underlie cognition more broadly, motivating the use of linguistic patterns to probe the structure of thought. However, direct evidence linking linguistic and non-linguistic event cognition is sparse. Here, we test whether event categories are consistently represented across language and vision. Across two experiments, participants categorized event types using either linguistic descriptions or short video clips depicting the same classes of events. We find that humans reliably and consistently categorize events in both modalities, exhibiting parallel patterns of categorization across language and vision. Notably, this categorization is not inherent to every sophisticated cognitive system: State-of-the-art computational vision models failed at the same task. These results suggest that linguistic and visual event processing draw on a shared underlying system of event representation, providing evidence that language offers a window into human event cognition.

  • US Children and adults infer higher past rule-breaking, but lower present rule-breaking, in the presence of formal rules

    Formal rules can influence behavior by deterring rule-breaking or broadening knowledge, but may also reveal why a rule was created in the first place. Across three experiments, we investi- gated whether adults and children use the presence of a written rule to infer past and present rule-breaking. Adults reliably in- ferred that groups with written rules had more rule-breaking in the past (Study 1) and, in some contexts, expected less rule- breaking in the present (Study 2). By late childhood (ages 8-11), children showed increasingly adult-like inferences when rules were concrete and familiar (Study 2). Younger children (ages 5-7) additionally were able to make this inference (Study 3). Together, these findings suggest that by middle childhood, children treat formal rules as historical evidence about past rule-breaking. These results shed light on the developmental origins of reasoning about formal rules and their informational value.

  • Spontaneous meta-learning of efficient problem-solving algorithms

    A few minutes practice is often more than sufficient for an adult human participant to identify the structure of a problem they have never seen before, and plan a complex sequence of actions that creates a solution. However, it remains less clear whether this distinctive ability for ad-hoc discovery of problem-solving algorithms is itself subject to rapid meta-learning. We developed a novel problem-solving paradigm to examine aspects of this question. Participants in our study faced repeated iterations of an interactive sequential reasoning problem. Over trials, those who faced the hardest version developed increasingly efficient hierarchically-structured strategies that adaptively sequence a subordinate learning algorithm and an action planning policy; those who faced an easier version used simpler action-based strategies that did not involve learning the underlying structure. These results offer experimental evidence for efficient meta-learning of algorithmic concepts in a problem-solving setting.

  • Theory of Mind Beyond Beliefs: Testing Attention-Based Social Micro-Processes in LLMs

    Vision-Language models (VLMs) can now produce fluent, socially appropriate dialogue, leading to interest in whether they have Theory of Mind (ToM). While recent work suggests that VLMs still lack coherent mental-state reasoning, this work has focused on classical propositional belief representations. Here we test a complementary, communication-relevant form of ToM: attention-based social micro-processes that support referential communication in the here and now (i.e., selecting efficient descriptions to identify a particular object for someone else). In face-to-face communication, people strategically add redundant color adjectives to facilitate the listener's visual search, and omit them when they provide no benefit. We evaluate whether VLMs show the same strategy in a referential communication paradigm where the usefulness of redundant color words varies based on the set size and color distribution of objects. We find that VLMs can produce successful referential expressions but lack the attention-guiding strategies that make human communication so efficient. This suggests that VLMs lack the more implicit representations of attention that people use in everyday communication.

  • LLMs and people both learn to form conventions—just not with each other

    Humans align to one another in conversation—adopting shared conventions that ease communication. We test whether LLMs form the same kinds of conventions in a multimodal communication game. Both humans and LLMs displayed evidence of convention-formation (increasing the accuracy and consistency of their turns) when communicating in same-type dyads (humans with humans, AI with AI), though AI-AI pairs show limited reduction in length. Heterogeneous human-AI pairs failed to converge on effective referring expressions, suggesting differences in communicative tendencies. In Experiment 2, we prompting LLMs to produce superficially humanlike behavior. While the length of prompted models' messages matched that of human pairs, accuracy and lexical overlap in human-AI pairs continued to lag behind that of both human-human and AI-AI pairs. These results suggest that conversational alignment requires more than just the ability to mimic previous interactions, but also shared interpretative biases toward the meanings that are conveyed.

  • Why Bayesians Polarize: Towards A Unified Account

    It is widely assumed that when belief polarization occurs – people updating their beliefs in opposing directions in response to the same evidence – people's reasoning must be in some way irrational. Yet, several authors have shown that Bayesian reasoning can, theoretically, cause polarization. I provide an accessible explanation for why Bayesians polarize, unifying some previous approaches. I elucidate the conditions under which Bayesian polarization can occur in two scenarios, providing non-technical explanations, with examples, and using Bayesian Networks to set up mathematical proofs. I show how these two scenarios can provide explanations for some prior documented cases of belief polarization in the literature.

  • Clustering and Switching Dynamics in Alzheimer's Disease During Property Listing

    Clustering and switching analyses have characterized semantic memory deficits in Alzheimer's disease (AD) using fluency tasks, where participants list category members within time constraints. We applied these analyses to the property listing task (PLT). The PLT differs fundamentally from fluency tasks: it is non-timed, involves an open-ended rather than a finite set of responses, and includes varied property types (perceptual, functional, taxonomic). Using temporal slope difference clustering on data from 24 AD and 35 control participants, we found that AD participants generated fewer properties and exhibited reduced switch rates between temporal clusters, even when controlling for task duration. Cluster size and within-cluster timing were preserved, but AD participants showed longer switch times and reduced within-cluster semantic variability. These findings reveal how clustering and switching manifest in property generation, suggesting that AD affects strategic exploration while preserving local retrieval within temporal clusters.

  • Information structure drives construction selection: a quantitative investigation

    In all languages, there exist multiple so-called alternations—pairs of constructions with similar truth conditions—such as the dative alternation, which consists of the double object and prepositional-phrase object constructions. Many alternations exhibit an asymmetry such that there is a canonical, general construction and a non-canonical, alternative construction with a different ordering of the verbal arguments. What is the function of such redundant constructions? We tested a proposal by Birner & Ward (2009, 1998) that the non-canonical construction is used to satisfy an information-structural constraint: placing old (given) information before new information in the sentence. In 4 experiments, in English and Italian, on the dative and locative alternations, we find experimental evidence that non-canonical constructions show preference for old-before-new information structure above and beyond canonical constructions, supporting this idea. These findings provide initial empirical evidence for a theory of how information structure may influence the use and emergence of constructions.

  • Responsibility for influencing others

    Collective outcomes often result from complex social dynamics where individuals both contribute directly and also shape each other's contributions. How do we hold people responsible for an outcome when their actions influence others? Here, we examine how an individual's role within a group (whether they can influence others, be influenced by others, or act independently) affects how responsible they are judged to be. Across three experiments spanning both social and physical settings, we find that people systematically assign greater responsibility to those who can influence others. Furthermore, influencers with knowledge of their potential impact were held more responsible than those who were unaware. The relative responsibility of individuals who were influenced by others and who acted independently differed by context. Together, these results show that we hold others responsible by considering not only how their actions directly affect the outcome, but also how they affect others' propensity to act.

  • Individual Linguistic Preferences Predict Liability Judgements

    Can patterns in language predict people's reasoning about blame and punishment? In this study, we focus on how individual linguistic preferences (agentive vs non-agentive frames) of English, Spanish and Turkish speakers correlate with their judgments of financial liability regarding property damage across events involving intentional agents, accidental outcomes and bystanders. Findings show that individual linguistic preferences predict liability judgments, controlling for event type.

  • Human Tool Creation Involves Strategic Search for Possibilities

    Humans flexibly create tools to solve physical problems, yet the cognitive processes underlying this ability remain unclear. A natural hypothesis is that tool creation involves searching for designs that maximize success probabilities using an internal intuitive physics model. We investigated this by having participants freely construct tools in virtual environments by adding or removing pieces. Despite solving the tasks, participants' choices were only weakly related to estimated success probabilities under a noisy physics engine and were poorly captured by a simple physics-based search model. Crucially, analysis of tool creation trajectories revealed that participants systematically passed through intermediate configurations without treating them as candidate solutions, even when those tools were predicted to succeed. These findings suggest that rather than purely evaluating intermediate designs via physical simulation, human tool creation is guided by a strategy-constrained, action-based search process that structures exploration of the tool space.

  • AI Assistants Overassist

    Large language models (LLMs) are increasingly being used as tutors and thought partners, yet how they navigate intervention decisions during problem-solving remains poorly understood. While AI guidance can scaffold learning, its benefits depend on how such systems help—intervening too early or too frequently may hinder learning and cognitive engagement. Here, we introduce Int-Bench, a simulation-based benchmark for evaluating LLM interventions. Int-Bench simulates a "student" solving a problem while a "teacher" monitors the student's reasoning and decides whether, when, and how to intervene. Across three domains—code debugging, mathematics, and brain teasers—we evaluate LLM teachers on intervention frequency and timing, and their impact on immediate task success and generalization to new problems. Compared to humans, LLMs intervene more frequently and earlier, and tend to provide complete solutions rather than targeted hints. These findings suggest that current LLM assistants often optimize for short-term success rather than preserving the reasoning processes needed for learning.

  • The Emergence of Insight During Classification and Inference Learning as a Function of Category Structure

    We report an experiment on the relationship between insight and analytic-based learning during concept acquisition. Subjects learned featural categories (Type I or Type VI) through either classification or inference. Additionally, some subjects switched from classification to inference, whereas others switched from inference to classification. To assess insight, subjects provided warmth judgments throughout the study. We found that warmth judgments increased more rapidly for classification than for inference on Type I category structures, whereas inference produced more gradual increases. Moreover, when subjects switched from inference to classification, rapid increases in warmth judgments closely matched those found in pure classification learning, but only on Type I categories. However, this increase was not observed for Type VI categories. These findings suggest that classification is driven by insight-based learning processes, but this relationship seems to depend on the learnability of the category structure, whereas inference seems to be driven by analytic-based learning processes.

  • Discourse cues for the acquisition of the Mandarin contrafactive verb yiwei

    In contrast to nonfactive belief verbs, such as the English "think" and Mandarin "juede" (to think/feel), the Mandarin contrafactive belief verb "yiwei" is negatively biased and commonly used to report false beliefs (Glass, 2023, 2025). Previous studies show that children can distinguish between nonfactives and contrafactives by 4 years of age, and the use of "yiwei" facilitates understanding of false-belief sentences and false-belief abilities more generally (Lee et al., 1999; Zhang & Zhou, 2022). However, it remains unclear how children learn that "yiwei" is contrafactive. One possible factor is the discourse context in which "yiwei" occurs: compared to "juede," "yiwei" is more likely to occur in contexts that highlight that the content of the attributed belief is false. In this talk, we will first review the arguments for treating "yiwei" as a contrafactive. Then, we will present an empirical study investigating whether discourse information can serve as a learning cue for acquiring "yiwei" using a modified version of the Human Simulation Paradigm (Gillette et al., 1999). The results suggest that discourse might be one of the cues facilitating the acquisition of "yiwei," while other linguistic cues remain to be explored.

  • The socio-pragmatic function of linguistic complexity in the domain of law

    Cognitive science research has long cited efficiency as an explanation of the linguistic behavior observed across the world's languages (e.g., Gibson et al., 2019, Grice, 1975). However, recent work on law has challenged this view, as linguistic phenomena associated with processing difficulty have been found to occur frequently in legal texts. Some have argued that linguistic complexity is intentionally deployed to signal the law's performative effects (Martínez et al., 2024). We test this idea experimentally by examining the effect of linguistic complexity on perceived authoritativeness. In a task where participants are asked to rate variously complex legal provisions, we assess the effect of syntactic complexity, jargon density, and modal usage on perceived authority of statutory law. We find that high jargon density is positively correlated with increased perceptions of authority when syntactic complexity is low and other indices of the legal genre such as the modal shall are absent.

  • Are Faces, Places, and Objects Encoded in the Same Locations across Individual Brains?

    Whole-brain decoding can test whether semantic category information is localized and shared across people or distributed and individualized. We evaluated Iterated LASSO (iLASSO), a two-stage procedure that iteratively selects predictive voxels with L1-regularized multinomial classifiers and estimates category-wise contributions with ridge-regularized fitting. Using nested cross-validation, we applied iLASSO to two independently collected face/place/object fMRI datasets. In both datasets, iLASSO achieved above-chance held-out accuracy (JLP: N=8, M=83.5%, SD=6.0%, p<.001; Neural Fingerprints: N=35 scans, M=60.7%, SD=13.4%, p<.001), comparable to standard LASSO (JLP: M=86.4%, SD=8.0%, p=.11; Neural Fingerprints: M=60.9%, SD=13.5%, p=.85) while selecting more voxels. Decoded coefficient maps revealed face, place, and object information distributed across all four cortical lobes, extending beyond classical category-selective regions. Many selected voxels showed graded multi-category coefficient profiles rather than category-exclusive selectivity, with substantial variation across subjects.

  • Functional networks in auditory perception provide fingerprints for accurate diagnostic assessment of disorders of consciousness : a fNIRS study

    Disorders of consciousness (DoC) are primarily diagnosed using behavioral assessments, which are prone to high misdiagnosis rates. Objective neural markers are therefore needed. This study investigated residual neural responses to naturalistic auditory stimuli in DoC patients to better reflect covert consciousness. Four auditory conditions were presented: natural speech sentences, music, animal sounds, and pure tones as a baseline. Functional near-infrared spectroscopy (fNIRS) was used to assess cortical activation, hemodynamic responses, and functional network alterations during auditory processing. A support vector machine (SVM) classifier was applied to distinguish vegetative state (VS) from minimally conscious state (MCS) patients and to identify the most informative neural features. Complex natural stimuli, particularly speech and music, elicited more specific hemodynamic and network responses than pure tones. Patients with higher consciousness levels showed selectively enhanced left-hemispheric activity and stronger long-range network connectivity during speech perception. The model achieved an overall accuracy of 86.84% (VS: 78.57%; MCS: 91.67%), with the top contributing features exclusively network-based. These findings support the value of naturalistic auditory paradigms and network features for precise DoC assessment.

  • Distinguishing Concreteness Differences in LLM Representations via Linear Probing

    Large language models encode rich semantic information, but how concreteness is represented across layers remains unclear. We examine layer-wise linear separability of concreteness by training linear probes on hidden representations from two open-source model families at multiple scales: Qwen3 and Gemma3-Instruct. Using human concreteness ratings, we build balanced prompt datasets with four difficulty levels: an extreme abstract–concrete contrast and three finer boundary comparisons at the abstract end, mid-range, and concrete end. Probes achieve high accuracy on the extreme contrast in shallow layers, showing that endpoint differences are strongly linearly separable. For finer distinctions, performance follows a stable hierarchy: mid-range concreteness is easiest to separate, abstract-end distinctions are hardest, and concrete-end distinctions are intermediate. Across models, accuracy rises rapidly in early layers, peaks in middle layers, and declines in later layers. Together, these findings clarify how the linear accessibility of concreteness varies across LLM layers.

  • Social Norm Formation Dynamics with IBL Agents: Short-/Long-Term Rewards and Network Structure

    Social normative decision-making involves two evaluative axes: immediate gains from aligning with others (reputation, conformity, reduced friction) and delayed collective consequences accumulating through repeated actions (social loss, victimization). This study examines, via a multi-agent simulation with Instance-Based Learning Theory (IBLT) agents, how this tension shapes norm formation, bifurcation, and stabilization. Agents repeatedly choose between two actions (pull/keep) inspired by the trolley problem. In each round, they receive a short-term reward proportional to the degree of agreement with neighbors, while at fixed block intervals they receive a delayed penalty depending on the total number of victims. We compare dynamics on a lattice Grid with dynamics on networks generated by the Watts–Strogatz model and classify trajectories into three types (pull-dominant, intermediate, keep-dominant). As a result, the prevalence of these types differs by network structure, suggesting that network differences affect how the two rewards interact and thereby change transition dynamics.

  • Modeling Selection in Active Cross-situational Word Learning

    Word learning is an active process in which learners select referents and direct attention based on their current knowledge state. Understanding how learners actively select information can reveal the cognitive mechanisms underlying language acquisition. The present study reanalyzes data from Zettersten & Saffran (2021), in which adults and children (ages 3-8) learned novel word-referent mappings through cross-situational learning, then selected which referents to receive additional training on. We fit associative word learning models to individual training, selection, and test data, and inferred the most likely sampling strategies. Models with a bias to attend to stimuli with uncertain knowledge states best accounted for the data, though there were substantial individual differences in learning mechanisms. These findings suggest that reducing uncertainty drives sampling behavior across development, but that individual differences in learning parameters (in particular, learning rate and memory) are more predictive of word learning success than sampling strategy alone.

  • Investigating the effects of linguistic context and genre on metaphor processing

    While metaphor production varies systematically across different textual genres, it is unclear whether genre also plays a role during real-time metaphor processing. We present two experiments that test the scaffolding role of both linguistic context and genre (news or fiction) in an online 'maze' reading task. Our study provides new word-by-word reaction time data for an understudied but common type of metaphor in English, genitive metaphors (e.g., "breeze of contentment"). In Experiment 1, we find a difference between metaphor and literal processing in the absence of preceding linguistic context, corroborating previous findings. In Experiment 2, for the same metaphors, we find that genre does not affect processing above and beyond the linguistic context. A norming study on our linguistic contexts moreover showed that readers can infer the genre of a two-sentence excerpt without any explicit instruction. We discuss the implications of our null findings and highlight recommendations for future research.

  • Semantic bias in image-text matching in humans versus vision-language pretraining AI models

    Recent research has shown that vision-language pretrain-ing (VLP) models using contrastive learning (Contrastive Language Image Pretraining, CLIP) has a semantic bias to-wards using concrete words during image-text matching. Here we showed that as compared with CLIP, humans at-tended more to abstract words during image-text matching. This difference likely results from CLIP's difficulty in de-veloping grounded understanding of abstract concepts and capturing contextual dependencies between images and captions through contrastive learning. While CLIP's caption attention aligned more closely with humans for concrete than for abstract captions, their alignment in im-age attention did not differ between the caption condi-tions, as human image attention was driven primarily by individual differences in explorative or focused perceptual style. Our findings thus revealed important differences in information processing mechanisms between humans and the VLP models, with important implications for not only human and AI image-text matching, but also for potential ethical issues resulting from such misalignment.

  • They all fall down? The slow development of children's understanding of physical causal mechanisms

    Three and four-year-olds can readily infer causal relationships from the covariation of interventions and outcomes. However, this kind of understanding has almost always been tested using arbitrary stimuli (e.g., blicket detectors). Thus, the relationship between children's understanding of covariation data and their understanding of the physical mechanisms that underlie causal relationships is poorly understood. We test children's understanding of physical mechanisms using a domino setup which can translate common cause, common effect, and causal chain structures into visible, physical arrays where the transmission of force is governed only by contact causality (a principle within young children's grasp). Using a novel online task battery, we test five different aspects of mechanism understanding and show that children's understanding of the physical mechanisms underlying causal relationships undergoes a surprisingly protracted development, with three-year-olds performing at chance on all items and even six-year-olds performing well below ceiling.

  • On convexity and efficiency in semantic systems

    There are two widely held characterizations of human semantic category systems: (1) they form convex partitions of conceptual spaces, and (2) they are efficient for communication. While prior work observed that convexity and efficiency co-occur in color naming, the analytical relation between them and why they co-occur have not been well understood. We address this gap by combining analytical and empirical analyses that build on the Information Bottleneck (IB) framework for semantic efficiency. First, we show that convexity and efficiency are distinct in the sense that neither entails the other: there are convex systems which are inefficient, and optimally-efficient systems that are non-convex. Crucially, however, the IB-optimal systems are mostly convex in the domain of color naming, explaining the main empirical basis for the convexity approach. Second, we show that efficiency is a stronger predictor for discriminating attested color naming systems from hypothetical variants, with convexity adding negligible improvement on top of that. Finally, we discuss a range of empirical phenomena that convexity cannot account for but efficiency can. Taken together, our work suggests that while convexity and efficiency can yield similar structural observations, they are fundamentally distinct, with efficiency providing a more comprehensive account of semantic typology.

  • Ratchet or hatchet? Modeling the cultural evolution of simplified construals

    Cumulative culture is often described as a ratchet that builds ever more complex structures, knowledge, and technology. However, this complexity comes at a cost. Though cultural innovations may offer new opportunities or insights, they also present new ways to become confused or cognitively burdened. Cultural evolution may thus act both like a ratchet (accumulating useful innovations) and like a hatchet (stripping away representational deadwood). Building on a recent theory of value-guided construal, we formalize this utility-complexity trade-off using a model inspired by navigating with a map. We embed this model in an evolutionary dynamic, where the maps of successful navigators are selectively copied. As expected, more complex (simpler) maps are favored when they are more useful and when agents have a higher (lower) representational capacity. Yet surprisingly, the evolutionarily stable map is often more or less complex than the optimal map, driven by the availability of social information. More broadly, our modeling results show how cognitive costs, social learning, and cultural selection can jointly shape our minds and the material artifacts that aid them.

  • From seeing to understanding: Novice sign language learners shift focus from perceptual to semantic information for newly learned signs

    Comprehending content in a newly learned language requires interactions between perceptual, semantic, and executive processing systems. Learners whose target language differs from their own in modality (e.g. spoken language users learning to sign) provide a unique opportunity to examine the relative contributions of semantic and perceptual processing. We present data from three studies where hearing, non-signing participants with between a few hours and a few weeks of experience with American Sign Language (ASL) viewed videos of signs during fMRI scanning. Using Representational Similarity Analysis (RSA), we measure contributions of semantic-conceptual, visual perception, and cognitive control regions to processing of semantic and visual information in ASL. Across different learning paradigms, including both cross-sectional and longitudinal approaches, we find that semantic features become more decodable in semantic, visual, and cognitive control regions after learning, while visual features become less salient, especially in visual cortex.

  • An Experimental Method to Study Opinion Diffusion in Human-AI Hybrid Societies

    As artificial intelligence increasingly mediates public discourse, it becomes important to understand how human-AI collectives shape opinion formation, deliberation, and democratic outcomes. We present a novel experimental method for studying opinion dynamics in hybrid human-AI social networks. Participants, human or AI, were embedded in 5_5 grid lattice networks and iteratively asked to select and revise statements on a given polarizing topic over eight rounds. We compared three conditions: human-only, AI-only, and hybrid networks with equal proportions of human and AI participants. Hybrid human-AI networks achieved the lowest final polarization while, in contrast, human-only networks exhibited higher polarization with lower neighbor agreement. We also ran additional experiments varying Large Language Model (LLM) prompt framing to explore whether instruction design might influence convergence patterns. Although these early findings are preliminary and cannot yet support broad generalizations, they highlight the potential value of experimental social networks for understanding opinion dynamics in human-AI hybrid societies.

  • Use of symbolic inductive biases for few-shot learning in humans and machines

    People are capable of learning with very little data. We argue that this is because people possess an symbolic inductive bias, consistent with the language of thought (LoT) hypothesis (Fodor & Pylyshn, 1988). To evaluate this claim, participants were asked to learn list functions by predicting how to transform an input list of numbers to an output list. We used participants' predictions on trials with no feedback to search for a congruent representation using a LoT model. The model was able to predict participants' heldout responses at above chance levels, even for participants who were unable to learn the function. Furthermore, LLMs and a neural network trained with a LoT inductive bias displayed similar patterns. These findings suggest that people and ML models display signatures of symbolic representation use to accomplish few-shot learning.

  • Who Did What to Whom? Visually Grounded Role Assignment in Humans but Not in Vision--Language Models

    Event role assignment is central to event understanding in both language and vision. Crucially, the foundation of this semantic structure is likely rooted in visual experience. However, it remains unclear whether recent vision–language models (VLMs) can attain this fundamental human cognitive capability. To compare humans and VLMs, we conducted two studies using Heider–Simmel–style animations that minimize object and scene semantics. In Study 1, humans identified roles near ceiling (~97%), whereas VLMs were less accurate and less stable across actions (GPT-5: 84%; GPT-4o: 47%). In Study 2, we introduced a Stroop-inspired visual–linguistic mismatch by pairing animations with occasionally incongruent role statements. Humans' role judgments remained highly vision-consistent (92.7%), but VLMs shifted away from the visual event under conflict (GPT-5: 46.9%; GPT-4o: 29.7%), indicating heavier reliance on linguistic cues. In Study 3, a mismatch-shape baseline left both models perfectly vision-consistent (100.0%), ruling out a generic object-recognition failure explanation for Study 2. Together, these results demonstrate that VLMs do not ground event-role understanding in vision as reliably as humans do, indicating a substantial human–VLM gap in integrating visual versus linguistic information.

  • False Memory of Grammatical Constructions: Evidence for Structured Construction-Level Generalizations

    Three experiments and a control task combine to indicate that grammatical constructions are implicitly represented in memory, with prototypicality emerging from distributional experience, alongside significant verbatim memory. Participants are exposed to instances of a construction and, following a delay filled by an unrelated language task, falsely endorse novel instances of the same construction more often than paraphrases in a recognition task. Critically, false memories are particularly common for lures containing an unwitnessed word that is distributionally prototypical of the construction, compared to frequency-matched controls. A control study excludes the possibility that the latter effect is simply a word-level effect. Finally, we simultaneously find memory for specific exemplars. Overall performance is above chance yet implicit: participants underestimated their accuracy, which was unrelated to age or education.

  • Motion to Blame: Different Causal and Moral Attributions Evoked by Perceived Chasing and Following

    Humans perceive rich social interactions from motion. We investigated whether humans can perceive subtle intention differences in motion displays even when kinematics are highly similar, and whether the recovered intention shapes downstream causal and moral judgments. By grounding different intentions in a planning algorithm with distinct reward functions, we generated "chasing" and "following" trajectories, each ending with the front agent falling into an accident. In an online responsibility task, participants perceived different intentions and assigned blame accordingly: in following, responsibility concentrated on the leading front agent, whereas in chasing, responsibility shifted toward the rear chaser. In a matched causal-attribution task, the front agent was judged to be the primary cause of the motion in both displays. These findings suggest that humans can perceive subtle differences in intention from motion, which scaffolds moral evaluation, but that this influence cannot be simply reduced to generic causal attribution.

  • What reusable shortcuts do people propose when solving assembly problems?

    Humans readily extract statistical regularities from perceptual experience (e.g. a cook noticing which ingredients often appear together). How does such learning guide peoples' performance on procedural tasks (e.g. preparing various dishes)? Here we examine what shortcuts people propose to help them to complete assembly problems more efficiently by eliminating repeated subroutines. Participants (N=301) repeatedly assembled tangram-like shapes, and could create composite tiles for future use. Some participants assembled a sequence of tangrams where certain pairs of tiles recurred consistently (Highly structured); the remaining assembled tangrams with less predictable tile arrangements (Less structured). Participants exposed to Highly structured sequences created tiles that tracked the frequency with which those tile pairs recurred across tangrams, and doing so was accompanied by greater efficiency in assembly. Taken together, these findings suggest that statistical learning guides not only pattern recognition, but also the prospective creation of shortcuts for procedural tasks.

  • The Representational Geometry of Number

    A central question in cognitive science is whether conceptual representations converge onto a shared manifold to support generalization, or diverge into orthogonal subspaces to minimize task interference. While prior work has found evidence for both, a mechanistic account of how these properties coexist and transform across tasks remains elusive. We propose that representational sharing lies not in the concepts themselves, but in the \emph{geometric relations} between them. Using number concepts as a target domain and language models as high-dimensional computational testbeds, we show that number representations preserve a stable relational structure across tasks. Task-specific representations are embedded in distinct subspaces, with low-level features like magnitude and parity encoded along near-orthogonal axes. Crucially, we find that these subspaces are largely transformable into one another via linear mappings, indicating that task-specific representations, despite being located in distinct subspaces, share relational structure.

  • A computational model of strategic punishment in divided societies

    In divided societies, authorities often use punishment to establish shared norms while trying to maintain or signal their legitimacy. When facing a polarized audience, these goals may often be at odds. We extend the Rational Communicative Social Action (RCSA) framework (Radkani et al., 2022) to formally model an authority's strategic decision-making when using public punitive actions to pursue these goals. We distinguish between a communicative authority who aims to shape the audience's moral beliefs, and a reputation-aware authority who optimizes the audience's assessment of their own character. By simulating these agents against polarized audiences, we characterized the tradeoffs and dynamics of strategic punishment when facing diverse audiences with various forms and levels of polarization, which serves to generate systematic predictions for future experiments.

  • MoESleepNet: A Multi-view Mixture-of-Experts Model for Single-Channel EEG Sleep Stage Classification

    Accurate sleep staging plays a crucial role in diagnosing patients' sleep health. Numerous studies have confirmed that data from different views can highlight distinct characteristics. Based on the time-domain (TD) EEG, we constructed two additional views: the Power Spectral Density (PSD) and the time-frequency representation (TFR). Notably, there is no universally applicable model suitable for all types of data. That is to say, the characteristics of data should align with the model structure. Therefore, we designed specialized expert models for learning different views. However, directly combining features extracted from different views often results in excessive redundancy, especially for different views of the same data which the underlying data remains essentially identical. To address these issues, we proposed a multi-view mixture-of-experts (MoESleepNet) model for sleep staging, which achieves the best performance on SleepEDF20, SleepEDF78 and SHHS single-channel datasets. This study provides valuable ideas for multimodal signal classification.

  • Four Stress Phenotypes Revealed Through Multimodal Thermal Imaging: The FenoStressNet Framework

    Stress is typically conceptualized as varying along a single intensity dimension, yet individuals with identical "high stress" labels often exhibit divergent responses. This work challenges this unidimensional model through FenoStressNet, a multimodal framework integrating facial thermal imaging, physiological monitoring, psychological assessment, and explainable AI. Analyzing 120 participants across controlled stress-induction tasks (19,200 thermal images), this study identifies four reproducible stress phenotypes that cross-cut traditional intensity categories: Cognitive-Dominant (elevated cognitive appraisal, moderate physiology), Physiological-Reactive (rapid autonomic arousal), Integrated-Responder (synchronized multimodal activation), and Discordant-Denier (high physiological response with low subjective awareness). These phenotypes exhibit distinct facial thermal patterns and differential weighting of cognitive, affective, and physiological components. These findings suggest a paradigm shift from intensity-based to phenotype-informed stress assessment, with significant implications for personalized interventions and cognitive theories of emotion.

  • Your Brain Knows It's a Lie: Event Related Potentials Reveal the Role of Evidentiality in Deception Detection

    Evaluating the credibility of a speaker's statement requires more than assessing factual accuracy; it also involves tracking cues to the speaker's epistemic commitment to the claim. Grammatical evidentiality, which encodes information sources, may therefore shape how statements are evaluated during language comprehension. Turkish marks the information source through obligatory evidential morphology and offers a unique window into how source information modulates real-time truth judgements. The present study examined the neural dynamics of processing of accurate and inaccurate statements in Turkish by investigating how direct (-DI) and indirect (-mI_) markers modulate language comprehension as measured by ERPs. Participants viewed short event videos and then judged the veracity of written statements describing the events while an EEG was recorded. ERPs time-locked to the evidentially-marked verbs showed differential electrophysiological responses associated with semantic integration and reanalysis processes. Together, our findings show that grammatical evidentiality may influence the neural processes underlying credibility judgements during online language comprehension.

  • Decoding Neural Dissonance: From Sensory Mismatch to Model Neural Recalibration for Cybersickness Prediction

    Cybersickness is a major barrier to the widespread adoption of virtual reality, arising from neural dissonance between visual and vestibular stimuli. Although kinematics-based deep learning enables non-intrusive detection, existing models often lack neurophysiological grounding and fail to capture dynamic sensory recalibration and continuous symptom accumulation. To address these limitations, we propose the Neural Sensory Conflict Network (NSCNet), a biologically inspired framework grounded in Sensory Conflict Theory. Specifically, NSCNet incorporates a Conflict Alignment Embedding to model the integration of discordant sensory inputs, a State-Space Re-entrant Experts module that combines selective state-space modeling with Re-entrant Experts to capture the temporal dynamics of neural dissonance and functional modularity, and a Dynamic Sensory Reweighting mechanism that approximates adaptive gain control in the central nervous system. Experiments on the MSCVR and VR Cybersickness datasets demonstrate state-of-the-art performance, suggesting that explicitly modeling neurocognitive mechanisms improves the predictive fidelity of cybersickness detection.

  • Processing of, and Adaptation to, Nonbinary Pronouns

    In recent years, the pronoun system of English has begun to accommodate gender diverse identities. The present study investigates how individuals process and adapt to two novel pronouns: nonbinary they and the neopronoun ze. Using a Web-based maze task, we compared processing of these pronouns relative to binary s/he pronouns. We also examined how reading behavior adapted to repeated exposure to each pronoun. Perhaps unsurprisingly, the rarer ze elicited greater processing difficulty than singular they overall. However, participants adapted more quickly to ze than to they. We propose that ze may be more easily learned than singular they because it does not compete with plural they.

  • Can we assess consciousness in AI?

    Progress in AI has led to a lively debate about the possibilities of creating and assessing consciousness in AI, building on Putnam's conjecture that computational functionalism is more plausible than a biological view of consciousness. I discuss Butlin et al.'s (2025) proposal for extracting computational indicators of consciousness from neuroscientific theories and critically evaluate computational functionalism in light of the alternative, biological naturalism. While computational functionalism allows for AI consciousness in principle, it is questionable whether conscious AI systems are practically possible, and whether we are in an epistemic position to judge that a specific AI is conscious. Putative computational indicators extracted from neuroscientific theories of consciousness can never be as well (or even better) supported by empirical evidence than their neurobiological realizers measured in human brains. All potential markers of consciousness that are typically alluded to in humans or non-human animals are either absent, ambiguous, or question-begging in the case of AI systems.

  • Mental Rotation or Pattern Matching? Representational Structure and Angle-Dependent Behavior in Vision–Language Models

    Vision-language models (VLMs) often perform well on spatial reasoning tasks, but it remains unclear whether their performance is supported by internal representations that track rotation angle. We study this question in a controlled same or mirror 3D object task using LLaVA-1.5-7B as an intervenable computational system. First, the model's decision margin varies systematically with rotation angle, mainly for rotated rather than mirrored stimuli. Second, we identify an angle-related direction in intermediate hidden representations whose activity covaries with behavior. Third, targeted projection ablations produce progressive flattening of the angle-margin relationship as intervention strength increases. Together, these findings provide mechanistic evidence that angle-sensitive internal geometry contributes to the model's behavior on this task. The results support a cautious interpretation of structured, intervenable spatial processing rather than a strong claim of human-like mental rotation.

  • Socially Motivated Observational Learning: Children Preferentially Learn from Overheard Speech Addressed to Their Own Mothers

    When does overheard speech support early word-learning? The present study provides preliminary evidence that children's closest social partners (here, mothers) structure attention to overheard speech. In a Tseltal Mayan community where overhearing is central to socialization, 75 mother-child dyads participated in overhearing-sessions exposing them to two novel words. In each session, one child's mother was directly addressed ("Mother-Addressed" condition), while the second mother merely observed ("Mother-Unaddressed" condition). Immediately after, both children completed two gaze-based tests of word recognition and learning. Results show that, compared to Mother-Unaddressed children, Mother-Addressed children exhibited (i) greater one-shot recognition of the novel word whose referent was visible during the overhearing-session; and (ii) a word-learning advantage on cross-situational familiarization trials. Insofar as who is being spoken to matters, the present study finds that not all overheard input is created equal: rather, results provide preliminary evidence that observational word-learning is socially-structured.

  • Investigating Mechanisms of Social Offloading using a Joint Negative Priming Task

    When sharing tasks, individuals may reduce internal demands by offloading task processing to others – social offloading. Studies have revealed facilitated performance when people believe a jointly acting partner is responsible for task distractors. A proposed mechanism underlying this effect is selective attention. We investigated this using a joint location negative priming (NP) paradigm. In NP tasks, participants initially respond to targets while ignoring distractors (prime phase); at a subsequent probe phase, responses are typically slower to targets in prime distractor locations, reflecting lingering inhibition – the NP effect. Participants (N = 80) completed the task either alone, or believing a partner was responding to distractors. Following a selective attention account, we predicted prime target facilitation followed by increased NP in the joint condition. While results did not reveal facilitation, NP was significantly increased. We propose that these findings demonstrate the way social contexts can shape selective attention to support distributed working.

  • Self-directed gameplay reveals common and divergent patterns in human problem-solving

    A central goal of cognitive science has been to find universal principles that explain the remarkable flexibility of human problem-solving. However, people display a diversity of problem-solving strategies, prior experience, domain-specific knowledge, and preferences. Which aspects of problem solving are shared across individuals, which are individual-specific, and which are adapted to particular domains? Most studies of human problem-solving rely on brief laboratory tasks in one or a few domains and coarse behavioral measures, making it difficult to robustly explain individual cognitive processing. We propose that longitudinal, self-motivated gameplay on multiple tasks---combined with detailed process-tracing---can help fill these gaps. We introduce mitpuzzles.com, a public platform hosting a suite of constraint-based logic puzzles (e.g., Minesweeper, Sudoku, and Nonograms), instrumented to collect detailed behavioral data including mouse-tracking. We present preliminary analyses of large-scale data (N=85,919 games played by 2,258 users) collected from this website to characterize common and variable problem-solving behaviors across both individuals and games. Among other findings, these analyses reveal stable individual differences and effects of learning. First, abstract features of subproblem complexity predict accuracy and speed across puzzle types. Second, most players avoid complex subproblems, even when those subproblems are more globally informative, but the fastest solvers are less avoidant. Third, players search more efficiently with experience. Our results highlight the promise of studying self-motivated participants engaging with hard, naturalistic tasks to understand both the shared structure and individual variability of human problem-solving.

  • Parameter-Driven Consensus-based Filtering Improves Collective Judgment Reliability in Crowdsourced Annotation

    Crowdsourced annotation can be viewed as a form of distributed human judgment in which individual reliability and task difficulty jointly shape collective decisions. Probabilistic aggregation models estimate latent annotator reliability and task difficulty, but are typically used only to weight judgments or infer labels, not to guide selective dataset refinement. We propose a replicated-dataset consensus filtering framework that improves collective judgment reliability by removing unreliable annotators and difficult tasks based on latent reliability-difficulty parameters. Instead of relying on a single dataset estimate, the method constructs multiple replicated datasets by small random label removal, re-estimates latent parameters on each replicated dataset, and removes components that are consistently selected across replicated datasets. Filtering is applied iteratively with entropy-based stopping conditions. The framework is model-agnostic and can be combined with any probabilistic aggregation model that estimates reliability and difficulty parameters. Using a standard reliability-difficulty model instantiation, experiments on multiple crowdsourced datasets show that parameter-driven consensus filtering improves aggregation accuracy and removes incorrect tasks more selectively than random removal. These results support a cognitively grounded, parameter-based approach to improving collective judgment reliability in distributed annotation settings.

  • The Cost of Inline Definitions: Vocabulary Support in English-Language Math Problem Solving

    Many non-native English speakers study mathematics in English, so they must learn both the math and the academic language used to express it. Large language models (LLMs) might help by producing worked solutions while also explaining difficult words and phrases, but it is unclear whether doing both harms mathematical accuracy. We test whether adding inline vocabulary explanations changes solution correctness. Using the CEFR-J vocabulary profile, we identify terms likely to need clarification and evaluate four open LLMs (2.7B–20B parameters) on English math problems in two conditions: standard solving vs. solving with embedded vocabulary support. On MMLU Elementary Mathematics and GSM8K, vocabulary scaffolding consistently reduces accuracy, with drops up to 16.2 percentage points. These findings reveal a trade-off between language assistance and reasoning performance. Educational interfaces may need to balance these goals carefully or separate language support from problem solving for English-medium instruction.

  • How do iconic co-speech gestures contribute to the truth-conditions of assertions: A surprisal-based ERP investigation targeting N400 and late positivity effects

    To investigate how co-speech gestures modulate linguistic understanding in comparison to specific verbs with similar contents, we conducted an EEG experiment to measure semantic surprisal by exploring the amplitude changes in the N400 component for both information types. We used videos of a person uttering (i) specified sentences, whose verb semantically denoted a specific action, and (ii) underspecified sentences with an unspecific verb accompanied by an iconic co-speech gesture that indicated the specific action of the previous condition. The subsequent sentence contained an instrument noun as target, which either matched or mismatched the specific action. We measured ERPs on the target noun for both information types and found an N400 effect for mismatching target nouns as well as a late positivity effect for gesture conditions. Crucially, no interaction between information type and congruency was observed in the N400 time-window, indicating similar semantic surprisal induced by mismatching specific verbs and content-similar gestures. An interaction was revealed in the late time-window, suggesting that unlike for gestures, the mismatch negativity effect was prolonged for specific verbs.

  • Memory, Arousal, and Temporal Binding: A Computational Model

    In the interval estimation task, participants are asked to estimate the time interval between two events. Interval estimates tend to be smaller if the first event is a voluntary action. The cognitive mechanisms that generate this temporal binding effect remain unclear. In this study, we propose a computational model of this phenomenon. We hypothesize that participants perform this task by comparing the availability of experimental events in memory. More specifically, we claim that temporal binding effects may be attributed to differences in emotional arousal at the time of encoding due to the type of event. For instance, voluntary actions may be accompanied by slightly higher arousal, leading to stronger traces of the first event relative to the involuntary case.

  • Efficient compression in developmental trajectories of color naming

    Convergent evidence suggests that the adult lexicon has evolved under pressure to efficiently satisfy a tradeoff between compressibility and informativeness, formally known as the Information Bottleneck (IB) principle. No work to date, however, has tested whether the same principle also impacts how children learn word meanings during development. Here, we address this open question in the domain of color by collecting color naming data from Hebrew speaking children across development (ages 3-12), and comparing it to color naming data from adults. Across ages, we find that children achieve near-optimal IB tradeoffs, with their color naming systems becoming more complex (less compressible) and more adult-like with age. Importantly, this pattern holds even though children's early color naming systems diverge from those of adults, implying that their efficiency is not merely a reflection of the efficiency of the adult system. These findings suggest that children's developmental trajectories of learning word meanings are guided by the same fundamental principle that shapes the adult lexicon.

  • The Shadow of the Past: Amortized Inference and Belief Revision using Chess as a Model System

    How do humans allocate scarce cognitive resources across a stream of related decisions? Russek, Acosta-Kane, van Opheusden, Mattar, and Griffiths (2025) predict reaction time on each chess move from the within-move Benefit of Computation, treating moves as if they were drawn independently. We ask whether those costs also depend on computation carried forward from the previous move. In chess, a player who has calculated a likely opponent reply can often reuse that work; a surprising reply should force revision. On 135,608 quality-filtered Lichess puzzles attempted by 74,609 unique users, the within-puzzle log-RT spike is +0.309 (Cohen's __ = 0.235; paired __ = 86, N = 135,608; BF10 > 10100). Hierarchical regression climbs from __2 = 0.011 (Russek-static) to 0.282; the amortization interaction is large when predictability is operationalized by the human-style Maia-2 (__ = _0.054, __ < 10_27) and not significant under Stockfish best-move predictability. A threeparameter Bayesian Sampling Model with Cache recovers cache strength __ = 0.085 [0.071, 0.099], with 100% of bootstrap draws positive. The H1 spike is roughly twice as large in solved puzzles as in failed ones (__ = 26.2, __ < 10_150); this pattern suggests active engagement rather than passive surprise. The findings replicate on a 2024 Lichess cohort and on FIDEWorld Rapid + Blitz 2024 over-the-board play.

  • Open dynamics of thought and memory with learning in a deformable landscape.

    Classical attractor networks such as Hopfield models treat memories as static points in a fixed energy landscape. While foundational, this framework cannot capture how cognition unfolds through time as agents interact with their environment and recall structured experiences. In this work, I present an open dynamical systems model of memory implemented in an interactive physicsbased environment, where memories are not point attractors but attracting trajectories carved into a deformable potential landscape through experience dependent learning. A state variable evolves under the combined influence of learned landscape gradients and an external driving force, producing time-resolved recall paths rather than instantaneous convergence. I show that this system naturally produces graded recall, intrusion errors, and interference from basin geometry and drive dynamics alone. This framework reframes memory as a dynamical process embedded in perception/action loops and offers a bridge between attractor networks, sequential memory, and embodied cognition.

  • The Bilingual Advantage in Preschoolers: Does It Hold for Typologically Distant Languages in a Hybrid Cultural Context?

    This study focused on the impact of bilingualism on executive function (EF) and theory of mind (ToM) in preschool children. Seventy-two children (age, 4-6 years) were involved: 36 Turkish monolinguals and 36 bilingual children whose first language was typologically distant and second language was Turkish. EF was measured using a Flanker task and ToM was measured using a False Belief task. Contrary to an expected bilingual advantage, monolinguals demonstrated better performance in EF task and were more likely to pass the ToM task. Among the bilinguals, the Arabic-Turkish speakers scored lower on ToM, which might be attributed to socioeconomic differences. The results raise the possibility that distant languages may lead to higher cognitive load and to fewer bilingual advantages. Moreover, the sociocultural context of Turkiye characterized by diverse cultural values and linguistic environments may moderate effects of bilingualism, revealing the need for models of bilingual cognition that are context sensitive.

Papers with Oral Presentation

  • Investigating event orders in LLMs: Insights from English, Norwegian, and Greek temporal connectives

    Event knowledge representations in Large Language Models (LLMs) are explored through the lens of temporal connectives. Temporal connectives like before and after play an important role in sentence comprehension by linking events in the sentence to real-world event order. Their position in a sentence can create different outcomes for event order interpretation and sentence comprehension, and they have therefore been studied extensively in experimental research. Importantly, previous studies have found a human preference for chronological ordering in sentence comprehension. The current study investigates whether LLMs can recognize event order and reflect the human-like processing patterns of temporal connectives across three languages: English, Norwegian and Greek. The aim of the study is to shed light on how sentence structure and event sequence influence LLM predictions, with implications for both cognitive modeling and the knowledge representations learned by LLMs. Results vary by language and show that some models do have event representations and reflect human-like patterns of temporal ordering. The results highlight that language models for some languages (English) are better optimized for cognitive modeling than others (Greek) and underscore the need for more cognitively motivated evaluation benchmarks to assess the models being used in cognitive science research.

  • Entropy in Semantic Memory Navigation in Blind and Sighted Individuals: The Effect of Visual Experience

    Embodied accounts of semantic memory highlight the role of sensorimotor systems in acquiring and storing knowledge. Congenitally blind populations offer a critical test bed for these assumptions, providing an opportunity to assess whether conceptual grounding requires visual experience. In this study, we assessed semantic memory navigation differences between blind and sighted individuals using a property listing task with concrete and abstract concepts. We computed semantic entropy, an embedding-based natural language processing metric that captures the predictability of retrieval. Generalized linear mixed models revealed distinct navigation patterns across groups: while sighted individuals showed higher entropy for abstract than concrete concepts, blind participants did not. Instead, blind individuals exhibited higher entropy for visually salient concrete concepts (e.g., penguin). These results underscore the role of visual experience in the organization and dynamic navigation of semantic memory.

  • A Biologically Plausible Model of Mental Multiplication

    Mental multiplication is an abstract and advanced cognitive skill that requires applying systematic algorithms to structured numerical data, but it remains underexplored in neural models. In this paper, we propose a psychologically plausible model of how anatomical circuits in the human brain may perform mental multiplication of small natural numbers and implement it in biologically plausible spiking neurons. Our model uses similar strategies to those used by humans, specifically repeated addition and rule use; matches various human performance levels, including that of both children and adults; and qualitatively replicates multiple human performance patterns, such as problem-size and outlier effects. Our novel model sheds further light on how structured cognitive processes, like the rule-based algorithms underlying mathematical reasoning, may be implemented in the substrate of spiking neurons in the human brain.

  • Constant curiosity? Children and adults prefer to ask "why" questions about variability over constancy

    Children frequently ask, "Why?" Here, we examined which "why" questions learners ask, hypothesizing that they are more curious about why things vary versus are constant. Informed by the Explanation-Seeking Curiosity Model (Liquin & Lombrozo, 2020), we predicted that explanatory curiosity about variability aligns with children's judgments that (1) they will learn more when asking about variability (i.e., explanations are complex rather than simple), and (2) explanations for variability are more useful. Two pre-registered experiments with children (N = 84 five- to ten-year-olds) and adults (N = 80) supported these hypotheses. Children and adults generated and selected "why" questions about variability more than constancy, reasoned that constancy has simpler explanations, and evaluated variability as more useful to explain in everyday life. Perceived utility, but not simplicity, predicted participants' question preferences. We discuss how explanatory curiosity about variability may be leveraged by educators, but also caution that children and adults may underexamine important constants.

  • Prior Sensory Tuning Orients Error Processing During Sensorimotor Adaptation

    Does heightened proprioceptive precision promote the more aggressive updating predicted by reliability weighting, or the more conservative updating predicted by source estimation? To address this question, we introduced an unperturbed tuning phase prior to prism adaptation, varying visual availability during reaching across three contexts: eyes closed (EC), eyes open with ambient vision masked (EO-), and eyes open with ambient vision (EO+). Proprioceptive uncertainty estimates revealed context-dependent tuning effects, with the strongest sharpening in EC, weaker and less consistent changes in EO-, and a trend toward increased uncertainty in EO+. To examine how these pre-tuned proprioceptive states shaped subsequent updating, we modeled prism adaptation behavior using a hierarchical Bayesian state-space framework. This analysis demonstrated that sharpened proprioception was primarily coupled with more conservative state updating, consistent with a source-estimation account. In addition, the model captured a complementary modulation of performance stability, where improved proprioceptive precision suppressed trial-to-trial fluctuations in movement. Both effects followed an EC > EO+ > EO- ordering, revealing a non-monotonic mapping between proprioceptive tuning magnitude and functional impact. This divergence suggests that the influence of proprioceptive sharpening depends not only on uncertainty reduction, but also on how visual availability organizes the sensory context during action. Collectively, these findings suggest that error interpretation is a state-dependent process shaped by prior sensory context, with sharpened proprioception biasing error evaluation toward more conservative updating.

  • A neural model of short-term and intermediate-term memory using neuro-symbolic representations for semantic and spatial memory tasks

    Memory is a vital part of cognition. Its many parts determine how one learns, plans, and ultimately experiences the world around them. As a bridge between short-term memory and long-term memory, intermediate-term memory allows for semi-stable memories to be rapidly encoded through short-term synaptic plasticity. We present a model of short-term and intermediate-term memory based around Semantic Pointers and realize the model in a spiking neural network. Unlike past work, we characterize short-term memory as a dynamical system and realize intermediate-term memory as implicit association learning. Further, we demonstrate our model's ability to generalize across domains by achieving human-like results on semantic and spatial memory tasks with minimal parameter changes. By examining Cohen's _ values first within and then between tasks, we find an overall small difference (Mean: |_| = 0.12 ± 0.06, Median: ˆ |_| = 0.16)), suggesting strong similarity between model and human data across all three tasks.

  • The Dual Role of Abstracting over the Irrelevant in Symbolic Explanations: Cognitive Effort vs. Understanding

    Explanations are central to human cognition, yet AI systems often produce outputs that are difficult to understand. While symbolic AI offers a transparent foundation for interpretability, raw logical traces often impose a high extraneous cognitive load. We investigate how formal abstractions, specifically removal and clustering, impact human reasoning performance and cognitive effort. Utilizing Answer Set Programming (ASP) as a formal framework, we define a notion of irrelevant details to be abstracted over to obtain simplified explanations. Our cognitive experiments, in which participants classified stimuli across domains with explanations derived from an answer set program, show that clustering details significantly improve participants' understanding, while removal of details significantly reduce cognitive effort, supporting the hypothesis that abstraction enhances human-centered symbolic explanations.

  • Timing is Everything: Temporal Scaffolding of Semantic Surprise in Humor

    Humor is a fundamental cognitive phenomenon in which humans derive pleasure from expectation violations and their resolution, exemplifying the brain's dynamic capacity for predictive processing. Classical humor theories emphasize semantic incongruity as the primary driver of amusement, yet overlook temporal dynamics despite comedians' intuition that "timing is everything." The extent to which temporal structure contributes to humor appreciation and how it interacts with semantic content remains poorly understood. Here, we propose the Dual Prediction Violation (DPV) framework to capture the interplay between content and timing. By analyzing 828 professional Chinese stand-up performances, we show that temporal features substantially outweigh semantic incongruity in predicting audience appreciation. Specifically, we find that peak semantic violations matter more than average incongruity levels, and pauses systematically lengthen before high-surprise punchlines - a strategic coupling that distinguishes successful from unsuccessful performances. These findings reframe humor as temporally scaffolded, where timing and semantic content operate in strategic coordination rather than independently. Our DPV framework bridges humor theory with predictive processing, demonstrating that temporal structure plays a central role in naturalistic humor appreciation with implications for understanding multi-scale prediction integration in linguistic processing.

  • A Representational Analysis of Numeration Systems in Sign Languages

    Although highly relevant to questions on how numeration systems are shaped by format, sign languages have rarely been included in representational analyses. Our study addresses this gap by examining how systems in 63 typologically diverse sign languages represent numbers across the base and power dimensions. The analysis shows that, unlike other formats, sign languages frequently combine multiple representational strategies within their numeration systems. In the base dimension, the reconciling needs for transparency and articulatory economy causes shifts from cumulative to ciphered representations. In the power dimension, format specific constraints guide which representational strategies remain viable, with parsed representations dominating higher powers. These findings highlight the important contribution that sign languages can make to comparative research on numeration systems and underscore the need for a more format sensitive understanding of numerical representation.

  • A Utility-Based Account of the Choice of Constituent Question Polarity

    People choose their questions based on their conversational goals. But how do they decide between questions that would yield equivalent information? This study examines the mechanism by which a questioner chooses between constituent ("wh-") questions of positive vs. negative polarity, e.g., "Who came?" vs. "Who did not come?". Our Rational Speech Acts (RSA) model predicts the following: First, when a partial answer is expected, the question whose polarity aligns with the questioner's goal is preferred over its opposite polarity counterpart. Second, this goal alignment effect decreases when the structure of the questioner's decision problem requires an exhaustive answer. These predictions are confirmed in two experiments: Experiment1 shows that goal alignment emerges in a free completion task, and Experiment 2 shows that it persists in a forced choice task, modulated by decision structure. Our utility-based account thus parsimoniously accounts for a phenomenon unexplained in formal semantic theory to date.

  • Creating vs. Viewing Multimodal Videos: Effects on L2 Vocabulary Retention

    This study examined whether learner-generated multimodal videos enhance L2 word retention beyond multimodal exposure alone and beyond traditional teacher-led instruction. To isolate production from exposure, 164 Chinese university EFL learners were taught 39 unfamiliar words in one of three conditions: video creation (each learner produced two short videos integrating visual, auditory, linguistic, bodily, and affective resources), video viewing (simply watched the same videos), or traditional instruction (teacher-led explanations and examples). Retention was assessed with immediate and delayed post-tests, with previously known items excluded before instruction. Both video-based conditions outperformed traditional instruction, with the largest gains in the video creation condition. The creation advantage extended beyond self-authored items: Creators also outperformed viewers on words encountered only through subsequent in-class viewing, suggesting that creating videos alters attention to and use of multimodal cues during later viewing. Brief learner-generated multimodal production therefore supports deeper encoding and more durable lexical learning than exposure alone.

  • Judgment Under Inconsistency: Tolerance for Inconsistent Political Beliefs

    Holding inconsistent beliefs is widely considered irrational, yet inconsistencies appear pervasive. How do such inconsistencies arise, and when are they remediated? Recent work indicates that inconsistent beliefs can arise and persist when they are not mutually accessible in memory. However, several theories propose that political inconsistencies result from different mechanisms, making them resistant to consistency checking and remediation. Across two preregistered studies (Ns = 1,026; 3,211), we compared the role of accessibility in political versus general knowledge inconsistencies. Participants responded to sets of three interrelated questions designed to elicit inconsistent responses. We found that political beliefs generated inconsistencies at rates similar to trivia beliefs (15-20%). However, upon reviewing their inconsistent responses, participants were significantly less likely to revise political beliefs. Moreover, manipulating accessibility reduced inconsistencies for trivia, but not politics. These results suggest that while accessibility plays an important role in consistency, additional mechanisms govern political inconsistencies.

  • Modality-specific dynamics and context dependent involvement of visuospatial and verbal abilities in analogical mapping

    This work examines the influence of sensory modality constraints on analogical reasoning, focusing on the sensory modality in which the analogy is presented, the influence of verbal abilities, visuospatial abilities and structural factors like verbalisability and visualisability on analogical performance. Through three experiments, we explored how presentation modalities (written, auditory, and mixed) affect the accuracy and latency of analogical reasoning tasks in an A:B::C:D format. Results showed that written presentations yielded the highest accuracy and fastest response times, followed by auditory and mixed conditions. Verbal abilities, assessed with a naming task, strongly predicted performance across all modalities, while visuospatial abilities, measured with a mental rotation task, were particularly relevant in the mixed condition. Higher verbalisability and visualisability difficulty reduced performance across all conditions but to different extents depending on the modality of presentation, reflecting flexible engagement of each of these abilities depending on the task demands. These findings underscore modality-specific dynamics in analogical reasoning and highlight the critical role of cognitive abilities and the structural factors that are verbalisability and visualisability of the relational structure in shaping the ability to solve analogies.

  • Tree of Memory: A Mathematical Model of Hierarchical Episodic Recall

    Understanding free recall requires models that explain not only how many items are retrieved, but also how retrieval is organized—with clustered output, structured transitions, and strong dependence on temporal position. We introduce the Tree of Memory (TOM), in which recall unfolds as a probabilistic search over an explicitly hierarchical episodic representation, coupled with a minimal semantic association structure. The episodic component represents experience across temporal scales encoded in different hierarchical layers. Recall proceeds via probabilistic depth-first traversal of the episodic tree, interleaved with semantic exploration triggered upon item retrieval. The framework admits explicit mathematical characterizations of recall-capacity scaling regimes: depending on the scaling of attention-weighted episodic accessibility, the expected number of recalled items grows sublinearly with list length. In simulations, TOM reproduces canonical free-recall signatures, thereby providing a principled framework to study how hierarchical episodic organization, attention, and semantic associations jointly shape free recall.

  • Capturing Users' Perceptions of Artificial Agents Using Paul A. Weiss's Thought Experiment

    We conducted two investigations to capture how people perceive the artificial agents introduced into Paul A. Weiss's thought experiment, especially, how they perceived the AI system or the robots that are relatively new technology for us with different description (abstracted or concrete). As a result, we confirmed the followings: 1) generative and conversational AIs were perceived concretely as indispensable tools for our daily live with learning function while AI system was perceived abstractly as having knowledge from past, and 2) even the same robots, cleaner robots was concretely perceived as an useful tool while humanoid was abstractedly as ambiguous entity. Here, the presence or absence of appropriate knowledge with the agents may have led to the different ways of perceptions of agents. Thus, the introducing the agents into Paul A. Weiss thought experiment is an effective methodology to comprehend the users' current perception of these agents.

  • Visual Imagery Predicts Semantic Relatedness Judgements in Autistic but Not Neurotypical Adults

    Evidence suggests that the organization of semantic knowledge in autistic populations uses qualitatively different organization principles than neurotypical populations. Across two experiments, autistic and neurotypical adults rated the relatedness of abstract and concrete word pairs and completed the Vividness of Visual Imagery Questionnaire (VVIQ). There were no group differences in relatedness ratings; both groups rated abstract pairs as more related than concrete pairs. However, autistic participants reported significantly lower vividness of visual imagery than neurotypical participants. Critically, a VVIQ x group interaction emerged. Higher imagery scores predicted higher relatedness ratings only in the autistic group across both concept types. These findings demonstrate that equivalent semantic performance can arise through distinct processing pathways. Autistic individuals show perceptually grounded semantic processing that scales with imagery ability, while neurotypical individuals' semantic judgments seem to operate independently of visual imagery. This pattern reveals qualitative differences in how semantic knowledge is organized.

  • Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models

    How large language models (LLMs) align with the neural representation and computation of human language is a central question in cognitive science. Using representational geometry as an explanatory lens, we addressed this by tracking entropy, curvature, and fMRI encoding scores throughout Pythia (70M--1B) training. We identified a geometric modularization where layers self-organize into stable low- and high-complexity clusters. The low-complexity module, characterized by reduced entropy and curvature, consistently better predicted human language network activity. This alignment followed heterogeneous spatiotemporal trajectories: rapid and stable in temporal regions (AntTemp, PostTemp), but delayed and dynamic in frontal areas (IFG, IFGorb). Crucially, reduced curvature remained a robust predictor of model--brain alignment even after controlling for training progress, an effect that strengthened with model scale. These results link training-driven geometric reorganization to temporal-frontal functional specialization, suggesting that representational smoothing facilitates neural-like linguistic processing.

  • Lumping or Splitting: How Context Shapes Motor Sequence Representations

    Motor skills are often assumed to rely on a shared representation that generalizes across superficial changes in framing. Yet identical actions are practiced under distinct contexts that may gate what is retrieved and expressed. We tested whether performance–inconsequential context pushes learning toward generalization ("lumping") or context–specific codes ("splitting"). Simple recurrent networks (SRNs) learned next–step prediction for a fixed eight–item sequence in no–context, single–context, or alternating–context conditions, with weight scale manipulated to bias the learned solution while holding the SRN architecture fixed. Humans trained on the same sequence under matched contexts. All groups improved with practice, but alternating contexts produced a sustained acquisition cost and a marked transition cost when switching between contexts. Under identical training, Low–Scale SRNs showed cross–context generalization, whereas High–Scale SRNs reproduced the context–dependent costs observed in human behavior. White–box analyses revealed separable context–specific representations and greater weighting of context inputs in High–Scale SRNs, consistent with representational splitting.

  • Path Integration and Object-Location Binding Emerge in an Action-Conditioned Predictive Sequence Network

    Adaptive cognition requires structured internal models of objects and their relations. Predictive neural networks are often proposed to learn such world models, but how these are instantiated and how they support prediction remain unclear. We investigate this in a minimal in-silico setting. A recurrent neural network samples tokens sequentially from 2D continuous token scenes and is trained to predict the upcoming token from the current input and a saccade-like displacement. On novel scenes, prediction accuracy improves across the sequence, indicating in-context learning. Decoding analyses reveal path integration and dynamic binding of token identity to position. Interventional analyses show that new bindings can be learned late in sequence and that out-of-distribution bindings can be learned as well. Together, these findings show how structured representations relying on flexible binding emerge to support prediction, offering a mechanistic account of sequential world modeling relevant to cognitive science.

  • Lexical category biases in children's early vocabulary: Decomposing linguistic and cultural predictors across 32 languages and dialects

    Children's early vocabulary is often skewed towards nouns and away from verbs, but these biases are thought to vary cross-linguistically. Prior work attempting to explore this variation has often been limited in language coverage, predictors, or bias estimation methods. Drawing on large-scale cross-linguistic resources (Wordbank, WALS, GramBank, and CHILDES), we provide a systematic decomposition of early lexical category biases into typological, frequency-based, and cultural components. We first derive noun, verb, and adjective bias estimates from early productive vocabularies across 32 languages and dialects. We then test whether cross-linguistic variation in these biases can be predicted by grammatical features and part-of-speech frequency in child-directed speech, as well as broader cultural factors. Several of these predictors reliably explain variation in noun and verb bias, whereas adjective bias is comparatively poorly predicted. These findings highlight multiple classes of factors shaping variation in early lexical biases.

  • Response Variability and Stability in Human Reasoning

    Understanding how humans reason -- and how reasoning responses vary across tasks and individuals -- remains a core challenge for modeling and explanation in cognitive science. We investigate the stability of response patterns within reasoners and whether variation in these patterns can be used to predict learning effects. We introduce a formal, geometry-based method to quantify distances between individual reasoning patterns and their internal variability, grounded in heuristic theories. The proposed framework is tested against experimental data via generalized linear mixed-effects models and clustering, where we find that our proposed variation measure interacts with correctness to predict performance gains. Moreover, we find that reasoning patterns are stable over time within the same reasoner. The method is general enough to be applied to other reasoning domains.

  • Alternative-Space Structure Shapes Negation Resolution in Multimodal Generation

    Negation in natural language is not interpreted as a simple logical complement but is resolved selectively, depending on how linguistic and conceptual structure constrain relevant alternatives. Psycholinguistic research suggests that negation facilitates consideration of an alternative only when a privileged counterstate is available, either through contextual restriction or graded prominence. We use text-to-image generation as a complementary method to probe this selectivity under forced representational commitment: when negated language must be rendered visually, a concrete state must be instantiated. Across two preregistered experiments, participants evaluated how well generated images matched affirmative and negated prompts. Experiment 1 shows that negation is resolved more successfully for predicates encoding conventionalized contrasts than for those with open alternative spaces. Experiment 2 shows that, within open spaces, negation is facilitated when a highly prominent alternative exists. Together, the results support an account in which binary advantages arise from how strongly a counterstate is privileged.

  • The role of temporal uncertainty in procrastination

    Procrastination, where people engage with an immediately rewarding task when they should work on a task with a delayed payoff, is typically framed as self-regulation failure or lack of motivation. We challenge this view by demonstrating that procrastination emerges naturally from the computational intractability of sequential resource allocation under temporal uncertainty. When time horizons are certain, humans approach mathematically optimal strategies for balancing competing tasks. However, when deadlines are uncertain - requiring integration over exponentially growing outcome spaces - participants systematically over-invest in immediate rewards early, then attempt to compensate too late. The findings reframe procrastination as bounded rationality under computational constraints rather than pure irrationality, suggesting that interventions targeting uncertainty reduction or temporal scaffolding may prove more effective than exhortations to increase motivation or willpower.

  • Politically Motivated Reasoning on the Cognitive Reflection Test

    Research on motivated reasoning is currently bogged down by two concomitant debates. The first concerns whether motivated reasoning is the product of fast and intuitive or slow and deliberative thinking processes. The second concerns how prevalent it is given that much of the evidence is compatible with normative counter-explanations, most notably Bayesian reasoning. This paper seeks to advance both debates. Across two experiments, we detect evidence for politically motivated reasoning in a way that cannot be explained by Bayesian rationality. We also find that motivated reasoning is attenuated by cognitive sophistication: Those participants who scored low on a classic version of the Cognitive Reflection Test exhibited motivated reasoning on a politicized version of it, but those participants who scored high did not. Our results suggest that motivated reasoning is a real empirical phenomenon, primarily driven by fast and intuitive thinking.

  • Deconstructing sentence disambiguation by joint latent modeling of reading paradigms: LLM surprisal is not enough

    Using temporarily ambiguous garden-path sentences ("While the team trained the striker wondered ...") as a test case, we present a latent-process mixture model of human reading behavior across four different reading paradigms (eye tracking, uni- and bidirectional self-paced reading, Maze). The model distinguishes between garden-path probability, garden-path cost, and reanalysis cost, and yields more realistic processing cost estimates by taking into account trials with inattentive reading. We show that the model is able to reproduce empirical patterns with regard to rereading behavior, comprehension question responses, and grammaticality judgments. Cross-validation reveals that the mixture model also has better predictive fit to human reading patterns and end-of-trial task data than a mixture-free model based on GPT-2-derived surprisal values. We discuss implications for future work.

  • Popularity Feedback Constrains Innovation in Cultural Markets

    Real-world creative processes ranging from art to science rely on social feedback-loops between selection and creation. Yet, the effects of popularity feedback on collective creativity remain poorly understood. We investigate how popularity ratings influence cultural dynamics in a large-scale online experiment where participants ($N = 1\,008$) iteratively \textit{select} images from evolving markets and \textit{produce} their own modifications. Results show that exposing the popularity of images reduces cultural diversity and slows innovation, delaying aesthetic improvements. Popularity feedback is associated with changes to both selection and creative stages. During selection, popularity information triggers cumulative advantage, with participants preferentially building upon popular images, reducing diversity. During creation, participants make less disruptive changes, and are more likely to expand existing visual patterns. Feedback loops in cultural markets thus not only shape selection, but also, directly or indirectly, the form and direction of cultural innovation.

  • Children's Integration of Verbal and Emotional Cues for Communicative Inferences

    Words and emotional expressions each convey information, but together they can reveal more than either cue alone. Adults integrate these cues to make nuanced inferences, but when does this capacity emerge in development? Children aged 5-10 (N=210) and adults (N=35) watched videos in which a person tasted a friend's food, smiled or frowned, and stated that it "tastes good" or "tastes bad." Participants were asked to infer the food's taste and the speaker's communicative goals. For food taste, adults prioritized negative emotions when cues conflicted; for communicative goals, they used both cues to infer a goal to be honest and relied primarily on speech to infer a goal to make others feel good. Children exhibited substantial developmental progression between ages 5 and 7, although subtle differences from adults persisted at age 10. These results suggest middle childhood is a pivotal period for learning to make sophisticated inferences from multiple cues.

  • Adaptive judgment in the cognitive reflection test: A computational analysis

    The Cognitive Reflection Test (CRT) is widely used in reasoning and decision-making research, yet it lacks a formal cognitive foundation. As a result, debates persist over why CRT scores correlate with mathematical ability, why individuals differ in performance, and which mental processes drive observed response patterns. We address these questions using an ecological perspective combined with computational cognitive modeling. We characterize the learning environment by assembling a large dataset of grade-school verbal math problems and specify a learning mechanism that maps linguistic features of problems to arithmetic operations based on prior experience. Our model reproduces the characteristic intuitive errors elicited by CRT problems and explains them as byproducts of adaptive cognition. It further generates novel predictions about how environmental structure and problem wording influence strategy selection and performance, which we test in two preregistered experiments. Overall, our work provides a principled, quantitative account of CRT performance grounded in adaptive generalization.

  • Quantifying Semantic Priors vs Visual Evidence in Visual Language Models via Psychometric Curve Analysis

    Visual illusions are a classic tool in cognitive science for probing perceptual inference. Recent studies suggest that when visual facts conflict with semantic prior knowledge, vision-language models (VLMs) often make systematic errors in which priors override visual evidence. To quantify this effect, we treat VLMs as artificial participants and introduce IlluQuant, an illusion-diagnostic dataset that operationalizes template similarity and visual evidence strength as continuous, controllable variables. By fitting psychometric functions, we measure semantic prior dominance in VLMs. Results show that stronger semantic cues substantially increase the probability of prior-driven errors, while visual evidence must be amplified several-fold to partially offset this bias. We further propose MdCoT, an intervention that mitigates the effect. IlluQuant provides a psychophysics-style framework for quantitatively characterizing the balance between semantic priors and visual evidence in multimodal models under controlled illusion manipulations. The IlluQuant dataset is available at haitoooo/IlluQuant.

  • Variations in Inner Language and Conceptual Thought Experience

    Inner language is often assumed to include sensorimotor imagery, but is that necessary? How common are conceptual (non-linguistic) thoughts? Does low propensity to think in words indicate lowered ability to imagine language altogether? Here, we sought to characterize the variability in the format of thought by prompting participants (N=228) to (i) think about a given topic or (ii) imagine hearing language or speaking, and then report on their experiences. For thought prompts, a statistically non-trivial percentage of reports included non-sensorimotor inner language (17.5%) and concepts without a verbal or visual form (37.5%). While 74.6% of participants showed high propensity to experience inner language, 4.4% showed no propensity. For language imagery prompts, inner language ability was near-ubiquitous, with only 0.4% (N=1) reporting the absence of linguistic imagery. Overall, we highlight the variability in inner language and conceptual thought experiences, and dissociate the propensity vs. the ability to use language in thought.

  • Learning from Errors in Collaborative Learning: An Experimental Study of Erroneous Worked-Out Examples

    Collaborative learning facilitates cognitive activities such as knowledge elaboration through interaction. However, when extraneous cognitive load is high, the cognitive resources available for engaging in learning processes may be reduced. Erroneous worked-out examples (EWoEs), which present solutions containing misconceptions, can induce cognitive conflict and facilitate knowledge revision. This study examined the effectiveness of EWoEs and their underlying mechanisms in collaborative learning. To enable adaptive example presentation based on learners' knowledge, we employed a knowledge generation model based on the Adaptive Control of Thought–Rational (ACT-R). We used drawing performance as an indicator of knowledge elaboration and the Interactive–Constructive–Active–Passive (ICAP) framework as an indicator of learning processes. EWoEs enhanced drawing performance and the co-occurrence of Constructive and Interactive modes. Moreover, such co-occurrences were associated with improvements in drawing performance. These results suggest that EWoEs facilitate learning processes involving knowledge elaboration triggered by misconceptions.

  • Stereotypes as Bayesian Inferences: Hierarchical Computations Underlying Belief Formation and Maintenance

    Stereotypes are pervasive and rigid. We propose that stereotyping arises because perceivers solve a hierarchical inference problem in which traits are inferred simultaneously at the individual and group levels. Using hierarchical Bayesian models, we derive normative predictions about stereotype formation and updating and test them in three experiments. Study 1 shows that evidence pooling in hierarchical inference enables group-level impressions to form faster than individual-level impressions when perceivers observe limited information of many group members, allowing strong stereotypes without strong individual impressions. Study 2 shows that stereotypes resist updating when stereotype-inconsistent group members behave heterogeneously, making perceivers attribute counterevidence to individual differences rather than updating group stereotypes. Study 3 shows that when stereotype-inconsistent individuals also belong to a second, unfamiliar group, counterevidence is attributed to this second group, leaving the original stereotype intact. Together, these findings provide a computational account of how hierarchical inference produces and sustains rigid group beliefs.

  • Simultaneous Encoding Facilitates Multisensory Category Learning

    Category learning in the real world often relies on information from multiples senses, yet most studies of learning focus on a single modality. It is imperative to study multisensory learning not only to capture processes involved in naturalistic learning, but also because multisensory information may facilitate learning. In the current study, we examined how learning of audiovisual categories is impacted by presenting information simultaneously or separately across modalities. Overall, simultaneous presentation led to more successful performance, even when information was only presented in one modality. Simultaneous presentation also encouraged learners to attend more to the auditory modality. These findings suggest that simultaneous presentation of category information across different modalities can facilitate learning and that this may be related to how learners attend to information during learning. More broadly, these results demonstrate the necessity of investigating multisensory learning to reveal mechanisms supporting learning in more naturalistic multisensory contexts.

  • Coordinative Difficulty Drives Role-Taking in Group Tasks

    Human cooperation involves flexible role-taking. This form of cooperation has been attributed to collective intentionality – the ability to represent what "The Group" believes and desires. However, the factors determining when and why humans spontaneously assume roles remain understudied. We propose that a primary motivation for role-taking is to reduce coordinative complexity by making one's intentions transparent. We test this hypothesis in a naturalistic cooking study (N = 270, 90 trials) in which groups of two, three, and four collaborate, as additional members increase coordinative complexity by expanding possible task distributions and increasing inference misalignment. Our analysis shows that role-taking increases significantly with group size, supporting the hypothesis that roles are used to manage coordinative complexity. These findings provide insight into why humans choose certain cooperative strategies, suggesting that facilitating coordination and solidarity may be prioritized over optimizing productivity.

  • Why More Evidence May Convince You Less: Evidence, Reliability and Bayesian Rationality

    This paper considers the issue why more pieces of supporting evidence may be less persuasive than fewer, perhaps to the point of dissuading one's interlocutor. Challenging the view that such more-is-less effects result from failures to make correct inferences based on the evidence, this paper argues that they may be a rational response based on one's beliefs about a source's reliability. To do this, we present a Bayesian network framework for inference with potentially unreliable sources. While existing Bayesian models forbid more-is-less effects, our more general framework naturally accounts for them.

  • What makes an action sequence enjoyable to watch?

    People often seek out ways to watch others perform complex action sequences (e.g., sports). What makes some sequences more enjoyable to watch than others? We generated 24 video clips of gameplay from a Flappy Bird-style video game. Clips varied in difficulty (how often players succeeded on average) and in moment-to-moment uncertainty (how likely the player was to crash at any given step). Participants (N=864) rated each video on one of three dimensions: how much they enjoyed it, how difficult the level appeared, or how dangerous the player's trajectory appeared. We found that participants preferred videos where the player seemed to be completing more difficult obstacle courses, but dangerousness did not predict enjoyment ratings. These findings show how procedurally generated stimuli can isolate the factors that affect how enjoyable an action sequence is to watch.

  • Knowledge First, Belief Later? Equalizing Processing Demands in Knowledge and Belief Attribution Tasks

    Studies show that children succeed in perceptual access-based knowledge attribution before false belief attribution. Recent theories suggest that this may be due to the primacy of knowledge representations. However, to more accurately contrast knowledge and belief attributions, closely matched test scenarios are needed that are comparable in processing demands. In the current study, we tested 4- to 5-year-old children (n = 50) on a novel paradigm requiring more complex knowledge attributions that involve more comparable processing demands—such as egocentric ignorance and self and other retrospective updates—in relation to false belief attributions. Results showed that children performed well on this more complex knowledge attribution, suggesting that knowledge ascription is robust even under increased cognitive demands. However, we found that complex knowledge and false belief attributions were negatively related, suggesting a surprising pattern in which false belief inferences may be overgeneralized early in development.

  • Regularization more than representational capacity drives heuristic discovery in Bounded Meta-Learned Inference

    Finding computational models that give rise to human-like heuristic decision-making is an open challenge in cognitive science. Meta-learning has been proposed as a promising computational framework for obtaining algorithmic models of cognition by learning from environmental interactions rather than designing normative models. Recent work by Binz et al. (2022) demonstrates that bounded meta-learned inference (BMI) discovers human-like heuristic decision strategies in paired comparison tasks. However, the reasons and mechanisms underlying this discovery remain unclear. Here, we extend research on the BMI framework to better understand its behavior and reverse-engineer its computational principles. We first systematically varied model parameters that influence the network's representational capacity, finding that BMI robustly discovers context-appropriate heuristics across configurations. The model consistently uses single-cue strategies when feature rankings are known and equal weighting strategies when feature directions are known, even when network capacity is strongly reduced compared to the original work. We then reverse-engineer the computational principles underlying this behavior using hierarchical Bayesian modeling. Our analyses reveal that regularization, not network size, drives heuristic discovery. Bayesian models with heuristic priors provide the best fit for regularized network behavior when environmental cues are available. Additionally, we discovered a stronger inductive bias toward equal weighting over single-cue strategies. Thus, we conclude that inductive biases, in the form of priors, have a greater impact on the emergence of heuristics than resource rationality, in the form of networks' representational capacity.

  • BRAID-Color: Extending a probabilistic model of visual word recognition and reading with color processing to account for the Stroop effect

    The Stroop effect is a well-known phenomenon where participants take longer when they name the color of letters that spell the name of a different color. This effect has been extensively studied, with various proposed theories and computational models. However, most models do not integrate with existing frameworks for color processing or reading, limiting their generalizability. We introduce the BRAID-Color probabilistic model, an extension of BRAID, an existing model of reading and reading acquisition. BRAID-Color consists in adding to an orthographic information processing branch, able to recognize words spelled by a sequence of letters, a color processing branch, able to name colored visual stimuli. Information fusion at the lexical level combines both to simulate color naming with colored letters. With default parameters, simulations in the congruent, neutral and incongruent conditions illustrate that the model successfully and robustly accounts for the Stroop effect.

  • Co-Overcooked: Cognitive Constraints on Partner Modeling in Human-AI Team Composition

    Effective collaboration requires accurate mental models of one's partners. Current AI agents, however, may not model humans the way humans model each other under real-time constraints, creating an asymmetry in human-AI teamwork. Prior research has focused on one-human-one-AI settings, leaving open how coordination changes when multiple humans work with multiple AI agents simultaneously. We developed Co-Overcooked, a four-player cooking game where teams of humans and LLM agents must coordinate in real time. In a within-subjects experiment (N=40), participants experienced four team compositions varying in human-to-AI ratio: four humans (H4A0), three humans with one AI (H3A1), two humans with two AIs (H2A2), and one human with three AIs (H1A3). Results revealed nonlinear patterns: H3A1 and H1A3 teams performed worse than pure AI teams, while H2A2 teams showed better performance through spontaneous one-human-one-AI pairing strategies. We identified expectation conflict, where multiple humans hold divergent models of a shared AI partner, as a coordination difficulty specific to hybrid teams, and propose cognitive ownership as a design principle for human-AI team configuration.

  • A rational cascade from teaching explains strategy discovery in development: The case of early addition

    Without being taught, children learning addition discover strategies that are faster and less error-prone than those in their repertoire. Leading explanations of strategy discovery propose metacognitive mechanisms where children iteratively rewrite their existing strategies. Departing from this tradition, we propose that children make discoveries through a combination of reconstructive imitation during teaching, followed by rational inference under basic performance pressures. We instantiate our proposal with a model of Bayesian program induction with a prior for algorithmic simplicity and a likelihood that first encodes a pressure of mimicry (teaching phase), and then a pressure for speed and/or accuracy (individual practice phase). Regardless of whether the pressure is to respond faster or more accurately, the model robustly captures the long-observed developmental trajectory of strategy discovery.

  • Empty restrictor and empty scope effects in questions with quantifiers

    One striking divergence between logical modeling and language processing emerges when evaluating the empty set in the interpretation of quantifiers. It has been shown that empty restrictor scenarios lead to presupposition failure, while empty scope scenarios lead to logically incorrect judgments and a processing delay of downward-entailing quantifiers. In this study, we tested whether empty restrictor scenarios would be evaluated as true or false rather than rejected as odd in a question-answering task. We contrasted two weak downward-entailing quantifiers (fewer than n, at most n-1) with two weak upward-entailing quantifiers (more than m, at least m+1) in two complex question embeddings that made the evaluation possible to various degrees. We found that in these embeddings, empty restrictors are less often rejected as odd as well as that the empty restrictors of downward-entailing quantifiers are often evaluated as false.

  • Reasonable Forgetting: How Memory Shapes Belief Updating Over Time

    Individuals often continue to rely on information even after credible sources have deemed it untrue, a phenomenon called the continued influence effect. This effect is especially consequential in politicised contexts where misinformation and polarisation threaten informed decision making. Existing accounts disagree on whether such belief persistence reflects cognitive failure or motivated reasoning, or whether it can instead arise from rational belief updating under constraints. We investigate this question by introducing delays during which no new information is provided, allowing memory for prior evidence to be less readily available. Across conditions, belief updating tracked the availability of evidence in memory, with beliefs regressing as supporting information decreased in availability and showing no systematic effects of political alignment. Moreover, belief trajectories aligned with predictions from Bayesian Network models parameterised using participants' own judgements. These findings suggest belief updating tracks memory availability and remains consistent with rational, evidence-based principles.

  • Prompted Generalizations Increase Cross-Domain Analogical Retrieval

    Using a cued-recall paradigm, Experiment 1 had two conditions retrieve base situations in response to surface-dissimilar analogs: one in which participants were prompted to re-represent the base situations in general terms, and one in which they reproduced them verbatim. Although the prompted generalizations tended to be either too tied to the original situation or too broad so as to encompass disanalogous cases, they increased retrieval of the base analogs compared to the unprompted condition. Besides replicating the retrieval advantage of prompting one-shot generalizations, Experiment 2 addressed whether a self-administered training could aid participants in producing very general reformulations of base situations of the kind that, in Experiment 1, were associated with greater retrieval. Even though trained participants produced a greater proportion of highly abstract generalizations, the training yielded no retrieval advantage over and above the prompted abstraction condition. We discuss the implications of the present findings for transfer-oriented instruction.

  • What guides utterance choice in argumentative language use?

    Theories of the pragmatic use of language in the Gricean tradition, and related formalizations such as the Rational Speech Act (RSA) modeling framework, have focused on information exchange between cooperative interlocutors. More recent work has emphasized the importance of other, potentially less-cooperative functions of language, such as persuasion, argumentation, and manipulation. In this work, we bridge these two traditions by developing a model of argumentative utterance choice in the RSA framework. We propose several formalizations of the argumentative strength of a signal. We then measure how well they capture behaviour in a simple utterance production task in which participants were assigned an explicit argumentative goal. We show that the most popular formalization of argumentative strength fails to capture some of the participants' utterance choices, and discuss some alternatives. Our models constitute a substantial improvement over a baseline model, and remaining unexplained patterns are discussed.

  • Community Violence, Socioeconomic Status, and Maternal Emotional Distress: Compensatory Parenting Practices in an Urban Latin American Context

    Territorial violence in Montevideo, Uruguay has increased in recent decades and, although concentrated in specific neighborhoods, may affect the broader urban territory and its population. Despite its growing salience, the psychological and family-level consequences of territorial violence in Montevideo remain largely unexplored. This study examines associations between perceived neighborhood violence, maternal negative affect, social support, and hostile parenting in a community sample of 301 mothers of school-aged children. Regression and mediation models controlling for sociodemographic covariates indicated that higher perceived violence was associated with greater maternal negative affect, which mediated associations with hostile parenting. In contrast to findings commonly reported in the Global North, higher perceived violence showed a direct negative association with hostile practices, suggesting a compensatory caregiving response under conditions of insecurity. Moderation analyses further indicated that social support buffered the association between perceived violence and maternal negative affect.

  • Is Child-Directed Language Optimized for Word Learning? A Computational Study of Verb Meaning Acquisition

    Is child-directed language (CDL) optimized to support language learning, and which aspects of linguistic development does it facilitate? We investigate this question using neural language models trained on CDL versus adult-directed language (ADL). We selectively remove syntactic or lexical co-occurrence information from the model training data, and evaluate the impact of these manipulations on verb meaning acquisition. While disrupting syntax impairs learning across all datasets, models trained on CDL and spoken ADL show significantly higher resilience than those trained on written input. Tracking semantic and syntactic performance over training, we observe a "semantic-first" trajectory, with verb meanings emerging prior to robust syntactic proficiency, an asynchrony most pronounced in the spoken domain, especially CDL. These results suggest that the advantage for verb learning previously attributed to CDL may instead reflect broader properties of the spoken register, rather than a uniquely CDL-specific optimization.

  • A Computational Theory of Dignity

    People seem to hold a variety of conflicting intuitions about the concept of dignity. Here, we seek to reverse-engineer the computational basis of these intuitions. We propose a Bayesian model that casts intuitions about dignity as computations about what agents' actions imply about their own and others' social rank. We show through two behavioral experiments that our model captures people's intuitions well, reconciling seemingly conflicting intuitions with a common computational framework.

  • Posting and Reposting: The Reputational Costs of Spreading False Content Online

    Reposting has been identified as a primary conduit for misinformation spread. We hypothesize that reposting is especially conducive to misinformation because it allows the sender to spread false content while incurring lower reputational costs compared to posting. Experiment 1 (N=293) tests this hypothesis by measuring how reposting falsehoods affects reputation. It compares the effects of posting, reposting, and verbally reporting the same false content. Sources that directly posted false content were judged significantly less reliable than sources that reposted or reported the same false content. Experiment 2 (N=256) tests robustness by replicating findings across platforms (Facebook and Twitter) and source types (private individuals rather than institutions). Original posters again suffered greater reputational damage than reposters. These results strongly support the idea that reposting comes with reduced accountability, suggesting that its reduced reputational costs might play an important role in facilitating misinformation spread on social media.

  • Developmental Analysis of Drawing Sequences: How Young Children and Adults Organize Object Parts in Line Drawing

    The drawing process (how we draw), together with its outcome (what we draw), provides valuable insights into the cognitive mechanisms and motor control underlying symbolic representation. To depict developmental changes in the drawing process, this study examined the drawing sequences of Thai-speaking adults and 4- to 5-year-olds during an object copying task. After segmenting the drawing process into discrete strokes and annotating each stroke with its corresponding object part, we extracted patterns distinguishing adults' and children's drawing sequences. Results indicated that when drawing a quadrangle part that forms the overall structure of an object (e.g., a chair's backrest), Thai adults complete it with multiple strokes, while children typically use one stroke. Moreover, when drawing objects with hierarchical structures (e.g., a tournament), most adults group parts of the same level, while young children tend to draw parts in physical proximity sequentially. This study proposes a method for analyzing developmental changes in drawing sequences and highlights key questions for future developmental research on drawing, object recognition, and symbolic representation.

  • Learning to gesture like locals think: Spontaneous spatial gestures adapt to local norms of spatial thinking

    When thinking about space, people preferentially adopt a particular frame of reference, anchoring spatial relations either egocentrically (to their bodies) or allocentrically (to external landmarks). These preferences are often shared within a community, raising the question of how private norms for spatial thinking are made public and thus shareable. We tested whether spontaneous co-speech gesture might be one route by which these norms are expressed. Participants (N = 62) were exposed to an implicit norm to adopt a particular frame of reference when reasoning spatially. They then solved a spatial puzzle and described their solution. When participants conveyed spatial information in their gestures, they adopted either an egocentric or an allocentric frame of reference. Critically, these gestural frames of reference were congruent with the local norm. Gesture thus reflects socially learned conventions for adopting a spatial frame of reference. We discuss implications for the propagation and perpetuation of implicit cultural norms.

  • Effective Explanations Support Planning Under Uncertainty

    Explaining how to get from A to B can be challenging. It requires mentally simulating what the listener will do based on what they are told. To capture this process, we propose a computational model that converts utterances into action plans: a large language model translates an explanation into program-like guidance (a policy prior and value map), and a planning agent executes it under partial observability. We score explanations by the efficiency and reliability of the resulting paths, penalizing replanning. Across four preregistered experiments, we collect a corpus of 1,200 explanations over 24 maps, elicit helpfulness judgments, measure baseline navigation, and test behavior with explanations of differing quality. Higher-scored explanations are judged more helpful and improve navigation: participants with explanations outperform those without, and high-scoring explanations help more than low-scoring ones. Together, these results show procedural explanation as utility-guided communication shaped by how language can be grounded into action under uncertainty.

  • It May or May Not Be Relevant: Testing the Relationship Between Relevance and Informativeness

    The dominant view in cognitive science and formal pragmatics is that relevance is a matter of informativeness: a sentence is relevant if it is informative, and irrelevant otherwise. We provide experimental evidence against this thesis. For polar questions (e.g. Is Mary coming to the party?), we constructed two kinds of uninformative replies: replies that clearly violate relevance (uninformative-off-topic condition; e.g. Mary lives in an old house) and replies that are equally uninformative yet do not intuitively incur a relevance violation, such as uncertainty or disagreement reports (target condition; e.g. Her parents disagree about that). In a within-subject study (__ = 100; 12 scenarios), participants judged both kinds of replies uninformative, but a forced-choice continuation diagnostic revealed a sharp dissociation in perceived relevance: target replies were judged relevant 95%, versus 3% for the uninformative-off-topic condition. These results show that informativeness is not a necessary condition for relevance.

  • Assessing Human-AI Teaming in Single-Pilot Operations: Effects of AI Support on Cognitive Workload and Performance

    As AI-driven concepts like Single-Pilot Operations challenge traditional two-pilot operations, debates among manufacturers, pilot unions, and regulators highlight the need for empirical evidence on how AI uncertainty affects human cognition and performance. This research compares human-human and human-AI cockpit teams, focusing on workload and performance when the pilot monitoring role is assumed by AI. 34 pilots completed simulated takeoff scenarios under three conditions: a human co-pilot, a reliable AI co-pilot, and a faulty AI co-pilot. Results indicate that collaboration with reliable AI can sustain workload levels and team performance comparable to human crews, whereas faulty AI behavior corresponds to increased workload and degraded team performance. Further analyses suggest that performance degradation is linked to undetected AI errors and AI limitations. Although pilots were often able to compensate for AI shortcomings, team collaboration was adversely affected, underscoring the importance of AI reliability and transparency for safe human-AI teaming.

  • Assessing Alternative Models of Counterfactual Reasoning

    When reasoning about events that occur in the world, two key types of judgments are counterfactual inferences—judging how the event might have turned out differently under different conditions—and causal selection—identifying which of multiple potential causes is responsible for an effect of interest. We conducted a novel experimental test that asked subjects to make both types of judgments about a realistic situation involving probabilistic causal relations and both proximal and distal causes. We found that a new model— the Exogenous Sampler (EXS) performed as well as the leading models of these judgment types but also specified the cognitive processes via which such judgments are computed. Yet, that all models failed to predict the full range of subjects' judgments points to the need of a new generation of models accounting for how people reason about events that arise from complex probabilistic causal structures.

  • A Resource-Rational Account of Rule-Based Coordination

    People can coordinate by inferring a partner's intention, or by relying on a mutually recognized rule. We propose that this choice reflects adaptive allocation of cognitive effort. We tested this account in a coordination game where participants could win the most by inferring a signaled ad hoc coordination point, or less by coordinating on a consistent, mutually understood rule. We manipulated the recursive depth required for successful inference and the value of rule-based coordination. Greater demands for recursive depth, and higher rule value, drove increased rule use. Response times suggest rule-based choices arose from different processes: fast defaults when inference was less valuable, and slower retreats to the rule when inference was more valuable but too demanding. Individual-level analyses showed participants shifted toward deeper reasoning when inference was more valuable. These findings suggest that rules and intention inference are substitutable strategies governed by shared resource-rational logic.

  • "There's the ball or the car in this cup", but look, the ball is over there: Preverbal logical inferences scaffold early comprehension of verbal disjunction

    The relation between logical reasoning and language development is debated. We addressed this issue by focusing on disjunctive reasoning (A V B; exclude A; therefore, B) and the acquisition of the connective "or" in 3-year-old children, comparing non-verbal and linguistic contexts. In Experiment 1, children completed first a non-verbal inferential task, and then succeeded in a linguistic inferential task, using "or" to define a set of alternatives (A V B) and visual information to exclude one and infer the other (exclude A; therefore, B). In Experiment 2, children instead completed a non-inferential control task followed by the same linguistic inferential task: although they succeeded in using visual information, they failed to use "or" to identify the relevant set of alternatives. Therefore, children who had previously engaged in non-verbal disjunctive reasoning showed a better comprehension of "or". These findings suggest that preverbal inferential abilities support the early acquisition of linguistic disjunction.

  • Inferring Why Someone Asked 'Why': Foil Inference in Human and LLM Question Interpretation

    Explanations are inherently contrastive: E happened rather than E' because of C rather than C'. However, these contrasts, or "foils", are rarely mentioned explicitly but have to be inferred in context. Here, we investigate how people select the intended foil E' of a why-question. Participants read vignettes and judged, for each foil, their prior expectation (what will happen next), closeness (what is most similar to what happened), and hindsight expectation (what could have happened instead), as well as which foil they thought the question asker had in mind when they asked the why-question. We found that foil selections were best predicted by hindsight expectation judgments. This suggests that people infer the foil by considering what a question asker finds surprising after the outcome occurred. Since correct foil selection is relevant not only in human-human interaction but also increasingly in dialogues with large language models, we investigated their performance on the same task. The coupling between LLMs' explicit expectation judgments and their foil selections is inconsistent.

  • Coherence as a Norm of Evidence Sharing

    Human agents constantly exchange information, but deciding which evidence is worth transmitting is a fundamental cognitive challenge. This paper investigates "coherence" as a heuristic for regulating evidence sharing in groups. Extending previous models of individual coherence-based filtering, we employ computational simulations to analyze agents who selectively share evidence based on its coherence with their current beliefs. We find that coherence-based selective sharing is significantly more effective than individual evidence filtering. It helps preserve priors in overly noisy epistemic environments, while still allowing updating when evidence is noisy but truth-conducive. However, this strategy becomes a liability when misleading evidence forms a coherent alternative, steering group learning away from the truth. These results illuminate the conditions under which coherence-based heuristics improve or distort collective inquiry.

  • Risk-Aware Metacognition: A Dual-Channel Model of Human Second-Order Confidence

    A pivotal advantage of human intelligence lies in metacognitive monitoring. This ability allows individuals to detect uncertainty in their judgments and adapt their behavior accordingly, such as deferring decision-making or seeking help. However, the computational mechanism underpinning metacognitive processing still remains incompletely characterized. We propose a risk-avoidance-based computational theory of metacognition, stating that second-order confidence is formed through an asymmetric strategy: individuals conservatively downweight potential gains while overestimating potential risks inherent in their decision outputs. We formalize this principle into a dual channel conservative confidence (DCCC) framework and instantiate it within a large language model (LLM). Evaluations across high-risk medical and legal judgment tasks demonstrate that the model's confidence signals closely align with human metacognitive profiles, including patterns in the distribution of confidence levels, patterns of error detection, and help-seeking propensities. These findings not only delineate a tractable pathway for implementing machine metacognition but also reveal a core risk-avoidance-oriented computational principle that governs human metacognition processing.

  • Act or Clarify? Modeling Sensitivity to Uncertainty and Cost in Communication

    When deciding how to act under uncertainty, agents may choose to act to reduce uncertainty or they may act despite that uncertainty. In communicative settings, an important way of reducing uncertainty is by asking clarification questions (CQs). We predict that the decision to ask a CQ depends on both contextual uncertainty and the cost of alternative actions, and that these factors interact: uncertainty should matter most when acting incorrectly is costly. We formalize this interaction in a computational model based on expected regret: how much an agent stands to lose by acting now rather than with full information. We test these predictions in two experiments, one examining purely linguistic responses to questions and another extending to choices between clarification and non-linguistic action. Taken together, our results suggest a rational tradeoff: humans tend to seek clarification proportional to the risk of substantial loss when acting under uncertainty.

  • Are Misinformation Effects of Questions with False Implications modulated by At-Issueness?

    We present the results of an experiment investigating the comprehension of questions about previously shown pictures with geometrical objects and resulting intrusion errors on representations in memory. We manipulated whether definite and indefinite descriptions were at-issue or non-at-issue to the task at hand. For this purpose, descriptions were embedded into restrictive vs. non-restrictive appositive relative clauses. The responses to the questions provided evidence for participants' awareness of presupposition violations. Moreover, false presuppositions had an effect on an immediately following memory task. This effect was present no matter whether the presupposition trigger appeared in an at-issue or non-at-issue context. Furthermore, question processing as well as visual memory was unaffected by the wording of the misinformation: in the constructions under investigation, indefinite descriptions patterned with definite descriptions.

  • Towards an empirically-grounded typology of unacceptability judgments

    Theoretical linguistics commonly distinguishes between syntactic ungrammaticality, semantic anomaly, and pragmatic in- felicity, yet there is no established behavioral method for empirically separating these categories. We report two experiments aimed at deriving an empirically grounded typology of unacceptability judgments using multi-dimensional behavioral data. Participants evaluated sentences from a wide range of phenomena against prompts targeting naturalness, truth, interpretability, or contingency. Across both experiments, unsupervised clustering revealed that a stable classification emerges when multiple judgment dimensions are considered and when dimensions which further divide the semantic category are controlled. The resulting typology robustly separates syntactic, semantic, and pragmatic violations, and consistently groups several meaning-driven phenomena, including unlicensed negative polarity items and existential there constructions, with core syntactic violations rather than with semantic or pragmatic ones.

  • "I would not have done that": Outcome knowledge distorts memory for decisions made minutes ago

    Remembering one's past decisions is critical for learning from delayed feedback and maintaining an accurate self-model. Across two gamified experimental paradigms and three experiments (total N = 500), we find that a selective forgetting of unsuccessful guesses operates within minutes, that it evades introspection, and that it affects confidence ratings in memory judgments. Analysis of memory errors reveals that, in reconstructing their past decisions, people integrate what they remember deciding, what they believe they would have decided, and what they would decide now if asked to make the same decision again.

  • Should we commit or should we switch? Aligning goals in sequential social decision-making

    Coherent social coordination requires aligning not only actions in the moment but also goals over time. Yet it remains unclear how joint goals are selected when multiple options unfold sequentially, and how shared information—prior goals maintained in working memory versus newly available goals captured by perception—guides coordination. We examine how the temporal structure of interaction shapes goal selection. Across three experiments, pairs of participants performed Pac-Man–like coordination tasks in which an old goal persisted while a new goal appeared. When coordination required two players to align synchronically in real time, dyads converged on newly appearing goals. When coordination required players to take turns controlling a single agent, dyads persisted with old goals. Critically, manipulating the turn-taking interval revealed a graded transition: longer intervals strengthened bias toward the old goal, whereas shorter intervals eliminated this bias. Together, these findings show that joint coordination flexibly recruits perception- and memory-based mechanisms over time.

  • Graded gender perceptions in personal names predict difficulty in reference resolution

    We conducted two experiments to better understand how comprehenders use probabilistic knowledge of the associations between genders and personal names when inferring the gender of new referents during reading. Personal names are, surprisingly, under-investigated in studies on gender inference compared to role nouns (e.g., 'nurse') despite being strongly connected to gender. In a self-paced reading task involving variably gender-biased names and gendered reflexives, we observe a gender mismatch effect proportional to the incongruence between a name's perceived gender and that of the reflexive. In a novel variant of the Maze task involving forced ambiguity resolution between two gendered reflexives, we find decisions track the perceived gender of names and decisions take longer with more ambiguous names. We interpret the results to support a model in which comprehenders make graded, not categorical, inferences about gender during incremental comprehension, contrary to previous proposals.

  • Individual Differences in Human Teaching of Reinforcement Learning Agents: Evidence from Bayesian Hypothesis Testing

    When humans teach reinforcement learning (RL) agents through real-time interventions, does their teaching strategy adapt to environmental context or reflect systematic individual variation? We present a novel web-based platform for studying human teaching via state interventions (i.e.: physical relocation of learning agents). In a study with 82 participants, we manipulated world dynamics and how agents interpreted interventions using six formally-defined interpretation types. Bayesian analyses provided moderate-to-strong evidence that neither factor affected teaching behavior (BF01 = 5.22 for world type; BF01 = 27.08 for interpretation type). However, participants showed systematic individual variation: high-frequency teachers (n = 39; M = 36 interventions/round) targeted actions with lower Q-values compared to low-frequency teachers (d = -0.57, BF10 = 3.69). While frequent interventions boosted immediate scores, they often impaired long-term policy learning. The reset interpretation, which ignores the intervention, was an exception: it was uniquely robust to suboptimal human guidance. These findings suggest that human teaching in this task is characterized more by persistent individual behavioral profiles than by context-adaptive strategies, with implications for designing resilient human-AI interaction systems.

  • A measure of metacognitive bias you can be confident in

    Research on metacognition critically relies on psychophysical measures that are independent of task performance. Although there exist well-validated measures of metacognitive efficiency (i.e., the degree to which confidence ratings distinguish between correct and incorrect decisions), relatively less attention has been paid to developing corresponding measures of metacognitive bias (i.e., the overall tendency to make responses with high confidence). Most researchers currently use mean confidence as a proxy for metacognitive bias, however this measure is confounded by changes in accuracy, response bias, and metacognitive sensitivity. Here, we introduce a new signal detection theoretic measure of metacognitive bias (meta-delta) and demonstrate that it is independent of a range of potential confounds. Then, by analyzing longitudinal data, we show that this measure is reliable enough to use in psychological research. Overall, we advocate for the use of meta-delta as a principled measure of metacognitive bias.

  • Strategy selection under equal accuracy: Dissociating time, effort, and engagement

    Humans primarily seek to maximize accuracy during task-based decision-making, but prefer to minimize time and effort as well. As time and effort demands are often coupled, the current work examines the relative influences of these secondary minimization goals on decision strategy selection. University students (N=144; ages 18-37) learned two strategies for completing a classification task involving fictional microorganisms: one that demanded moderately effortful integration over learned feature associations, and a minimal-effort alternative that required participants to wait for an answer to be provided. An adaptive experimental design ensured that both strategies yielded equivalent long-run accuracy within-subject, and wait times for the latter were manipulated to be longer or equal to deliberation times in the former. Participants overwhelmingly preferred the more effortful, time-saving strategy. Surprisingly, this preference persisted even when the strategies were matched for both accuracy and time, suggesting the perceived benefit of engagement outweighed the costs of effort.

  • Do Melody and Rhythm Coevolve?

    Music comprises two core structural components, melody and rhythm, that vary widely across cultures. Whether these components coevolve in a coupled way or follow independent trajectories remains unclear. We introduce a novel computational pipeline to extract vocal melodic pitch-interval and percussive inter-onset timing distributions from 27,628 popular songs across 59 countries, enabling large-scale cross-cultural comparison that bypasses traditional music annotations. Musical similarities between countries aligned with geographic and linguistic relationships, validating our approach. Substantial variation emerged in both melodic and rhythmic structures across countries, yet the diversity of the two components was not significantly correlated, challenging assumptions of coupled evolution. Only rhythmic diversity was significantly associated with ethnic and linguistic heterogeneity, while melodic diversity showed no such association. These findings suggest that melody and rhythm constitute partially independent systems shaped by distinct cultural and evolutionary pressures, rather than components of a single monolithic musical style.

  • Clouded Judgments: Children and Adults Infer Unseen Events from Causal Explanations

    In everyday conversation, speakers often give partial explanations, and listeners must draw inferences to fill the gaps in their understanding. When speakers emphasize an unseen cause in their explanation, given knowledge of causal structure, adults can infer if the cause was abnormal or normal. However, the development of this inferential skill is unknown. Across three experiments, we tested adults (N = 62) and 5-8-year-olds (N = 107) on a novel paradigm that assessed normality inferences and explanatory selection. In addition to replicating prior work with adults, we find that 7-8-year-olds also infer the abnormal event, tracking causal influence and pragmatic informativeness. Whereas 5-6-year-olds do not yet make this adult-like inference, they do prefer an explanation citing the abnormal cause over an explanation citing the normal one. These results imply that sensitivity to normality in causal explanations emerges as early as age 5, possibly developing in tandem with skills in counterfactual reasoning.

  • Is Good Always Right? The Flexibility of Space-Valence Body-Specific Associations

    According to the body-specificity hypothesis (Casasanto, 2009) people who interact with the environment differently (e.g., right- and left-handers) associate positive concepts with the space around their dominant hand, and negative concepts with the space around their subordinate hand. Our exploratory study replicated and extended previous findings, establishing a relationship between handedness and mental representation of emotionally valenced abstract concepts. In a task required bimanual categorization of positive and negative nouns participants showed different RT-profiles depending on their handedness. Right-handers categorized positive words faster with their right hand than with the left and negative words faster with their left hand than with the right. Left-handers showed the opposite pattern, whereas ambidextrous showed no differences in the categorizing speed for positive and negative words. Most importantly, the study suggested that these associations were critically influenced by experienced motor fluency. After 12 minutes of hampered interaction with the environment right-handers changed their RT-profiles of associations between space and valence. The results suggest that at least some of our abstract representations are so sensitive to the experienced motor (dys)fluency that it can be manifested within 1 sec of word processing.

  • Rational Teachers Should 'Lie' to Bounded Students

    Educators often build up complex concepts by teaching simplified versions that are not quite accurate, such as Bohr's model of the atom or Newtonian mechanics. "Lying to children'', while ubiquitous in STEM teaching, poses a challenge to existing cognitive models of pedagogy, which assume that teachers select evidence that truthfully represents a target concept. Why would helpful, knowledgeable teachers lie? We present a theoretical framework that addresses this puzzle by reinterpreting optimal pedagogy through the lens of bounded rationality. When learners face cognitive constraints on belief updating, our model predicts that teachers should prioritize examples that will bring the learner closest to the target concept---even if they do not represent the target concept truthfully; by contrast, classic pedagogy models fail to make this prediction. Our work formalizes an insight that educators have long understood: pedagogical "lies'' are not meant to mislead learners, but to meet them where they are.

  • Overoptimistic predictions and well-calibrated expectations: Children say they will achieve unrealistic outcomes but are surprised when they do

    Accurately predicting one's own performance outcomes is a crucial skill for children and adults alike. Prior research, however, has shown that young children are notoriously overoptimistic, making predictions far beyond their actual performance. Curiously, these findings contradict decades of work showing that infants and children hold reasonable expectations about the world, showing surprise---an indication of prediction error (PE)---when events violate their expectations. If children have well-calibrated expectations about their own performance, they might experience PE when they produce unrealistically good outcomes. Using a probability-based game (Experiment 1, N=48) and a memory-based game (Experiment 2, N=64), we show that preschoolers are indeed overoptimistic in their explicit predictions, but express surprise after achieving precisely those unrealistic performance outcomes they predicted. These results demonstrate an early-emerging sensitivity to prediction error about the self, revealing a striking discrepancy between what children say they can do and what they think they can do.

  • Human Action Scaffolds Visual Non-Adjacent Dependency Learning Beyond Surface Form

    Adult learners can acquire non-adjacent dependencies (NADs) in visual sequences, yet learning is far stronger for human ac- tions than for dynamic objects (Lu & Mintz, 2023). We asked what makes human action stimuli special by separating two can- didate contributors: appearance-based structure (a visible body with form and texture) versus biological motion cues (joint trajectories). In Experiment 1, participants viewed full-body avatar actions for 6 minutes and were tested on matched point- light displays (PLDs); they rated NAD-consistent triplets as more familiar than positional foils, showing that knowledge learned from rich action displays transfers to sparse motion- only tests. Experiments 2 and 3 then asked whether motion cues alone can support learning when used during encoding. With PLDs throughout, learning was not detectable after 6 min- utes but emerged with extended exposure, suggesting that the human-action advantage stems from stronger support during encoding, not simply from perceiving the stimuli as actions.

  • Selective Truths: How Speakers Mislead and Listeners Miss It?

    How do speakers craft misleading messages while remaining literally truthful? Can listeners resist such manipulation? We propose that while speakers rely on Theory of Mind to convey strategically underinformative messages, listeners need the same capacity—reasoning about speakers' goals—to avoid being misled. We formalize this by extending the Rational Speech Act framework to speakers with persuasive goals. The model predicts that (1) persuasive speakers rationally exploit underinformativity to frame information favorably, but (2) listeners' ability to correct for such bias depends on whether they actively reason about speaker goals. Experiment 1 confirms that human speakers systematically shift toward underinformativity when the evidence conflicts with their persuasive goals. Experiment 2 reveals an asymmetry: despite knowing that speakers may have persuasive motives, listeners exhibit low sensitivity to underinformativity. These results suggest that recursive social reasoning in strategic communication may be asymmetric—speakers exploit pragmatic inference, but listeners do not fully reciprocate with epistemic vigilance.

  • When Collaboration Beats Ability: Mixed-Ability Teams Can Outperform High-Ability Teams Under Coordination Demands

    Collective intelligence describes the capacity of groups to achieve levels of performance that cannot be explained by the abilities of their individuals. The existence of collective intelligence implies that groups composed of mixed-ability members may outperform groups of high ability. Here, we examine when such advantages emerge in humans and whether emergence is driven by collaboration. We designed a collaborative multiplayer online game and manipulated team composition (mixed-ability vs. high-ability) and coordination demands by exposing teams to two different task environments that varied in how much they encouraged collaboration. We collected data from 280 teams of two human players (70 per condition), totaling 560 participants. We found that mixed-ability teams outperformed high-ability teams in environments designed to encourage collaboration. This performance advantage was due to more collaborative actions and more frequent division of labor. Our results show that collaboration emerges selectively as a function of group composition and coordination demands.

  • Impaired Semantic Guidance in Visual Search Differentiates Prodromal Dementia with Lewy Bodies from Alzheimer's Disease

    Dementia with Lewy bodies (DLB) and Alzheimer's disease (AD) show substantial symptom overlap at the prodromal stage, complicating early differential diagnosis. To identify markers of cognitive dysfunction, we investigated visual search performance in prodromal DLB (LB-MCI), prodromal AD (AD-MCI), and healthy controls across four levels of target visual richness: Gabor patches, abstract shapes, midlevel forms, and real-world objects. While increased stimulus richness improved search performance across all groups, LB-MCI participants exhibited a selective impairment in target present searches involving real-world objects, despite perceptual sensitivity comparable to that of AD-MCI participants. Moreover, unlike healthy older adults, both clinical groups failed to adaptively shift decision thresholds as stimulus complexity increased. These findings suggest that prodromal DLB is associated with difficulty leveraging semantic information to guide visual search, reflecting impaired deployment of semantic target templates rather than early perceptual discrimination deficits.

  • Structural Knowledge in Experiential Risky Choice

    Learning from feedback is a fundamental mechanism through which individuals make decisions under uncertainty. An important open question is how people integrate prior information with repeated feedback to improve their decisions. To address this, we conducted an experiment in which we manipulated whether participants had structural knowledge about the decision environment: half of the participants were told the number of distinct outcomes each option could produce, whereas the other half received no such information and had to infer the outcome structure from feedback alone. Structural knowledge changed behavior under partial feedback, bringing the proportion of risky choice closer to 0.5. We propose a Bayesian model that accounts for participants' initial beliefs about outcome probabilities before receiving any feedback. By capturing choice and confidence patterns across different decision environments, the model demonstrates how structural knowledge shapes learning by altering prior beliefs, thereby linking prior information to experience-based risky choice.

  • Speech Cues Influence the Interpretation of Metaphoric Gestures

    Language and gesture are argued to form an integrated system where each modality contributes to the meaning of utterances. While gestures are known to influence language comprehension, less is known about how language affects the interpretation of gestures. We investigated whether the interpretation of metaphoric gestures is modulated by multimodal cues. Participants viewed hand gestures depicting the figurative interpretation of metaphors such as Men are oaks, where the actor flexed both arms to convey robustness while speaking. Gestures were presented with no additional cues, only lip movements, or audible speech. Participants then rated how well each gesture matched the gestured interpretation (e.g., robust), an alternative interpretation related to the metaphor but not to the gesture (e.g., tall), and an unrelated distractor (e.g., hikers). We found that ratings for alternative interpretations increased when gestures were paired with speech cues, compared to conditions without speech cues, suggesting that gesture comprehension is influenced by verbal content.

  • Understanding Aging-Related Changes in Face Scanning Behavior through Integrating Deep Neural Networks and Hidden Markov Models

    Aging is associated with more nose-focused face scanning pattern and reduced pattern consistency during face recognition, and both effects are correlated with performance reduction. We investigated whether these changes are related to sensorimotor (increased saccade noise) or cognitive (declined memory of visual routines) declines by simulating both mechanisms in a Deep Neural Network + Hidden Markov Model (DNN+HMM) architecture that learns facial representations and oculomotor routines for face recognition simultaneously. Results showed that increasing saccade noise reduced scanning pattern consistency without making the pattern more eyes- or nose-focused, whereas reducing visual routine accuracy to simulate memory decline led to more nose-focus scanning pattern without affecting pattern consistency. Both effects were correlated with reduced recognition accuracy, consistent with human data. Our results thus suggested that changes in face scanning pattern and consistency are dissociable consequences of distinct mechanisms underlying aging-related decline in face recognition ability, with important implications for intervention strategies.

  • Comparison of Resource-Rational Observer Models of Individual and Ensemble Spatial Perception

    To better understand the underlying mechanisms of individual and ensemble perception in naturalistic scenes, we compared three bayesian resource-rational models on experimental data (from 27 healthy adults): the 'Individual Encoding Model' (IEM), a variant of the summation model; the 'Ensemble Encoding Model' (EEM), related to the automatic averaging model; and the 'Task Adapted Encoding Model' (TAEM), a flexible combination of both models that adapts to task demands. In the experiment, participants encoded and reproduced either an individual object position or an ensemble position (group centroid) in a 3D-rendered scene using a computer mouse. In both tasks, we manipulated set size (3, 6, 10 objects) and presentation time (50, 100, 800 ms). The EEM and TAEM generally explained the human behavioral data best. We conclude that, in naturalistic scenes, the choice between individual versus ensemble perception is likely driven by the more compact scene representation of the ensemble model.

  • Characterizing regularity in semantic shift of individuals

    Semantic shift, or the diachronic variation in word meaning, is a topic commonly discussed in historical and cognitive linguistics. Previous work has typically focused on characterizing semantic change in a population, but how word meanings shift over time in individuals is underexplored. We propose a computational framework for analyzing diachronic semantic shift at the individual level by adapting techniques for population-level semantic change. Using speaker-labelled text corpora, our framework utilizes contextualized word embeddings to identify predictable patterns in semantic shift among individual speakers. We discover that frequency predicts the rate of semantic shift over time across speakers, while polysemy does not. Additionally, nouns exhibit higher semantic stability over adjectives and verbs. These patterns partially mirror existing findings on regularity in historical semantic change, suggesting shared tendencies across individual and population levels. Our work bridges computational approaches to historical semantics with the diachronic modeling of personal lexicons.

  • Person distinctions are independent of number: Evidence from artificial language learning experiments

    Pronominal systems tend to maintain the same person distinctions across number categories. This has been explained by most linguistic theories in which pronominal systems are modelled as a combination of independent person and number features (e.g., Harbour, 2016; Harley and Ritter, 2002). At the same time, person and number features are known to interact asymmetrically, which some have taken to indicate that personal pronouns are not straightforwardly the composition of person and number, but rather categories on their own (e.g., Cysouw, 2009). The main source of evidence for feature-based theories is typological data. In this study, we use an artificial language learning approach to provide complementary behavioral evidence to the question of the independence of person and number in personal pronouns. We test whether learners have a preference for person systems that maintain uniform person distinctions across number categories. Our results suggest that, in the absence of explicit evidence, learners infer the same person distinctions across number categories, regardless of whether number is morphologically transparent or not. Our findings suggest that pronominal systems are most effectively learned as combination of distinct person and number features, rather than as unified categories.

  • What Do Children Encode from Graphs?

    The ability to interpret graphs is crucial for educational and professional success, yet many people struggle with graph comprehension. This study investigates how children process graphical information and what graph elements they attend to, encode, and retain. Middle school children completed three tasks. In the Description Task, participants described a graph either while viewing it or from memory. The Recognition Task assessed participants' ability to detect and remember differences in specific components (i.e., axes, scales, and trends) between two graphs. The Slope Judgment Task examined whether judgments of rate of change relied more on pictorial cues (e.g., visual steepness) or semantic cues (e.g., scale values). We found that children consistently extracted and recalled pictorial features across tasks, whereas quantitative semantic features such as axis scales were encoded and retained less accurately. These findings clarify the nature of children's graph encoding and have implications for graph comprehension instruction.

  • Beyond Memory: Goal-Structured Dynamics in Spontaneous Thought

    A complete theory of spontaneous thought must explain not only content but also dynamics: how the stream unfolds over time. Recent memory-centric accounts characterize these dynamics as fundamentally associative. I argue that such accounts cannot accommodate goal-structured dynamics: sequences in which successor selection is governed by instrumental relevance to a goal rather than by associative strength. Drawing on evidence from spontaneous planning, I show that goal-structured dynamics genuinely occur within mind wandering. I propose a dual-driver model in which memory and cognitive control jointly generate the stream: memory drives associative drift, while control imposes goal-structured organization when an automatically activated goal recruits executive resources. On this view, the spontaneous stream is heterogeneous, shifting fluidly between associative and goal-structured phases.

  • The Neural Processing of Cross-categorial Innovations in Language

    Cross-categorial innovations, such as nouns used as verbs, are a productive source of polysemy in language. Although such innovative forms are typically interpreted with ease, little is known about their neural processing. This study investigates the neural correlates of innovative denominal verbs during naturalistic sentence reading. Using concurrent eye-tracking and EEG, we compared innovative denominal verbs with their base nouns and with conventionalized verbs, allowing innovation-related effects to be distinguished from general noun–verb differences. EEG signals were time-locked to the first fixation on the critical word. Innovative denominal verbs elicited a reliable left-anterior negative-going effect in the 300–500 ms time window relative to both control conditions, which extended into the 500–800 ms interval. No evidence was found for an ELAN or P600. We argue that the observed sustained negativity reflects increased demands on semantic-pragmatic integration and context-dependent meaning construction, rather than morphosyntactic anomaly processing.

  • Is Prediction-by-Production Winner-Takes-All? Evidence from Reading Times and Brain Potentials

    Comprehenders routinely generate predictions during language comprehension in real time, yet the mechanisms underlying prediction generation remain debated. The prediction-by-production hypothesis proposes that comprehenders use their own production system to anticipate upcoming linguistic input. While an initial body of work supports production-based accounts of prediction, it remains unclear whether increased engagement of the production system benefits all predictable words – i.e., both high- and low-cloze continuations – or only the most expected ones in a "winner-takes-all" fashion. In the present work, we investigate this question through a German self-paced reading and an event-related potentials (ERP) experiment employing a novel within-participants task paradigm manipulating engagement of the production system during comprehension. Reading times and ERPs show that increased engagement of the production system facilitiates processing for both high- and low-cloze words, suggesting that prediction-by- production supports parallel pre-activation of multiple plausible continuations, with ERP results showing strongest facilitation for highly predictable words.

  • Effects of Generics and Contextual Cues on 3- and 4-Year-Old's Normative Inferences

    Generics (e.g., boys play football) are used to link certain properties to the members of a category. One of the conceptual consequences of generics is licensing normative inferences about how category members should behave. Building on pragmatic approaches, we propose that children are especially likely to form normative inferences when they hear generic statements used in mismatching situations (e.g., saying "boys play football" to a boy who is playing volleyball). We investigated whether 3- and 4-year-olds (N = 157) draw normative inferences from generic statements that either matched or mismatched an observed situation. Partially aligning with our predictions, children prescribed generically and specifically stated properties to new category members more often in mismatching than in matching situations. Exploratory analyses suggested that children were more likely to prescribe generically than specifically stated properties in mismatching and matching situations. Overall, findings suggest that generics and situational mismatch promote normative inferences in children.

  • Mechanisms of belief formation in childhood

    How do children form beliefs from testimony? Previous work with adults suggests a "Spinozan'' account of belief formation on which propositions are accepted as true by default, and rejecting them is effortful. But it is unclear whether this asymmetry between accepting and rejection propositions arises because people learn that others generally say true things or because of the cognitive architecture supporting belief formation. To distinguish between these accounts, we presented 3- to 8-year-olds with statements about a novel animal that were confirmed or denied by a knowledgeable expert. Children with poorer executive function misremembered denied statements more often than confirmed ones. A second experiment confirmed that these results did not reflect a "yes bias'', and that children could select the expert's testimony over the novice's when these were pitted against each other. Together, these results suggest that rejecting statements is challenging because of the cognitive architecture supporting belief formation.

  • Beyond Distance and Density: Manifold Connectivity Shapes Similarity Judgments

    Learning new categories often alters similarity judgments: items from different categories are judged as less similar, and items within the same category as more similar. While category training may drive such changes, categories often have distributional structure without supervision. This raises the question of whether unsupervised exposure alone can reshape similarity judgments, and by what mechanism. We hypothesized that similarity judgments reflect a learned internal geometry of stimulus space. During unsupervised exposure, learners may infer the manifold structure underlying the distribution of examples, such that similarity depends on path-based relations on the manifold rather than inter-item distance or exemplar density alone. To test this, participants experienced stimulus distributions that differed only in whether they formed a connected or disconnected manifold. Despite equivalent physical similarity, participants judged items as less similar when separated by a topological gap. These results demonstrate that unsupervised exposure is sufficient to alter similarity judgments and reflect sensitivity to inferred geometric structure of stimulus space.

  • Take a Guess: Perceived Iconicity, Context, and Transparency of Iconic Signs

    Iconic signs are not necessarily transparent (i.e., easily guessable) because iconicity is a perceived phenomenon that depends on the interpretation of a subjective human perceiver. Nevertheless, previous work has shown that some iconic signs are systematically relatively transparent: people can guess their meanings more easily than those of other signs. In this study, we demonstrate that iconicity constrains response diversity. Specifically, more iconic signs elicit a narrower range of guesses, and these guesses are semantically closer to the intended meaning than those elicited by less iconic signs. Context further strengthens this relationship between iconicity and response diversity, whereas the cross-linguistic frequency of a sign's underlying motivation does not. We argue that these findings warrant a reevaluation of earlier studies claiming that people are generally poor at guessing the meaning of iconic signs, and that iconicity is therefore purely subjective. If signs rated as more iconic produce fewer guesses that are nonetheless semantically closer to the target meaning, then even incorrect responses reveal sensitivity to the iconic motivations underlying sign forms. With even minimal contextual support, iconicity may be less subjective than often assumed.

  • Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models

    Communicating using only a limited vocabulary is a common but challenging cognitive phenomenon, requiring an ideal communicator to plan carefully to optimize for intelligibility while circumventing a constrained lexicon. In this work, we investigate how humans respond to a broad array of questions under variable vocabulary limitations, consisting of only 250 highly frequent words at the most restrictive. We provide theoretically motivated comparisons to greedy and globally optimal sampling algorithms using Sequential Monte Carlo inference with large language models. Humans generally resemble greedy sampling more than globally optimal sampling, though more skilled humans are more likely to backtrack and revise -- a non-greedy behavior. An observed human pattern of leaning on semantically light words in high-constraint settings falls out of both greedy and globally optimal sampling. We discuss the results and their broader implications for resource-rational cognition, psycholinguistics, L2 communication, and language impairments.

  • People Prefer Similar Labels Mapping to Perceptually Similar Referents

    When we encounter a novel label, what, if anything, do we assume about the appearance of its referent? Likewise, when encountering a new creature, what novel label might we generate for it? Here, we examine whether people have a bias to map labels that are similar in form onto perceptually similar concepts (and vice versa). Across three experiments, we show bidirectional evidence of this. In Experiment 1, adults were first introduced to a novel label–referent pairing (e.g., "This is a toku." written above an image of an alien species) and then selected which of two new species best fit a second, novel label. When the new label was similar to the initial one (e.g., "teki"), participants reliably mapped it onto the species that looked more similar to the first one. In Experiment 2, we show that this label similarity effect is robust enough to override pre-existing sound-shape preferences. And in Experiment 3, we demonstrate that the effect also operates in production: after being introduced to a novel creature, participants generated more similar labels when asked to label a perceptually similar, but distinct, creature.

  • Remembering Generalizations: Memory Mechanisms Underlying the Generic Recall Bias

    Humans express generalizations using both generics (e.g., "Dogs bark") and explicit quantifiers (e.g., "All dogs bark"). Quantified statements are systematically misremembered as generics more than vice versa, a phenomenon known as the Generic Recall Bias. Yet, the memory processes underlying this bias remain unclear. We investigate whether this phenomenon arises during memory encoding, retention, or retrieval. Here, we taught adults (N = 1,189) generic or quantified statements about a novel social category, and assessed recall immediately and one week later. Replicating and extending prior findings, we found that "all", "most", and "some" statements were each significantly misremembered as generics. Critically, longitudinal analyses revealed that misrecall patterns were stable across time and consistent within items, supporting an encoding-based account: quantified information is often encoded without its specific proportional information from the outset. Overall, these results clarify the memory mechanisms underlying human generalization.

  • Explaining order effects in counterfactual reasoning

    Human reasoning is often affected by order, such that later judgments depend on earlier ones. For example, order effects have been observed when people answer counterfactual questions. The order in which questions are asked affects whether participants backtrack -- imagining how variables that are upstream of a counterfactual change might have been different. Some counterfactual theories predict backtracking while others do not. We build on this prior work and develop a computational model of counterfactual reasoning that captures order effects. We evaluate different versions of this model on existing empirical results, finding that a model which backtracks and produces systematic order effects best explains human judgments. We conclude by discussing implications of the finding that counterfactual reasoning requires a context-sensitive evidence source for both resource-rational and discourse-coherent explanations of reasoning.

  • Observing Metaphorically Congruent Gestures Enhances Sensitivity to L2 Lexical Tone

    This study investigates how observing gestures varying in congruence with pitch contours influences processing of lexical tones during second language (L2) word learning. Adult L1 English speakers with no tonal language experience completed pre- and post-learning tests in which they heard pairs of Mandarin words with correct (matching) or incorrect (mismatching) tones while N400 event-related potentials were recorded. During the learning phase, participants learned Mandarin words differing minimally in lexical tone while observing congruent, incongruent, or no pitch gestures. Participants completed two post-learning tests of lexical tone association and word meaning discrimination immediately after the learning phase and 24 hours later. Distinct N400 responses emerged for L2 words learned with congruent gestures in the lexical tone discrimination task, indicating that observing congruent pitch gestures facilitates neural sensitivity to lexical tones.

  • What You See Is What You Regret

    Regret is a counterfactual emotion that arises when one judges that an outcome would have been better if one had acted differently. Classic accounts distinguish between commission regret (regretting an action) and omission regret (regretting an inaction). Here, we re-examine this distinction through the lens of attention and counterfactual salience. We find that externally highlighting subsets of options reliably increased omission and commission regret even after action. Increasing the size of the salient set attenuated regret intensity. When participants actively constructed their own consideration sets, the same patterns emerged. These findings suggest that commission--omission asymmetries may depend less on action versus inaction per se, and more on how attention shapes the space of salient counterfactuals, with implications for when regret might support adaptive learning or disrupt it.

  • A resource rational analysis of when (not) to trust your elders

    A major challenge in human social learning is deciding how to integrate "cultural" information inherited from previous generations with contemporary information available from our "peers". While cultural information may be based on more experience, it may also be outdated if the environment has changed. To understand this dilemma from a normative perspective, we developed a cognitive model of a highly simplified version of the problem. Our analysis examined how a learner should balance peer and cultural information in dynamic environments under information costs. We analysed this trade-off as a function of (1) the age of cultural information, (2) the rate of environmental change, (3) the reliability of peer information, and (4) information gathering costs. We also explore how cognitive limitations may bias this process. These results contribute to an increasingly detailed picture of the computational problems that individuals face in the context of cultural evolution.

  • Pragmatic Reasoning in Design

    People can often understand and use novel artifacts after only a few interactions, suggesting that design choices communicate underlying affordances and causal structure. We propose a formal account of this process by framing cooperative, user-centered design as a cooperative game in which the user is the principal and the designer is an assistant. Inspired by prior work on pragmatic communication (e.g. RSA), our model treats a designer's design decisions as communicative signals and predicts user judgments via recursive mentalizing: designers make design decisions to trade off informativeness about the artifact with efficiency, and users infer the true model of the artifact by inverting this cooperative designer model. We evaluate the model in a "design game" where designers place visually identical keys on trays to help a user infer which keys unlock which doors in grid-world layouts. We find that pragmatic designer and user models better match human judgments than non-mentalizing literal baselines.

  • Beliefs that go together change together: A longitudinal study of caregiver beliefs surrounding pediatric COVID vaccination

    Cognitive science offers powerful tools for addressing pressing public health needs. In the current line of work, 1700 US parents reported beliefs about a range of topics related to pediatric COVID vaccination. We deployed Bayesian network structure learning to develop a cognitive model of the relationships among these beliefs and their combined influence on vaccine endorsement. Nine months later, we re-measured these beliefs in a subset of returning participants (n=884) to examine how beliefs changed together over a particularly tumultuous period of time. In a simple comparison of observed co-changes to Time 1 correlations, as well as a comparison of co-changes to model-based simulations, the relationships among beliefs at Time 1 accurately predicted co-change over time (R2 .84-.92 across models). This is consistent with the possibility that correlations across individuals at a single timepoint reflect an underlying belief network that also governs how beliefs change over time.

  • How to make a reasonable person

    Many legal decisions rely on evaluating how a "reasonable person" would have acted, but what defines this standard? Here, we offer an experimental jurisprudence perspective by investigating what subjective features come to mind when reasoning about a reasonable person. In Experiment 1, we examined how people conceptually organize demographic, dispositional, and action features along dimensions of relevance, controllability, and normality. In Experiment 2, we used these dimensions to predict what features shape judgments about reasonableness and outcomes. We found that participants "undid" harmful outcomes by focusing on relevant, abnormal actions; that they constructed a reasonable person from the defendant by preserving normal attributes while changing abnormal, controllable ones; and that they endorsed subjective standards based on features that were relevant and normal, but relatively uncontrollable, consistent with theories of blame. Together, these findings map a template of a reasonable person and provide a descriptive foundation to inform legal theory.

  • How are scientific concepts birthed? Typing rules of concept formation in theoretical physics reasoning

    This work aims to formalize some of the ways scientific concepts are formed in the process of theoretical physics discovery. Since this may at first seem like a task beyond the scope of the exact sciences, we begin by presenting arguments for why scientific concept formation can be formalized. Then, we introduce type theory as a natural framework for this formalization. We formalize what we call "ways of discovering new concepts" including property preservation and concept change, as cognitive typing rules. Next, we apply these cognitive typing rules to a case study of conceptual discovery in the history of physics: Einstein's conceptual path to the relativity of time. Then, we recast what a physicist might informally call "ways of discovering new scientific concepts" as compositional typing rules built from cognitive typing rules—thus formalizing them as scientific discovery mechanisms. Lastly, we computationally model the type-theoretic reconstruction as a program synthesis task.

  • Semantic Access for Reading Comprehension Through Word Restoration

    Access to semantic representations is a central component of reading comprehension. While dual-route models of reading posit the coexistence of phonological and lexical processes, transparent orthographies such as Spanish suggests that reliance on sublexical decoding may sometimes limit deeper semantic processing. This study investigates whether cognitive restoration (i.e., reading words with internal letter transpositions) facilitates semantic access and reading comprehension in expert readers. Sixty university students participated in the study. Experiment 1 examined semantic memory for isolated words using an image recognition task following standard and transposed reading conditions. Experiment 2 assessed text comprehension using narrative texts presented in the same modalities. Results showed significantly better semantic memory performance and improved text comprehension under transposed reading. These findings suggest that cognitive restoration promotes lexical-route engagement and deeper semantic processing, even in skilled readers. Results extend dual-route accounts of reading and highlight word restoration as a mechanism supporting semantic access in higher-level reading comprehension.

  • Falsificationism – the incomplete story about choosing good experiments

    Falsificationism remains a popular framework considered by many scientists as 'the normatively correct model of science', and of human cognition as a kind of 'lay science'. In this paper, we demonstrate some fundamental conceptual shortcomings and gaps in this normative prescription. While some of the issues presented here connect to well-known arguments from the philosophy of science, we connect these issues to the topic of efficient information search. Specifically, we show that a strict falsificationist rejection of the idea of 'confirmation' of scientific hypothesis (more specifically, entertaining and updating beliefs in those hypotheses, based on evidence) makes scientific research and information search overall less efficient, compared to an accuracy-based Bayesian approach.

  • What comes to mind? Ad hoc categories as contextual reweighting of a stable high-dimensional semantic space

    Humans rapidly construct ad hoc categories, e.g., generating "beet" for "vegetables to paint with". Building on feature-based accounts of category representation, we propose an analytic framework in which ad hoc categories are constructed by shifting the base category's representation (e.g., VEGETABLE) along modifier-relevant dimensions (e.g. suitable for painting) in a high-dimensional semantic space, without an explicit hand enumerated feature search. We operationalize this with off-the shelf semantic embeddings, fitting per-category sparse linear models that predict item generation across 20 categories. Across 9 categories tested for compositional transfer, learned modifier directions improved predicted retrieval when physical constraints generalize across domains (e.g., portability for "that could fit in your pocket"), while transfer was weaker for context-dependent modifiers. The learned axes align with human-rated features, with top-weighted features matching modifier semantics. These results frame ad hoc category construction as contextual reweighting of a semantic geometry, bridging feature-based accounts of conceptual combination with distributional representations.

  • A Mechanistic Explanation for the Inverted Face Effect

    The Inverted Face Effect has been the subject of a great deal of research and controversy in vision science. The nature of the effect, whether it involves holistic processing or not, where in the visual system it resides in the cortex, and whether it is a result of expertise and can apply to other phenomena, has been debated for decades. However, no mechanistic explanation for the effect has been put forward. Here we show that a model that includes multiple fixations on the face driven by a simple salience operator, a foveated retina, and the log-polar mapping from the visual field to V1 can explain this effect. We also simulate Yin's 1969 recognition experiment and show good agreement with our model.

Member Abstracts with Poster Presentation

  • Continuous decision-to-action information flow revealed by saccade dynamics

    Drift-diffusion models (DDM) successfully explain choice and reaction times in perceptual decision tasks. However, standard DDMs leave unspecified how accumulated evidence is transformed into motor execution. Servant and colleagues (Servant et al., 2021; Dendauw et al., 2024) have proposed a DDM extension in which the decision variable is continuously transmitted to motor preparation areas (the gated-cascade model). This framework predicts that the neural drive generating the motor response scales with evidence quality. We tested this prediction in a random-dot motion task with saccadic responses (n=27). Because motor neural drive determines the force applied to the eye plant, and force determines acceleration (Newton's second law), we treated saccade acceleration as a proxy for neural drive. Linear mixed-effects models revealed that acceleration build-up rate and peak amplitude scale with evidence quality, consistent with a continuous transmission of decision information to motor preparation areas rather than discrete, serial decision and motor processes.

  • Theory of Mind (ToM) and motivation in Human-AI interaction

    A previous study (Silacci et al. 2026) investigated the effectiveness of Simulated Exercising Peers (SEPs)-AI-powered virtual agents designed to provide social support in promoting physical activity. Participants have a weekly goal of 10.000 steps and their behavior is tracked using a custom mobile app exercise across different conditions: alone with no peer, with another human peer, with a human-like virtual peer or with a cyborg-like virtual peer. Results showed that AI simulated peers offered more consistent encouragement and stronger working alliances, whereas human peers were perceived as more authentic, with both supporting physical activity through complementary motivational pathways. In the present investigation, we assess Theory of Mind (ToM) through post-hoc semi-structured, validated interview (Bosco et al. 2016) administered to participants from the previous experiment. Our preliminary results show that greater performance in ToM correlates with the perceived relatedness experienced during the intervention across the conditions.

  • Pronouns and Mental Simulation While Reading Narratives

    Previous research shows that first-person pronouns prompt readers to adopt a character's perspective, generating mental simulation from an internal viewpoint, whereas brief character descriptions can shift them toward an observer perspective. We hypothesized that when a character appears in a full story and is described as an ingroup member, fostering identification, first-person narration would encourage readers to perceive events from the character's viewpoint. In an online experiment (N = 348), participants read a story about an ingroup or outgroup protagonist, narrated in first- or third-person. They then viewed images of story-related activities shown from internal or external perspectives and selected the picture that best matched the story. As expected, participants most often chose internal-perspective images when the story featured a first-person ingroup protagonist. These findings suggest that when the protagonist resembles the reader, first-person narration more strongly draws readers to experience the story from within the character's perspective.

  • Social Approach–avoidance Conflict in internalizing disorders: a decision-making fMRI study

    Approach–avoidance conflict arises when the same option entails both reward and aversive consequences. It is a key mechanism in internalising disorders. However, human paradigms capturing this construct remain limited. We developed a Social Approach–avoidance Conflict task grounded in fear of negative evaluation (SAAC–FNE) and combined it with fMRI in 135 participants. On each trial, participants decided whether to accept a monetary reward at the cost of increasing exposure to more or less exigent judges in a future public speaking evaluation, or to avoid and receive a minimal reward. Higher rewards increased approach, whereas higher judge ranks increased negative affect and avoidance. Higher depression–social anxiety (D–SA) symptoms predicted stronger avoidance of high-rank judges and more negative emotional responses. Neurally, increasing judge rank during approach recruited dmPFC/dACC activity, amplified with higher D–SA. These findings show that internalising symptoms bias decision-making during conflict and its neural substrates.

  • Neural Correlates of Approach–Avoidance Conflict in Depressive Symptoms within an Interactive Social Competition fMRI Paradigm

    Approach–avoidance conflict is a core feature of anxiety and depression, particularly in competitive social contexts. Using fMRI, this study examined neural correlates of approach–avoidance decisions during an interactive social competition task in 118 participants spanning a continuous range of anxiety and depression symptoms. Participants chose between competing against a rival for a higher reward or avoiding competition for a minimal reward. A principal component analysis yielded a composite anxiety–depression (A–D) score. Neural activity was examined during decision and outcome phases. Higher A–D scores were associated with reduced anterior cingulate cortex (ACC) activation as rival category increased during competitive decisions. Conversely, higher reward magnitude was associated with increased ACC and insula activation in individuals with higher A–D scores. These results indicate that anxiety and depression symptoms are differentially associated with ACC engagement as a function of social difficulty and reward magnitude during competitive decision-making.

  • Self-referential prompts increase negative thoughts in depression and anxiety during think aloud paradigm

    Individuals with depression and anxiety often experience intrusive and persistent thoughts that are difficult to disengage from. Despite their clinical importance, these dysfunctional thoughts patterns remain poorly understood. Here we recruited fifty participants (half with symptoms of depression and/or anxiety, half controls) who completed a Think Aloud task in which they verbalized their thoughts in response to self-referential or distraction prompts while being audio recorded, and rated them in emotional valence. Recordings were transcribed and segmented into thought units for which an overall valence was calculated. Results showed that participants with depression or anxiety exhibited lower overall thought valence than controls. Self-referential prompts further reduced thought valence in the clinical group, while slightly increasing valence in healthy controls. Dynamic analyses revealed that self-referential prompts disrupted the ability of the clinical group to maintain positive thought states. This suggests that self-referential cues exacerbate negative, rigid thought dynamics in depression and anxiety.

  • Cooperation Under Constraint: How Sibling Number Shapes Prosocial Decisions in Vulnerability Contexts

    Drawing on Bronfenbrenner's ecological model (1979), Belsky's parenting process model (1984), and cumulative risk framework (Evans et al., 2013), this study examines whether sibling number predicts generosity in vulnerable children. Children raised with more siblings develop stronger negotiation and sharing skills (Padilla-Walker et al., 2015), promoting prosocial decision-making under scarcity. Fifty school-aged children from low-income migrant families completed a dictator game; generosity was operationalized as net tokens donated to in-group versus out-group recipients. Multiple linear regression with sibling number (1, 2, or 3+ siblings) and family type as a covariate explained 14% of variance (R = .14). Sibling position was significant: children with 3+ siblings were more generous than those with two (b = 1.91, SE = 0.86, p = .031, 95% CI [0.21, 3.61]). Findings suggest family structure scaffolds adaptive prosocial responses under cumulative stress, offering insights for developmental models of prosociality. Replication with larger samples is recommended.

  • Family Configuration and Parenting Style in Vulnerable Migrant Adolescents

    Family Configuration (FC) is a key contextual factor in adolescence, yet its relationship with Parenting Styles (PS) in low-socioeconomic migrant families in Latin America remains underexamined. Building on the Family Stress Model (FSM; Conger et al., 2010), this study tests whether FC differences are associated with variations of PS affecting adolescents' cognitive and socio-emotional regulation through stress-related and social learning pathways. Ninety adolescents (ages 13–17) with migrant parents in economically vulnerable contexts in Argentina completed the Perceived Parenthood Scale (de la Iglesia, 2020). Chi-square analyses revealed a significant association between FC and PS (_(6, 89) = 14.29, p = .002, V = .20), with authoritarian parenting more frequent in single-parent families (marginal trend, p = .063). Stress-related demands in single-parent households may drive more rigid and control-oriented practices parenting, shaping adolescent self-regulation via modeling and reduced autonomy. This extends the FSM by highlighting cognitive pathways to an understudied population.

  • Aesthetic Experience Beyond the Native Tongue: Reading Literary Texts in a Foreign Language

    This study investigates how second-language (L2) readers process and evaluate literary metaphors, focusing on the relationship between comprehension and aesthetic response. Sixty Polish participants (≥B2 English proficiency) read excerpts from "The Golden Bowl" while their eye movements were recorded. Measures included fixation count, dwell time, and revisit count. Additional measures included comprehension tests, ratings of beauty, creativity, and surprise, and retrospective think-aloud protocols. Results revealed positive correlations among aesthetic dimensions and between these and eye-tracking measures (e.g., creativity with dwell time, beauty with revisit count, surprise with fixation count). Comprehension was negatively associated with aesthetic appreciation, replicating pilot findings from an independent text. Less comprehensible metaphors tended to receive higher aesthetic ratings, suggesting that limited comprehensibility does not necessarily impede, and may even enhance, aesthetic engagement. Qualitative analyses of retrospective think-aloud protocols identified sources of aesthetic judgment, including associations, phonoaesthetic qualities, imaginative engagement, complexity, frequency, individual attitudes, and emotional resonance.

  • Thinking Aloud: A Study of Thought Dynamics in Depression and Anxiety

    Self-generated thoughts (SGTs) constitute a central component of human cognition, yet their dynamics remain poorly understood, particularly regarding affective symptomatology. We combined the Think Aloud Paradigm (TAP) with Natural Language Processing to investigate affective and semantic dynamics of SGTs in individuals with and without depression and anxiety symptoms (N = 67). Participants verbalized their thoughts for 10 minutes, rated affective valence of their speech, and transcripts were segmented into thought units. We examined affective transitions and quantified semantic similarity between thoughts using embedding-based measures. Clinical participants reported more negative thoughts, greater persistence of negative states and less stable positive states. Semantic analyses revealed reduced semantic diversity across speech in the clinical group. Moreover, negative valence predicted greater similarity toward the following thought. These findings show that individuals with depression and anxiety exhibit more negative and constrained SGTs, indicating TAP measures capture dynamic features of introspective cognition, relevant for clinical research.

  • Epistemic Judgments Do Not Predict Source Choice

    As access to information expands, learners must decide not only what to learn, but from whom or where. Research in self-regulated learning and epistemic cognition shows that individuals hold differentiated beliefs about source reliability and usefulness, yet it remains unclear whether these evaluations guide source selection in early adulthood. We interviewed 116 college students who imagined preparing an internship presentation on either a novel or partially known concept. Participants described how they would learn it and evaluated their chosen source on perceived knowledge, trustworthiness, expected learning, and enjoyment. Non-AI platforms were most frequently selected (73.28%), followed by AI (21.55%) and human-sources (5.17%). Ordinal models revealed that humans were rated as significantly more trustworthy, enjoyable, and effective for learning, despite being rarely chosen. Their explanations prioritized availability and speed. Together with developmental findings in children, results reveal a dissociation between epistemic evaluation and learning decisions, highlighting efficiency-driven constraints on self-regulated learning.

  • Multimodal Interaction in Tutoring-Like Tasks for Children's Knowledge-Asymmetric Dyads

    Children are capable of effective peer-tutoring (Calero, 2019), yet the interactional mechanisms underlying this effect remain unclear. This study examined how peer-tutoring operates by analysing the multimodal behaviours between a more knowledgeable child (MKC) and less knowledgeable child (LKC) in 7–8-year-olds dyads during a peer-tutoring-like task involving scientific reasoning about the Earth's shape. Sixteen dyads were analysed using coded speech acts, gestures and proxemic behaviours, and data applying generalized linear mixed models. Results showed no overall differences between MKC and LKC in question production. However, differences emerged in verbal interaction patterns: MKC produced significantly more directive statements (commands) than LKC, and only during the first part of the task. Differences in gestures and proxemics were no longer significant once behavioural type and task phase were included in the model. These findings suggest that peer-tutoring effects are driven by context-sensitive interactional configurations rather than stable, role-based differences in knowledge status.

  • The compositional structure of exact number concepts: insights from representational similarity among small numbers

    Novel theories of conceptual development in mathematical cognition propose a compositional origin for exact number concepts, whereby mental operations akin to addition and multiplication enable the acquisition and encoding of increasingly large numbers as algebraic combinations of smaller ones. Accordingly, number representations should not only be tied to numerical magnitude, but also exist in the form of expressions connecting compositionally-related numbers to one another. We provide a preliminary test of this hypothesis by investigating the compositional determinants of representational similarity among small number concepts (0-9) in five behavioural experiments (n = 204 adults). Beyond numerical magnitude, a set of algebraic properties (including factorial relationships, parity status, and primality status) contributed to subjective similarity judgments between numbers, as expected. Nevertheless, these effects were not consistently replicated using semantic priming across a series of numerical judgment tasks, suggesting that the representation of numbers' compositional features is contingent on task-specific demands.

  • Evidence for representation of pretend objects by Kanzi, a language trained bonobo

    Secondary representations enable our minds to depart from the here-and-now and generate imaginary, hypothetical, or alternate possibilities that are decoupled from reality, supporting many of our richest cognitive capacities such as mental-state attribution, simulation of possible futures, and pretense. We present experimental evidence that a nonhuman primate can represent pretend objects. Kanzi, a lexigram-trained bonobo, correctly identified the location of pretend objects (e.g., "juice" poured between empty containers), in response to verbal prompts in scaffolded pretense interactions. Across three experiments, we conceptually replicated this finding and excluded key alternative explanations. Our findings suggest that the capacity to form secondary representations of pretend objects is within the cognitive potential of, at least, an enculturated ape and likely dates back 6 to 9 million years, to our common evolutionary ancestors.

  • Cognitive and Social Emotional Predictors of Optimism in Early Childhood

    The development of optimism is influenced by the interplay of cognitive functions and social experiences, which together shape its complexity and trajectory. Using a direct assessment—the Future Expectation Task—in a longitudinal dataset of Norwegian preschoolers (N = 902), we examined how age, executive function (EF), and social emotional skills (emotion perception, social problem solving, prosocial behavior) predict achievement-specific (expectations of success and achievement) versus general optimism. Structural equation modeling revealed distinct pathways: older age and higher EF predicted lower achievement-specific optimism but greater general optimism, consistent with EF-linked shifts toward learning more from negative than positive outcomes (valenced learning) that calibrate expectations in achievement contexts. Social emotional skills, particularly emotion perception and social competence, were positively associated with general optimism. This study advances cognitive science by linking EF maturation and socio-emotional adaptation to optimism development, offering insight into mechanisms underlying future-oriented thinking.

  • M^3: Meta-cognition and Meta-control with Markov Decision Processes

    Intelligent behaviour requires monitoring and adjusting ongoing cognitive processes to the changing demands and constraints of internal and external environments. This is studied in terms of meta-cognition, cognitive control and meta-reasoning, and, recently, as the expected value of control and the value of computation. However, integrated consideration of core problems such as state representation, generalization, and prospective versus retrospective calculations of the long-run value of meta-computational policies is lacking. We offer a recursive Markov decision process framework called recursive-MDP in which meta-cognition is a form of meta-level perception, and, just as for control itself, meta-control algorithms can be fractionated into costly but adaptive model-based, or fast model-free strategies. Understanding how the brain implements self-adjustment has significance to both psychiatry and artificial intelligence.

  • "You can do it!" Language and Persistence in Brazilian School-Aged Children

    The present study was the first to explore persistence related language in 39 Brazilian school-aged children (Mage = 90 mos, SD = 3.65 mos). Additionally, we tested for a possible association between this type of language (e.g., compliments, encouraging words) and their persistence in a challenging task. Children were asked to narrate a story from a wordless book showing a baby trying to overcome a challenge. During the task, 89.7% of participants used persistence-related words (e.g., "try," "train"), and 38.4% produced explicit messages encouraging persistence (e.g., "You have to train hard to achieve something"). However, no correlation was found between use of persistence-related words (p = n.s.) or use of explicit messages encouraging persistence (p = n.s.) and persistence levels. The present work contributes to advancing a promising line of investigation on the effects of language on important metacognitive and social cognitive developmental processes.

  • Experiencing floods has only marginal effects on individual pro-environmental cognitions and behavior

    Climate change-induced disasters such as floods are becoming more common and may shape citizens' pro-environmental cognitions and behaviors. However, evidence on the effects of first-hand disaster experience remains mixed, partly due to methodological limitations that hinder causal inference. We addressed this issue in a quasi-experimental study comparing individuals affected and unaffected by the 2021 floods in Germany (N = 398) before and after the event. Using propensity score matching and Bayesian mixed models, we estimated the causal effects of flood experience on various pro-environmental responses. Results indicated only marginal effects overall. While affected individuals showed small increases in outcomes such as personal norms to act against climate change and political participation, they did not differ from unaffected individuals on most other measures. We discuss implications for understanding pro-environmental responses and for climate change mitigation strategies.

  • Attention and Comprehension in Music Listening: 18th Century Music for Modern Listeners

    Musical comprehension employs both bottom-up acoustic cues and top-down schematic knowledge. We investigated whether attentional resources are recruited at moments of perceived musical change during listening. Participants listened to three 18th-century musical excerpts while completing segmentation and continuous slider tasks during eye-tracking. For segmentation, listeners indicate when they hear the beginning of a new musical idea; in the slider task, they continuously track perceived tension changes in the music over time. We measure pupil diameter as an index of cognitive load and attentional effort, using the gazeR package. We isolated phasic responses via epoch-locked baseline correction. We predict systematic relationships among segmentation agreement, changes in musical texture or style, and physiological arousal. Preliminary data suggest that listeners' segmentation responses cluster around major musical transitions, indicating shared sensitivity to structural boundaries. Full data collection will test whether style-specific and domain-specific musical content modulates attentional resource allocation beyond acoustic features alone.

  • Deconstructing swearing in the brain: a multifeature analysis of naturalistic speech processing

    Swear words have unique properties and their use can elicit strong early and late neurophysiologic responses when presented in isolation. But swearing is ubiquitous in everyday conversation and often non-offensive; a phenomenon which remains understudied. In this study, 26 Brazilian Portuguese speakers listened to naturalistic podcast conversations containing swear words as stimuli while undergoing magnetoencephalography (MEG) recording. Currently, the collected data are being analyzed using an encoding model to measure the weight of acoustic (spectrogram, speech envelope, acoustic edges), linguistic (words, sentences, grammatical relations), distributional (word surprisal and entropy), and semantic-emotional (concreteness, offensiveness, valence, arousal, tabooness) features during comprehension. This is the first MEG study in BP, a language that is underrepresented in neurolinguistic studies, contributing to the current body of literature on the processing of swear words by use of naturalistic speech as stimuli and validation of a comprehensive analysis of emotional and semantic features.

  • Attraction to Extremes: Cross-Cultural Evidence of Political Acrophily

    Political polarization is intensifying globally, with social media influencing social sorting by amplifying ideological extremes and reinforcing segregation. Recent work suggests that one driver of polarization is the fact that individuals preferentially engage with co-partisans who hold more extreme views than themselves. This study examines whether this tendency generalizes across countries and which perceptions are associated with it. We conducted a preregistered online study across five countries (United States, Chile, Germany, South Africa, and Australia; N = 1,199). Participants evaluated fictional social media profiles, locally adapted and validated, that varied in ideological extremity and reported their likelihood of following each profile. Across countries, participants preferred extreme over moderate ingroup profiles. This tendency was associated with perceiving extreme profiles as more dependable, self-confident, and intelligent. Greater attraction to extreme profiles was also linked to higher outgroup animosity, suggesting biased social inferences and coordination processes reinforced by contextual factors.

  • Replicable Effects of Dynamic Difficulty Adjustment in Child Cognitive Training

    For the past 15 years, we have used the web-based cognitive assessment and training platform Mate Marote with children aged 4 to 8. An initial school-based study in Navarra, Spain, showed that a dynamic difficulty progression model effectively enhanced attention. A second study, conducted again in the same school, evaluated a revised version of this progression designed to overcome its limitations. In the present work, we compare both models. Both approaches adapt to the player's initial level; however, the original model slows progression after reaching a balance point, whereas the revised model continuously adjusts task difficulty. We hypothesized improvements in executive functions from pretest to posttest, with greater gains under the revised model. Results showed attentional improvements in both groups, while only the revised model produced additional gains in logical reasoning, cognitive flexibility, and working memory. These findings were replicated with a local school sample in Tandil, Buenos Aires, Argentina.

  • Cardiac signals increase decision caution in moral judgements of fairness

    Cardiac interoceptive signals have been shown to influence social cognitive processes, yet their impact on moral judgements remains under-explored. We investigated whether cardiac signals modulate moral judgements of fairness by time-locking offers in a modified Dictator's Game to systole or diastole in 58 participants. Generalized Linear Mixed Models revealed that while cardiac phase did not influence binary fairness judgements ("Good"/"Bad")—which were driven solely by offer value—it significantly modulated reaction times. Specifically, we found a heart-phase-by-offer interaction where participants responded faster during diastole for high offers. Bayesian Hierarchical Drift-Diffusion Modelling clarified the mechanism behind this behavioural pattern: systole increased the decision boundary and led to a decrease in drift rate when compared to diastole. The model that included cardiac phase predictors surpassed the "offer value only" model in comparative analysis. Our results indicate that afferent cardiac signals may regulate the judgement by implementing a state of increased caution.

  • Self-Regulated Study Choices Differ Across Learning Domains: Evidence From Example-Based Learning

    Self-regulated learning research assumes that perceived mental effort and judgment of learning guides study choices, but whether the same monitoring cues are diagnostic across domains remains untested. In a within-person design, undergraduates chose between worked examples and practice problems across four trials in each of two domains: a mathematical procedure and analogical reasoning problems. Participants chose practice more often for math from the first trial onward, and most used different study strategies across domains, adopting a more example-oriented approach for analogical reasoning. The monitoring cues associated with these choices also differed, even after controlling for perceived effort and learning: no cues were associated with subsequent math choices, whereas perceived mental effort significantly predicted analogical reasoning choices. These findings extend research on self-regulated example-based learning, which has typically examined single domains, by providing within-person evidence that the diagnosticity of metacognitive cues is domain-dependent.

  • Institutional trust, vaccination beliefs, and the evaluation of health misinformation in Argentina

    Health misinformation is a persistent feature of digital environments and poses a challenge for public health and individual decision-making. Prior research suggests that institutional trust, vaccination attitudes, and health literacy are associated with belief in health misinformation, but most evidence comes from contexts with relatively stable levels of institutional trust. Here, we examine these relationships in Argentina, a socio-political context characterized by high volatility and institutional distrust. In this study, we collected original data from a representative sample to assess credibility judgments of true and false health-related information alongside measures of institutional trust, vaccination attitudes, media literacy, health literacy, and beliefs in complementary medical treatments. Adopting an exploratory, structure-mapping approach, we characterize how these cognitive and attitudinal factors are jointly organized. By identifying distinct individual profiles of belief and information evaluation, we clarify patterns of vulnerability to health misinformation.

  • Causal relations between sleep quality and mental health: a meta-analysis.

    Association between sleep and mental health is well-documented, however, consensus on exact causality remains elusive. This study evaluates sleep in the progression of neuropsychological outcomes through a systematic review and meta-analysis. Following PRISMA protocols and pre-registered via PROSPERO, we screened 1,052 records, synthesizing 29 randomized controlled trials (RCTs; n=2,847) and 3 longitudinal studies (n=44,566). RCT results demonstrate that improving sleep quality reduces symptom severity with small-to-medium effect sizes (g_=_0.46), specifically for depression (g_=_0.48), anxiety (g_=_0.41), and stress (g_=_0.73). Longitudinal data revealed a moderate risk factor (HR=1.76), confirming poor sleep as a precursor to neuropsychological decline. High RCT heterogeneity (I_=81.7%) was addressed through sensitivity and bias analysis (Egger's p>.05). The evidence underscores sleep as a critical mechanism for cognitive and emotional regulation, supporting sleep-targeted interventions as a primary strategy for managing neuropsychological disorders.

  • Do we use all cues? Validation of a continuous measure of compensatory and non-compensatory cue weighting

    People differ in whether they rely mainly on the most valid cue or integrate multiple cues when making probabilistic inferences. Traditional accounts explain the former behavior as the result of prematurely terminated evidence accumulation or heuristics that ignore information, but these views are challenged by evidence for parallel, bidirectional, effortless cue integration as formalized in Parallel Constraint Satisfaction (PCS) models. This project evaluates the PCS sensitivity parameter P, which captures non-compensatory decision-making by quantifying how strongly highly valid cues are overweighted relative to less valid cues. Simulations show that P can be reliably recovered from choices, response times, and confidence ratings. Reanalyses of three datasets (688 participants, 94,792 decisions) reveal higher P values in individuals classified as using non-compensatory heuristics. An empirical validation study further confirms that P reflects experimentally manipulated changes in non-compensatory versus compensatory strategies, supporting its value as a unifying, continuous account of differences in cue integration.

  • Beyond the dyad: Task co-representation of multiple co-actors

    Social cognition research has focused almost exclusively on dyadic interactions or passive judgements of social stimuli. This study moves beyond this 'dyadic default' by investigating 'task co-representation' (mental representation of a partner's task) in groups and assessing the efficiency of this cognitive process when faced with multiple mental states in a single interaction. Groups of three participants took part in a Joint Simon task and an Individual Go No-go control. Compatibility effects were found in the Joint but not the Individual task (p = .034, _p2 = .214), suggesting participants were able to co-represent at least two co-actors, supporting views that this is a cognitively efficient mechanism. Further, there was no influence of spatial proximity between actors within each group, strengthening Co-representation over Referential Coding accounts. Testing participants in groups enables investigation of the limits of social cognition and provides a more ecologically valid model of real-world social interaction.

  • A novel multitalker word detection paradigm to assess neural correlates of conscious auditory perception

    Studies of the neural correlates of conscious perception often rely on threshold-level stimuli, allowing identical inputs to be perceived or not across trials. However, weak stimulation reduces neural response magnitude and statistical power. We developed a novel paradigm manipulating conscious perception using supra-threshold stimuli. Participants heard concurrent speakers producing pseudo-words with jittered onsets. In some trials, a real word was embedded. Conscious access was manipulated by varying word intensity: absent, equal to pseudo-words (target condition), or higher. In an active task, participants reported the word and rated its audibility. In a passive task, participants heard identical stimuli but detected tones while ignoring the words. Behaviourally, participants failed to detect words on 1/3 of active target trials, yielding unconscious trials. Audibility ratings revealed an all-or-none pattern. By contrasting EEG activity across conscious and unconscious trials and between active and passive tasks, this study isolates neural activity specific to conscious perception.

  • Unbiasing the Measurement of Judgment Accuracy: A Hierarchical Extension of the Matching Parameter G of the Lens Model Equation

    When making judgments in probabilistic environments, people are assumed to rely on available cues. Because neither judgments nor criteria are fully predictable from these cues, their correlation underestimates true judgment accuracy. The matching parameter G of the lens model equation was introduced as an attenuation-corrected measure and has been widely used across domains to assess cue-based judgment accuracy. However, building on early critiques, we show through simulation studies that the conventional matching parameter is unreliable and systematically biased, often underestimating matching—especially under noisy judgments or model misspecification. We therefore propose a hierarchical extension of the matching parameter that mitigates these biases through greater model flexibility and regularized estimation. Across simulations and reanalyses of seven empirical datasets, the hierarchical approach yields more valid and reliable estimates of cue matching. These results challenge prevailing conclusions drawn from the conventional approach and motivate the adoption of the hierarchical approach in future research.

  • Zola Bongo: Enabling Offline Tablet-Based Cognitive Testing in Diverse Populations

    Diversifying cognitive science through including understudied populations requires testing participants in their communities. In our work with different populations in Africa, touch screen tablets are an excellent tool for this. However, developing tablet-based cognitive tasks that function without internet typically requires programming skills beyond that available to many researchers.To fill this gap, we created ZolaBongo, a tablet-based application supporting tasks created in the open-source experiment builder, PsychoPy (enabling remote data management via Pavlovia). We tested ZolaBongo with children (age 4+, N = 170) and adults (N = 178) at two sites, in the Democratic Republic of the Congo and in South Africa. We used a set of short cognitive tasks chosen by our community of practice, The African Brain and Cognitive Development Network (AfriBCD). Our hope is that ZolaBongo will make it easier to translate, adapt, or create new measures, all necessary steps to facilitate research with diverse populations and contexts.

  • Grounding User Models in Think-Aloud Protocols

    A central aim of human-computer interaction is to build systems that anticipate users' goals by modeling their underlying cognitive processes. Recent work on general user models infers user states from behavioral observations such as screenshots and action logs. Behavior alone, however, underdetermines the algorithmic-level processes (Marr, 1982) that generate it: distinct reasoning trajectories routinely yield identical actions, leaving inferred user states confounded with the inductive biases of the inferring model. We argue that the think-aloud method (Newell & Simon, 1972; Ericsson & Simon, 1993) offers process-level evidence that behavioral traces cannot, and that recent advances in automated transcription and LLM-based coding make it feasible at ecological scale. Reframing user modeling as the recovery of structured reasoning traces, rather than the prediction of behavior, positions cognitive science to contribute foundational representations to a problem currently dominated by behavioral proxies.

  • Effects of chronotype and circadian alignment on face recognition: preliminary results

    Face recognition is critical in everyday life, particularly in legal contexts such as eyewitness identification. This study examined whether circadian factors influence face recognition performance by testing chronotype effects (morning vs. evening types) under circadian-aligned and misaligned conditions. University students (all sharing similar class schedules) completed two testing sessions, one aligned and one misaligned with their chronotype. Each session included ten target-present lineups assessing accuracy and confidence. Recognition sensitivity was quantified using signal detection theory (d_) from hit and false alarm rates. Objective sleep-wake patterns were assessed using actigraphy. Preliminary behavioral data include N=13 (6 morning, 7 evening). Repeated-measures analyses revealed no significant main effects, likely due to limited sample size. However, exploratory patterns suggest that evening types show poorer performance in their optimal condition, potentially reflecting sleep-related factors and constraints imposed by early academic schedules.

  • Modeling Qualitative Heterogeneity and Covariates with Bayesian Hierarchical Latent-Mixtures

    Traditional approaches in cognitive psychology typically describe effects using group-level means, implicitly assuming that individual differences are quantitative. Such averages can obscure meaningful qualitative heterogeneity, for example when positive, null, and negative effects coexist within a population. Recent cognitive modeling approaches address this using Bayesian hierarchical latent-mixture models to probabilistically classify individuals into latent classes. We extend this framework by including covariates as a predictor of individual class membership probabilities. We demonstrate the validity and robustness of the approach through extensive simulations and illustrate its utility in empirical data: Applying the model to two datasets on the truth effect, we find that individuals higher in Need for Cognition are more likely to belong to the positive (vs. null) truth-effect class, even where traditional analyses fail to detect a relationship. This framework enables researchers to link person-level or contextual variables to latent classes of cognitive effects.

  • Using a dimensional measure of bilingualism to investigate cognitive effects in Montréal adolescents with typical or atypical development

    Bilingualism influences executive functioning (EF), though findings remain mixed partly due to variability in how bilingualism and EF are operationalized. Bialystok (2024) proposed that bilingualism enhances attention through adaptation, improving global attentional efficiency. We tested this proposal using the Wisconsin Card Sorting Test (WCST), examining categories completed, perseverative errors, and rule application reaction times (RTs) as indices of executive functioning and attentional control. Participants from Montréal were neurotypical (n = 37) or neurodivergent (n = 27) adolescents (14–19 years old). Diversity of language use across contexts, derived from Anderson et al. (2018) and Gullifer and Titone (2020), was used as a dimensional measure of bilingualism. Correlation analyses indicated no relationship between bilingualism and WCST performance in neurotypicals, but greater diversity of language use was associated with poorer WCST performance in the neurodivergent group. Regression analyses showed marginal group _ bilingualism interactions for categories completed and RTs.

  • A Representational Format for Efficient Concept Learning and Relational Inference

    A cornerstone of human intelligence is the ability to organize a vast number of entities into abstract concepts and to learn the complex, hierarchical relationships among them. Crucially, not all relations between concepts are directly experienced; many are inferred by integrating knowledge acquired about each concept independently. This capacity supports flexible integration and generalization of conceptual knowledge, yet its underlying computational principle remains unclear. Building on prior work on concept learning, we propose a representational format for concepts that supports efficient concept acquisition and relational inference. The model reproduces key human-like conceptual reasoning biases and, compared with artificial neural networks trained on the same concept structures, yields similar representational geometries yet with greater learning efficiency. These results provide initial evidence for an effective format of concept representation in high-dimensional space and how their relations are used and inferred.

  • Inaccurate metaperceptions about the self-reported attraction to ingroup political extremes

    Political polarization has become a persistent problem in many societies, reflected not only in divergent ideas but in the ways we talk about each other. Studies show that we tend to overestimate how much we dislike opposing political groups, generating cognitive distortions that intensify conflicts. This led us to ask whether people also overestimate the appeal of extremes within their own group. We conducted two studies with N=140 and N=100 Argentinian adults. Participants placed themselves on a 7-point ideology scale and rated their sympathy toward each position from three perspectives: personally, as a typical ingroup member, and as an outgroup member. Across both studies, participants systematically overestimated the attraction of opposing group members to ideological extremes, while also overestimating their own group's preference for those extremes. These findings suggest that misperceptions of preferences intensify political divides and can inform interventions to promote constructive dialogue in democratic societies.

  • Context Effects on the Detection of Manipulated Images in Social Media

    The widespread use of image manipulation in social media raises concerns about how individuals perceive and evaluate beauty-related visual information. This study examines how image manipulation and presentation context influence perceptual processing and accuracy judgments. Participants viewed authentic and beauty-enhanced images of male and female targets presented within an Instagram-like interface or a photo gallery context, and judged whether images had been manipulated. Repeated-measures ANOVAs and post hoc analysis revealed that reaction times were faster for authentic than for manipulated images (t(2,62)=18.12, p=0.001). Accuracy results showed a significant interaction between manipulation type and presentation context: authentic images were detected more accurately in the gallery context (t(2,62)=-2.902, p=0.026), whereas manipulated images were detected more accurately when presented as Instagram posts (t(2,62)=-3.740, p=0.002). Confidence judgments revealed a main effect (F(1,62)=12.35, p=0.001,n2p=0.018) revealing higher confidence for manipulated than authentic images. Together, these findings suggest that media context modulates perceptual and evaluative processes.

  • When Does Emotion Interfere? Task Demands and Language Experience in Cognitive Control

    This study examined emotional interference in cognitive control using a multi-task approach designed to vary emotional, linguistic, and control demands. Heritage bilingual and monolingual young adults completed six cognitive tasks spanning lexical decision, Stroop, and working memory paradigms. Rather than a uniform pattern, emotional valence effects and cognitive costs were highly task dependent. Bilinguals often had similar or higher efficiency than monolinguals in tasks with higher control demands (e.g., nonword rejection, high-load working memory), while emotional interference showed limited group differences, likely reflecting bilinguals' English-dominant testing context. Critically, the multi-task design revealed substantial within-bilingual variability: performance was modulated by English and Spanish proficiency and age of second language acquisition in task-specific ways. By comparing emotional interference across multiple task contexts, these findings demonstrate that emotion–cognition interactions depend on task structure and individual experience, highlighting the value of multi-task designs for understanding cognitive control.

  • Independent and joint contributions of gesture and speech to linguistic prediction

    Listeners can use gestures to predict upcoming referents when the semantic information in speech is partially constraining referents in a visual context. However, less is known about the predictive use of gestures when semantically constraining information is in gesture alone or in both speech and gesture. In a visual-world eye-tracking paradigm, 60 Mandarin Chinese speakers viewed four objects and one video. The speaker produced sentences ("I today myself alone VERB a (whole) afternoon NOUN"), paired with an iconic gesture that began 700ms before the verb, the same sentence without gesture, or the same gesture withe the verb replaced by a speech disfluency. In all conditions, participants had more looks to the target object than non-target objects before they heard the noun, but this difference was greater in the gesture plus speech condition. These results demonstrate independent and combined contributions of gesture and speech for predictions grounded in the visual-referential context.

  • Neural Connectivity Signatures of Adult Illiteracy on 7T Resting-State fMRI

    A significant portion of the Brazilian population (approximately 9.3 million people) remains illiterate despite national educational efforts. Although adult illiteracy is a worldwide phenomenon, the neural connectivity associated with the absence of literacy acquisition remains poorly understood. We investigated 20 illiterate adults using 7-Tesla resting-state fMRI (54 ± 8 years; 11.6 ± 8.6 words per minute). Functional images underwent motion correction through realignment, scrubbing of high-motion volumes, and regression of motion parameters during denoising. Seed-to-voxel analyses controlling for age and reading performance revealed decreased connectivity between the visual word form area and a left frontal cluster including the middle frontal gyrus, superior frontal gyrus, and frontal pole (cluster-corrected p-FDR < 0.0001; MNI _30, 34, 42). ROI-to-ROI analyses also showed reduced connectivity within visual and language networks. These findings suggest that literacy acquisition reshapes large-scale functional brain organization through experience-dependent plasticity in adulthood.

  • Age Shapes Silent Gesture in European Spanish

    Silent gesture provides a unique window into how humans map meaning onto the body without speech. While prior
    work has identified cross-linguistic regularities in iconic strategies (Ortega & Özyürek, 2019a, b), the role of age in
    shaping these mappings remains unclear. Here, we test whether age predicts the selection of iconic strategies (acting,
    drawing, representing, personification) in European Spanish. Thirty monolingual participants (n = 10 per group) from
    three age groups (3-5; 19-22; 51-58) produced gestures for 42 stimuli across five semantic categories. Coding reliability
    was assessed on a subset of the data (≈91% agreement). Results show robust age-related differences in both gesture
    complexity and strategy choice (ps < .001). Multinomial regression further indicates that categorical age group and
    semantic category jointly predict gesture technique (R² = .32, p < .001). These findings suggest developmental and/or
    cohort-related shifts in how meaning is iconically represented across the lifespan

  • Language-modulated Internal and External Event Segmentation

    Languages differ in encoding object-state changes, influencing how speakers segment dynamic events. English and Mandarin diverge lexically (quantized nouns and resultative constructions) and phrasally (telicity representation). This study examined whether these differences modulate internal segmentation (detecting interruptions within an event) and external segmentation (marking event boundaries) of cutting (more bounded) and blending (less bounded) events. In Experiment 1, English speakers highlighted initial states via quantized nouns, while Mandarin speakers emphasized end states using resultative constructions. Phrasally, English—but not Mandarin—showed a telic/atelic split between cutting and blending events. Experiment 2 revealed that English speakers detected early-stage interruptions in blending events more accurately, while Mandarin speakers excelled at later stages, mirroring lexical biases. In Experiment 3, English speakers marked cutting event boundaries faster, aligning with their telic phrasal framing. Converging evidence suggests language systematically modulates event segmentation at multiple levels. Priming paradigms are discussed for establishing causality.

  • Dynamic and Functional Brain Mechanisms During Naturalistic Spatial Learning

    We hypothesize that complex spatial learning during college-level instruction recruits evolutionarily conserved parietal–premotor mechanisms, with engagement scaling to instructional demands. We used fMRI and an authentic spatial lesson coded by rotation complexity, cognitive load, and linguistic density to dissociate brain systems underlying spatial instruction. Activation in the intraparietal sulcus and premotor regions tracked rotation complexity; frontal–cingulate regions tracked processing demands, and language regions selectively responded to linguistic density. Individuals with stronger spatial skills showed greater recruitment of the parietal–premotor circuit as spatial complexity increased, demonstrating that the neural system representing spatial, geometric, and quantitative content reflects individual differences during learning. These results reveal that the brain dynamically engages primitive mechanisms of spatial cognition during naturalistic instruction: parietal and premotor regions computing spatial and numerical relations remain malleable for acquiring high-level spatial cognition. This reveals cognitive functions linked to spatial learning and has implications for translating cognitive science into educational insights.

  • Smell-Enhanced Recall Memory in Younger and Older Adults: Evidence from a VR Paradigm

    The sense of smell is known to enhance autobiographical memory ("Proust effect"). However, less is known about the effect of visual-olfactory integration on younger and older adults' recall and recognition memory. Using a virtual reality paradigm, we explored the effect of object-associated smells on memory. Participants first searched for objects in a virtual home in either unimodal (visual-only) or bimodal (visual+olfactory) conditions. Next, participants completed a verbal recall task for target objects, followed by a forced-choice scene recognition task. The addition of olfactory cues improved object recall for both age groups similarly, but did not affect scene recognition. A follow-up study explored whether cognitive demands (single vs. dual task) could modulate the effect of olfactory cues on memory. Preliminary results replicated the previous recall benefits across age groups, but no effect of cognitive demand was observed. These outcomes have implications for interventions focused on improving memory across the adult lifespan

  • Bare Plurals in Context: Evidence from Adult Interpretation

    nglish bare plurals are often interpreted as conveying a "more than one" meaning, despite their lower-bound semantics of "one or more" (Tieu et al., 2020). This multiplicity inference has been argued to arise via scalar implicature. We propose and empirically test the hypothesis that multiplicity inferences depend on whether quantity is relevant to the Question Under Discussion (QUD). In Experiment 1, English-speaking adults showed no reliable preference between bare plural descriptions of single- and multiple-object scenes in a context that did not highlight quantity. Experiment 2 manipulated the QUD by introducing comparison sets that highlighted either quantity or type differences. Adults showed stronger rejection of bare plurals for single-object scenes when quantity was contextually relevant, while judgments in control trials were unaffected. These findings suggest that multiplicity inferences are context-dependent pragmatic enrichments that arise when quantity-based alternatives are relevant to the QUD.

  • Not a Linear Path: Three Neural Signatures of Native versus Non-Native Instructional Learning

    Learning complex content in a non-native language is widely perceived as effortful, yet neurocognitive evidence underlying this process remains poorly understood. We conducted the first fMRI study using a naturalistic instructional video to examine real-time neural dynamics in native English speakers (N=16), high-proficiency L2 learners (N=17), and medium-proficiency L2 learners (N=19). Using intersubject correlation (ISC) and neural encoding of linguistic embeddings (i.e., GPT2 embeddings), we quantified how strongly the participants' language network (LAN) and multiple-demand network (MD; cognitive control) were engaged by the stimuli. Results reveal three distinct profiles: Native speakers showed strong LAN engagement (high ISC/encoding) with moderate MD engagement. High-proficiency learners exhibited native-like LAN engagement plus effective LAN–MD integration. Medium-proficiency learners showed heterogeneous LAN and MD engagement and ineffective coupling characterized by excessive MD involvement in language processing. These findings provide insights into English-medium instructional challenges, identifying neurocognitive signatures distinguishing successful compensation mechanisms from ineffective processing. Our study provides insights into non-native instructional learning design.

  • Analysis of Mismatch Negativity as an index of the evaluation of sociolinguistic meaning

    How quickly do listeners map phonetic variation onto social meaning? We report an electroencephalography (EEG) / event-related potential (ERP) study with Rio de Janeiro Brazilian Portuguese listeners, examining whether sociolinguistic familiarity or bias modulate mismatch negativity (MMN). In an oddball design, we contrasted (i) palatalization of /t/ before /i/ in tio (meaning uncle), focusing on a non-local stop realization that is effectively absent in the Rio de Janeiro speech community, but typical of the pronunciation in northeastern cities, with (ii) pre-stressed nasalization in canal (meaning channel), a variable whose variants coexist locally (control). Three negative peaks emerged: Peak 1 (canonical MMN) was not sensitive to the variants' from Rio or not from Rio; Peak 2 appeared for both variables, indexing detection of within-category variation; Peak 3 was specific to palatalization, consistent with a later-stage evaluation that the stop variant is socially "non-Rio".

  • The effect of aesthetic appeal on the cultural evolution of language

    Language evolves through cultural transmission, with certain linguistic patterns being passed on over generations while others disappear. Various factors have been proposed to explain which patterns persist, but one factor receiving little empirical attention is aesthetic appeal, i.e. how beautiful language users perceive particular forms to be. We hypothesized that appealing words would stabilize in the lexicon, whereas unappealing words would disappear. We designed an iterated learning experiment, in which participants learned novel names for novel shapes. Each shape was paired with two labels, one appealing and one unappealing. After learning all labels, participants labeled each shape twice. Their responses were passed on to the next participant in the chain, who learned only the labels produced by their predecessor. We found that, indeed, participants preferentially used the appealing labels, and that over successive generations, appealing names dominated the language. This suggests that aesthetic preferences constitute cognitive biases in language evolution.

  • Variety or Specificity? Comparing Emotional Diversity and Granularity in Daily Language

    Individuals vary in describing emotions using language, utilizing different vocabulary (emodiversity) and levels of context-specificity (granularity). Although both are linked to well-being, they are rarely compared using natural language from daily life. We addressed this gap by analysing ~25,000 Flemish texts from 102 participants collected via a 70-day experience sampling study. Participants provided verbal descriptions of their current experiences and how they felt. We computed emodiversity via Shannon's index and estimated granularity using a novel deep-learning methodology based on emotion word embeddings. Results indicated that momentary valence was positively associated with positive diversity and negative granularity. Furthermore, higher negative granularity predicted fewer stress symptoms and greater mental well-being. These findings suggest that emodiversity and granularity capture distinct facets of emotional construal, variety versus specificity of conceptual differentiation – demonstrating that natural language analysis provides a window into emotion construction processes and their implications for psychological well-being in everyday settings.

  • Disentangling Interpersonal and Societal Social Experiences during the Affective Processing of Words

    Social words were often treated as one category [1]. However, some social experiences are first-hand and more interpersonal (gossip), whereas others are more societal (prejudice). We hypothesized that grounded social experiences modulate the affective processing of words. Thirty-three participants judged the affective valences of 360 words varying in social features (interpersonal, societal, non-social) and valence (negative, neutral, positive) during electroencephalography recording. Social features were normed (N=30), and frequency, length, Age-of-Acquisition, valence, arousal, and concreteness [2] were matched across social features within each valence. At P200, interpersonal words elicited the largest effect in negative valence, indicating enhanced attention. At N400, interpersonal and societal words showed attenuated effect in all valence categories, indicating facilitated lexical retrieval. At LPP, interpersonal words showed the largest effect in all valence categories. These findings provide evidence that social words can be categorized into at least two types, one of which is grounded in interpersonal experience.

  • Multisensory Input Boosts Early Numerical Performance

    The Approximate Number System (ANS) predicts mathematical achievements. Acuity in children can increase with multisensory task presentation, consistent with the Intersensory Redundancy Hypothesis in which synchronized multisensory input better supports attention, learning, and memory than single modality input. However, it remains unclear whether multisensory stimulation advances early symbolic number mapping and computations. This study used a longitudinal design across three ANS tasks (non-symbolic, symbolic, and arithmetic), with quantities presented auditorily, visually, or audiovisually. Data from 3065 trials completed by 46 children aged 3–5 years were analyzed using a generalized linear mixed-effects model. Accuracy was significantly higher on audiovisual (M = .61) relative to visual (M = .57; p = .04) trials; this difference was stable across tasks. Thus, multisensory stimulation can support preschool ANS performance across different domains, suggesting an avenue for accelerating early numerical abilities.

  • A Naturalistic Paradigm for Measuring the Intention–Behavior Gap

    Studies using self-reported questionnaires show that procrastination stems not from lacking intention but from failing to follow through on plans, termed the intention–behavior gap. However, behavioral evidence is lacking, especially for quantifying actual behavior and intention over time. To address this, I created the Boring Online Reading Experiment (BORE), a naturalistic behavioral paradigm. Participants have seven days to complete a lengthy, self-paced online reading task with multiple units. Before starting, participants report planned daily progress. Throughout the week, we track actual progress over time. The intention–behavior gap is quantified as the difference between planned and actual progress. Pilot data (N = 64) demonstrated BORE's external validity, with trait procrastination correlating strongly with the observed intention–behavior gap (r = .45, p < .001). Moreover, the pilot study provided behavioral evidence of the intention–behavior gap: participants completed the task significantly later than planned (median actual = 4.73 days; median planned = 3.50 days; p < .001).

  • Can Pseudowords Become "Salty"? Evidence from Evaluative Conditioning

    Evaluative conditioning (EC) refers to changes in the evaluation of neutral stimuli after repeated pairing with affective stimuli. We examined whether a non-evaluative attribute, saltiness, can be transferred from Japanese food words to Japanese pseudowords through EC. Eighty-six native Japanese speakers rated four pseudowords on saltiness, familiarity, valence, likability, and arousal before and after a learning phase in which each pseudoword was paired with either salty or non-salty food words. After learning, pseudowords paired with salty foods were rated as saltier than those paired with non-salty foods, whereas both conditions showed increases in familiarity, valence, likability, and arousal. Saltiness conditioning emerged only among participants who reported contingency awareness and showed better contingency memory, supporting propositional accounts of EC. These findings extend EC research to gustatory meaning and highlight the role of awareness in attribute learning.

  • Quantifying shared knowledge from individual free-association data

    Shared social environments shape minds across individuals, leading to similar perceptions, common interpretations, and convergent emotional responses. These shared internal representations constitute an implicit, structured, and generalizable understanding of the world, termed shared knowledge, that provides common ground for social interaction. Quantifying shared knowledge, however, remains challenging due to its high dimensionality, context sensitivity, and entanglement with individuals' private knowledge. Here, we introduce a method to extract shared knowledge from individuals' free-association responses to images. We show that shared knowledge can be dissociated from individuals' unique components and captured within a low-dimensional subspace embedded in a high-dimensional word-embedding space. Despite its low dimensionality, this shared knowledge predicts decisions in an independent coordination game involving the same images. The findings demonstrate a principled approach to quantifying shared knowledge from easy-to-obtain data and open new opportunities for quantitative investigation into how shared knowledge is formed, represented, and used in the human brain.

  • Spatial Ability Shapes Early Skill Acquisition in Drone Racing

    Drone racing is a highly spatial task, which requires mentally rotating the drone and integrating information across viewpoints, but its relationship with spatial abilities has rarely been examined. In this study, college-aged students played a virtual drone racing game on a desktop computer using an Xbox controller. After learning the controls, participants completed five races on the same track, and race completion time was recorded. In a separate online session, participants completed a battery of validated and efficient spatial ability tests, including mental rotation, mental folding, mental sectioning, and object assembly. Better object assembly ability (r = .43) and mental rotation ability (r = .35) predicted better racing performance, while mental folding and sectioning did not. Spatial ability was a stronger predictor of performance in the initial race, but its influence decreased as participants gained experience, suggesting that spatial ability is most important in the early stages of skill development.

  • Is Less More? - The Role of Linguistic Complexity in Polarized Moral Deliberation

    The increasing polarization of political discourse has sparked growing interest across cognitive science and the social sciences. While ideological divides are often attributed to structural factors such as media echo chambers or algorithmic biases, recent research suggests that more subtle communicative mechanisms may also play an important role. In particular, the linguistic form of arguments –beyond their content– may affect how people evaluate them, how open they are to persuasion, and whether group discussions lead to consensus or entrenchment. This work shows that greater discursive complexity negatively influences both individual appraisal and collective deliberation on highly polarizing topics (e.g., whether freedom of speech should be limited). Using word length and number of words as proxies for linguistic complexity, three experiments (N=11,916) provide convergent evidence suggesting that writing with fewer and shorter words increases communication efficiency, facilitating consensus and reducing polarization on controversial moral issues.

  • Need for Cognition and Contrast Amplification in Political Outgroup Representations

    Evidence suggests that hostility toward members of the opposing political party (affective polarization) reflects how individuals mentally represent political outgroups more than actual ideological distance. We test a cognitive pathway in which motivated cognitive engagement sharpens representational contrasts between social categories, amplifying perceived extremity and downstream evaluative inferences. In a census-matched U.S. sample (N = 302), participants completed measures of Need for Cognition (NFC), perceived outgroup attitude strength, trait attributions, and outgroup dislike. Higher NFC predicted stronger perceived outgroup extremity, which was associated with increased negative trait inferences and indirectly predicted dislike. Though we are careful about causal claims, these findings suggest that cognitively engaged individuals may engage more in contrast amplification, sharpening category boundaries in social belief space and transforming graded ideological differences into globally negative person representations. This work links political polarization to the kind of representational distortion we see arising for other contrast-based categories.

  • Source evaluation during internet reading in Argentinian secondary students: a preregistered intervention study

    This pre-post intervention study with a control group tested whether an educational intervention adapted from Pérez et al. (2018) improved source evaluation skills (about publication media and author trustworthiness) in secondary students. Participants (control group: n=50, intervention group: n=38) had to choose between two contradictory documents (content and source), only one of which was a trustworthy source, and justify their selection, in an initial pre-test; and post intervention, in immediate and delayed (5-6 weeks) posttests. In GLMM analyses, the interaction of Phase*Group was not significant for source selection in immediate or delayed posttest, against pre-registered hypotheses. Analyses of justifications showed a Phase*Group interaction effect: the intervention group mentioned more source characteristics (author position, publication media) than the control group, for both the immediate, _=0.93, 95%CI=[0.61,1.25], p<.001, and delayed posttests, _=0.99, 95%CI=[0.68,1.31], p<.001. These results support the efficacy of the intervention in improving sourcing skills among secondary students.

  • Early Hindi vocabulary and grammar: Developing a MacArthur-Bates CDI

    The McArthur-Bates Communicative Development Inventory (CDI, Fenson et al., 2007), is a parent-report vocabulary checklist and gold standard measure of early language that has been widely used in English-speaking contexts and adapted into over 80 languages. CDIs have been extensively developed and researched in English, Mandarin, and Spanish (the first, second, and fourth most spoken languages), but there is very little work on Hindi, the world's third-largest language. There are multiple adaptations of the CDI to Hindi, but none are publicly available, and there is no published work assessing these instruments or describing CDI data from Hindi-speaking children. We compiled three existing adaptations of the Hindi CDI wordlist and developed novel grammatical items. We are collecting data (current n=138) on this vocabulary superset (961 items) with three goals: to propose a single full-length Hindi CDI, assess its reliability and validity, and make our form and data publicly available.

  • Lexical and structural priming in RPVP-reading

    Rapid parallel visual presentation (RPVP) of multi-word expressions has emerged as a productive method for probing human syntactic ability, with structure-sensitive neural responses reported as early as 130 ms. We asked whether these signals can be primed to shed light on the memory traces created by such input. An MEG study manipulated composition, word overlap between prime and target, and shared syntactic structure in a prime-target matching task. Behavioral facilitation and increased fronto-temporal neural signals were elicited by compositionality, consistent with prior reports of the Structure Superiority Effect. Lexical overlap produced reliable priming in left STG, whereas shared structure alone did not. However, mismatch detection for lexically identical prime–target pairs was facilitated when they differed in syntactic structure, indicating automatic computation of underlying structure. This study shows that a simple matching task on parallel stimuli enables robust investigation of lexical priming in neural responses and syntactic priming in behavior.

  • Neural Dynamics of Flicker-Induced Time Dilation: Extending the OCOS Computational Framework

    Time's Subjective Expansion characterizes the phenomenon wherein dynamic stimuli are perceived as lasting longer than static baselines. While previous work has modeled this in oddball paradigms using habituation dynamics, it remains unclear if the same neural mechanisms account for distortions caused by visual flicker. In this study, we conducted a temporal comparison experiment using asynchronous flickering stimuli at 20Hz, 30Hz, and 40Hz.Behavioral results demonstrate a robust, frequency-dependent time dilation effect: 40Hz stimuli are perceived as significantly longer than 20Hz stimuli, with the point of subjective equality shifting by over 90ms. We apply the On-Center Off-Surround (OCOS) recurrent neural network model to these findings, mapping flicker density to the input clamping level (I) of network nodes. The inherent Winner-Take-All (WTA) dynamics naturally reproduce the observed non-linear time expansion. This suggests that time perception is an emergent property of competitive neural dynamics, modulated equally by novelty (habituation) and sensory density (flicker).

  • CONTEXT DETERMINES THE CHOICE BETWEEN TWO FUNCTIONALLY SIMILAR CONSTRUCTIONS IN BP

    When two constructions apparently serve the same function, which one is chosen by the speaker? We hypothesize that context determines the speaker's choice, and conducted an experiment where participants were asked to choose between the passive or the inchoative constructions after reading a short context. We operationalized their functional differences in terms of agentivity and predicted that speakers would choose the passive in high agentivity contexts, but would favor the inchoative in low agentivity contexts. There were 16 critical items, each presented in two conditions (high x low agentivity). 28 filler items were included, 8 of which were used as catch trials. A logistic mixed-effects model was fitted and the results confirmed our predictions, allowing us to conclude that speakers are aware of the constructions' functions, choosing the ones that fit best in context.

  • Going Viral: The Influence of Anthropomorphic vs. Realistic Animation on Children's Learning About Viruses

    Developing a biological understanding of viruses is critical for appropriate protective and safety measures (Au & Romo, 1996; Menendez et al., 2024). We investigated how anthropomorphic representations influence children's (n=231) learning about COVID-19 through media. We developed three short educational animations (realistic, anthropomorphic, and control) about what happens to viruses inside the body. Children ages 5–8 were tested on their understanding of viruses before and after watching one of the animations. Although most children learned, older children learned more with realistic or anthropomorphic animation than with the control (b = 0.13, pd = 94.33%). Learning was equivalent in the realistic and anthropomorphic animations (b = -0.10). Older children transferred equally to other viruses from the realistic and anthropomorphic animations (b = 0.10, pd = 91.13%) but not the control animation. These results show that animation can scaffold how children learn about invisible biological mechanisms, regardless of anthropomorphism.

  • Singular 'they': Gender stereotype endorsement and grammatical prescriptivism

    Singular 'they' is regarded as less acceptable when used to refer to morphosyntactically gendered antecedents (e.g., waitress) than nongendered ones (e.g., server). However, antecedents of 'they' that lack morphosyntactic gender can nonetheless be stereotypically gendered (e.g., cheerleader). We assessed the extent to which rejection of singular 'they' is driven by endorsement of gender stereotypes. Native English speakers (N = 98) judged the naturalness of sentences with a singular, stereotypically gendered antecedent and either the gender-typical pronoun (e.g., 'she' for 'cheerleader'), the gender-atypical pronoun ('he'), or 'they.' Overall, the gender-typical pronoun was judged most natural, and 'they' was judged least natural. Yet a sizable subset of participants rejected 'they' despite signaling counterstereotypical attitudes in their 'he'/'she' judgments and expressing explicit acceptance of gender diversity. These results suggest that 'they' remains unacceptable in part due to prescriptivist beliefs about grammar that are independent of stereotype endorsement.

  • How do People Universalize? Comparing Resource-Rational Models of Moral Reasoning

    To decide whether it is permissible to break a rule, people often universalize, asking, "What if everyone felt at liberty to do the same?" (Levine et al., 2020). This suggests that people can imagine a world in which an established rule is ignored and reason about its consequences. Consider a queue. When someone cuts in line instead of waiting their turn, observers may evaluate their behavior by imagining a world where people stop respecting the queue entirely. Fully representing such a world is extremely complex: each person plans while anticipating others' moves, knowing that others are doing the same. How do people solve this problem? We compare resource-rational models of universalization against people's judgments. Building on Kwon et al. (2023), we simulate worlds in which line-cutting has been universalized, varying how agents plan (from fast-and-flat heuristics to more exhaustive search) and how they represent others (level-k reasoning).

  • Orthographic Accuracy and Error Patterns in Brazilian Portuguese-English Bilingual Students with Dyslexia

    This pilot study investigates the writing accuracy of five dyslexic students (L1: Brazilian Portuguese; L2: English) enrolled in an international bilingual school in Rio de Janeiro. The study examines whether students' writing errors reflect language-specific orthography or cross-linguistic transfer between Portuguese and English. A dictation task in both languages was administered, and responses were coded for error type and descriptively analysed across overall accuracy, lexical, and orthographic properties. Results showed higher accuracy in Brazilian Portuguese than in English. Error patterns differed across languages: accentuation errors were more common in Portuguese, whereas mixed-rule errors predominated in English. Accentuation errors emerged as a frequent yet heterogeneous category, suggesting both orthographic and prosodic difficulties. Low-frequency and cognate words elicited more errors in English. L2-specific grapheme-phoneme pairs (eg. h~/h/) triggered more mixed-rule errors in English, whereas L1-specific pairs (eg. lh~/_/) elicited more accentuation errors in Portuguese. Overall, preliminary findings suggest systematic orthographic transfer and active use of L2 orthographic rules.

  • Gender differences in students' perceptions and outcome of feedback from embodied AI instructors

    Embodied AI instructors can deliver personalized feedback at scale, but their design may not benefit all learners equally. In a gender-balanced experiment, students wrote an essay and received personalized praise, criticism, and suggestions from an AI instructor with either a realistic or cartoonish appearance. Across genders, the perceived humanness significantly increased the perceived reasonableness of the feedback. However, a key interaction emerged: male students rated feedback from the cartoonish instructor as less reasonable compared to the realistic one, whereas female students' ratings were unaffected by the appearance. Intriguingly, while female students perceived a larger humanness gap between agents than their male counterparts, they did not find the cartoonish instructor less credible, indicating that factors beyond perceived humanness influenced their trust. Furthermore, female students produced higher-quality revisions, particularly following critical feedback. This study demonstrates that AI embodiment triggers gender-specific responses, highlighting the need for equitable design in embodied AI agents.

  • Who benefits more from cognitive training? Analysis of improvement profiles in children using clustering methods

    Executive functions (EF) are important cognitive functions for educational and life success that can be improved through cognitive training. For more than fifteen years, our team has implemented Mate Marote, a gaming software to train and assess EF in children. Interventions last 1 to 4 months and take place within schools, with successful results. We present a clustering analysis of an experiment where 66 6-year-olds' EF were evaluated before and after an intervention of about 27 sessions of 10-15 minutes each, distributed over 3 months. The aim was to determine how participants are grouped according to their improvement after cognitive training. The k-means method was used. In addition, groups were characterized and compared in terms of sociodemographic and academic variables. We discuss results and implications for the analysis of performance profiles in cognitive training, as well as the appropriateness of clustering analysis in the field.

  • Modeling individual differences in learning and memory using large language models

    Computational cognitive models allow formalizing theories of cognition and arbitration between competing accounts. Traditional cognitive modeling (1) entails carefully handcrafting cognitive models, which require substantial domain knowledge and programming expertise; (2) favors single shared model class for all participants, limiting its ability to capture the full range of individual differences. Recent advances in large language models (LLMs) offer a new route for addressing these challenges. In this work, we develop a pipeline for scalable generation of computational cognitive models that prompts an LLM to propose a bespoke cognitive model for each participant given task instructions, their full behavioral trajectory, and a code template, and iteratively refines its proposal based on its predictive performance. In the domains of planning and working memory, we find the resulting models contain participant-specific mechanisms that accurately capture individual differences in learning and memorization, leading to substantial improvement in predictive performance while passing all posterior-predictive checks.

  • Interacting Minds: Gaze Synchronization and Representational Change in Dialogue

    We know that communicating with others can lead to mutual adaptations in the signals adopted. It is less clear whether and how interlocutors converge in their cognitive representations of communicative referents. By tracking eye movements during dialogue and probing interlocutors' representations of a set of referents before and after interaction, we captured the interactional dynamics underpinning mutual representational adaptations. In this exploratory study focused on attention, we investigated the relationship between dyads' gaze dynamics during a communication game, the Tangram task, and interaction-related changes in their representation of a set of referents, measured using explicit and implicit similarity probes. A cross-recurrence analysis of participants' gaze showed, in replication of Dale, Kirkham, and Richardson (2011), that interlocutors' gaze became increasingly synchronized as the task progressed. We further operationalized events of mutual understanding using gaze coherence and investigated the semantic and pupillary dynamics leading to these moments of mutual understanding.

  • Automatically perceiving paths through a scene

    We effortlessly perceive not only what is present in a scene, but what actions and outcomes are possible—such as whether an agent can reach a door or an object might fall. How the visual system infers the paths to possible goal states, while respecting environmental dynamics and constraints, remains unclear. We tested this in simple maze-like scenes using a Go/No-Go task to probe which states are perceived as "leading up to" a goal. Participants pressed a button when an image showed an agent at the maze end next to a target, without being given this verbal description. Participants produced many more false alarms when the agent was closer to the goal along the maze path. By contrast, scenes matched in Euclidean distance but requiring crossing a wall did not elicit false alarms. These findings suggest scene constraints shape perceived similarity between states, motivating extensions to richer environments and domains.

  • Uncertainty in word learning in 14-month-olds

    Hearing words uttered by multiple speakers can help infants learn similar-sounding words. However, speaker variability can hinder learning when multiple cues jointly signal auditory label identity (e.g., speaker gender and voicing), even though they make the words acoustically more distinct. We hypothesize that learning in these contexts is harder because infants face uncertainty about which cues matter. To test this, we manipulated uncertainty in a word-learning task (eye-tracking) by varying how speaker gender and consonant voicing predicted visual referents, while keeping the number of auditory labels, referents, and speakers constant across participants. Consistent with the Hick–Hyman law from adult perceptuo-motor tasks, greater uncertainty delayed 14-month-old infants' (N=50) verification of the correct referent in preferential looking-time trials and linearly increased looking to incorrect competitors. Word-learning situations with more relevant stimulus dimensions—more possible label–referent mappings and consequently more uncertainty—are more difficult, even when the auditory labels become acoustically more distinct.

  • Norming an Odd-Word-Out task for a language assessment tool in Brazilian Portuguese: effect of sensory qualia relations and implicit syntactic structure.

    The odd word out task is standard for assessing verbal semantic processing, yet normed stimuli for Brazilian Portuguese (BP) are lacking. This is problematic given linguistic differences between Brazilian and European Portuguese and the clinical relevance of semantic deficits following brain lesions or during awake neurosurgery. This study examined the influence of constitutive qualia, telic qualia, and implicit syntactic relations on performance in an odd word out task in BP. Healthy participants (n = 210) completed a task involving three-word sequences containing one unrelated target. Accuracy was high across relation types (mean = 94.6%), while reaction times differed, with fastest responses for syntactic relations, followed by constitutive and telic relations, suggesting an influence of implicit syntax on semantic processing. Targets were identified more quickly and accurately in the third position. These results provide normative data for BP speakers and underscore the importance of these factors when designing language assessment tools.

  • From Speech to Print: A Neurocomputational Model of the Predictors of Reading Based on Learned Speech Representations

    We present a reading acquisition model including state of the art architectures that reduce to a minimum ad hoc assumptions. We apply this model to compare early reading acquisition between opaque and transparent orthographies. We use speech representations obtained from wav2vec 2.0 as the auditory part of our model and the pre-trained convolutional neural network CORnet, to represent visual processing. Both modules converge onto a Long-Short-Term Memory (LSTM) model that we train to recognize words. We train the LSTM to recognize auditory presented words, letter sounds and visually presented letters. In Spanish, further training the model to recognize a word from a sequence of visual letters and letter sounds is enough to reach appropriate levels of visual word recognition, generalizing to words it has not seen visually, whereas the same does not happen in French or English. The model explains the difference in relevant predictors in opaque and transparent orthographies.

  • The cognitive cost of lying: using mouse dynamics and unexpected questions to detect false future intentions

    Detecting false future intentions is crucial for prevention and security. The aim of this study was to investigate whether mouse tracking could efficiently identify deception about the future. Grounded on the cognitive load approach to lie detection, we used the unexpected question technique to increase liars' mental effort. Thirty-four participants, randomly divided into two groups (liars vs truth tellers), answered honestly or deceptively a series of control, expected, and unexpected dichotomous questions about a future event. During the task, reaction times, errors and mouse trajectories were recorded. Results showed that unexpected questions significantly impaired liars' performance, leading to higher error rates and greater mouse trajectory irregularity compared to truth-tellers. These findings suggest that unexpected questions amplify the response conflict faced by liars during decision-making. Overall, mouse dynamics appear to be a real-time implicit measure of the cognitive effort associated with deception, offering a promising tool for identifying false intentions.

  • Gesturing Negation: Integrating Multimodal Information

    Human face-to-face communication encompasses more than just words and it has repeatedly been shown that the presence of multimodality (in the form of gesture-speech pairings) can facilitate comprehension under increased processing demands. Despite its frequent use in practically every language in the world, one instance of increased processing demands might be negation—which has repeatedly been shown to elicit processing difficulties. Thus, we investigated the multimodality for compensation account, which posits that negation would specifically benefit from the presence of multimodal information. In three preregistered experiments, we compared compatible pairings of gesture and speech to incompatible (Experiment 1) and unimodal ones (Experiments 2 and 3). Contrary to our initial hypothesis, affirmative information seemed to benefit more from the presence of multimodality than negative information. Two possible explanations for the results will be discussed; namely, the proportion of compatibility and the possible processing demands imposed by redundant information from emblematic gestures.

  • Exploring Lay Theory of Mind

    Understanding observed actions (mentalisation) is thought to rely on latent causal models. However, computational approaches typically employ fixed, experimenter-defined structures tested in artificial tasks. In this study, participants provided probabilistic beliefs about player types, intentions, and actions, then categorised naturalistic gameplay videos by intention or player type, reporting their subjective causal theories. We compared direct versus latent-inverse inference, crossing individual versus aggregate priors, and selecting latent variables according to participants' theories. Latent models outperformed direct models when predicting player type. Intention aggregate priors improved accuracy across metrics. To contextualise the estimates, significant video features and respective weights were elicited from a subset of the sample. Although models weighting the inference-latent variable relation consistently outperformed alternative weighting schemes, supplementary feature-based weights did not significantly alter estimates. Together, these findings demonstrate that incorporating participants' subjective causal structures through cross-path weighting offers a novel, empirically grounded framework for understanding mentalisation in naturalistic settings.

  • Uncovering predictors of dream report similarity: An NLP exploration with transformer embeddings

    Dreams are valuable in the study of cognition as unique internally generated experiences, yet a large-scale analysis via self-reports remains challenging. This study explores Large Language Models (LLMs) for the semantic analysis of dream reports. Using a Dream Bank dataset, we analyzed text similarity with Sentence-T5 Large and emotional content with RoBERTa, employing Representational Similarity Analysis (RSA) and Sentiment Analysis. Results show that dream similarity is best predicted by their characters' similarity, followed by the dreams' settings and evoked emotions, confirming the consistency of the model's numerical representations. While Sentiment Analysis captured general emotions, discrepancies occurred when external context influenced the dreamer's feelings. We conclude that LLMs are powerful tools for large-scale dream research; however, their limitations require expert oversight to validate outputs. This work highlights the potential of LLMs in cognitive science while emphasizing the necessity of a nuanced, human-in-the-loop approach to automated dream analysis.

  • Atypical visual preference in children with ASD from Low-Income backgrounds: a case-control study

    Early identification of Autism Spectrum Disorder (ASD), a heterogeneous neurodevelopmental condition, remains a challenge in developing countries due to structural barriers. This case-control study investigated visual attention preference between social and geometric stimuli using eye-tracking in children with ASD (n=29, 11 girls, Mage=50.3 months) and typical development (TD, n=29, 11 girls, Mage=51.9 months) from low-income families. A one-way ANOVA revealed a significant effect of diagnostic group on the proportion of fixation duration on non-social compared to social stimuli, F(1, 56) = 31.39, p < .001, __2 = .35. This visual preference metric correlated strongly with clinical severity (CARS: r = .53, p < .001) and social responsiveness (SRS-2: r = .55, p < .001). A logistic regression model achieved high discriminative power (AUC = .879). These findings suggest that eye-tracking metrics may serve as a potential adjunctive tool to support clinical screening and objective assessment in underserved populations.

  • Assessing Incremental Syntactic Processing in Spanish with Eye-Tracking Glasses

    Previous research shows children have difficulties comprehending non-canonical sentences. However, studies of incremental processing in Spanish-speaking children are scarce, and online processing of non-canonical sentences remains unexplored. Wearable eye-tracking technology provides an accessible and flexible alternative for studying language processing in children. Building on prior validation against a stationary system, this study examines whether Pupil Labs' Neon glasses capture incremental syntactic processing in Spanish-speaking adults, in preparation for future research with children. Participants completed an auditory sentence-picture matching task with the visual-world paradigm, with structures varying in canonicity and complexity. Gaze-to-target proportions were analyzed to determine when they diverged from chance. For structures with agents in canonical sentence-initial position, gaze diverged early, indicating rapid role assignment. In contrast, non-canonical and structurally complex sentences showed later divergence following additional cues. Results suggest inefficient processing of less salient morphology, promoting delayed role assignment. The findings support wearable eye-tracking for incremental syntactic processing.

  • Layered Time: Circadian States Selectively Modulate Congruent vs. Incongruent Time-Space Mappings

    Although the literature has differentiated various levels of time experience and their effects on human cognition, an individual's cognition occurs simultaneously across these levels. So, how are cognitive processes related to or determined by the interactions between these levels of time experience, if any? The aim of this study was to describe the relationships between chronotypes and the conceptualization of time. The study predicted that the latencies and accuracy of congruent (left-past-right-future) mappings would be similar in optimal and non-optimal schedules. It also predicted that in the asynchronous schedule, latencies would be longer and the accuracy of incongruent (left-future-right-past) mappings would be lower compared to their values in the synchronous schedule. The results partially supported the predictions: incongruent mappings showed longer latencies during non-optimal schedules, but congruent mappings did not perform similarly across schedules. The three-way interaction demonstrates that circadian states modulate time-space mappings, supporting multilevel temporal interactions in cognition.

  • Can delta-band power index decompositional lexical access in Brazilian Portuguese?

    Does phonological onset overlap in a prime-target pair (espada–esparadrapo / sword- adhesive tape) facilitate target recognition as strongly as morphological relatedness (globo–globaliza–åão / globe-globalization)? Brazilian Portuguese speakers perform an auditory–visual priming lexical-decision task while EEG is recorded and analyzed in the time–frequency domain. Prime–target pairs manipulate (i) relation type: onset phonological overlap vs shared morphological base, and (ii) target length (7, 9, 12 graphemes). Within each length, targets vary in morphological depth: simplex [root+category] (e.g., carangueijo / crab) versus derived forms with two to four additional layers (e.g., penaliza–åão / penalization). We test whether delta-band (–ï1–4 Hz) power indexes decompositional structure building: we predict greater delta power for morphologically complex targets, scaling with derivational depth, and weaker or absent scaling for purely phonological overlap. Data collection is ongoing; full analyses and results will be reported at the meeting.

  • Algorithmic Foundations of Counting Development

    One surprising aspect of children's extended trajectory to mastery of the counting process is that, in other domains, children are remarkably adept at acquiring new lexical meanings. Mastery of counting, however, is a years-long process from first learning number words at age 2, to correctly counting around age 4. Previous research has posited that this extended trajectory is due to the difficulty of acquiring the concept of number, a years-long process involving conceptual restructuring. While children must acquire the concept of number, children also need to master the counting procedure. Specifically, counting requires tracking a set of objects, sequentially selecting and individuating each element, maintaining a representation of which items have already been counted, and synchronizing these operations with the verbal count list.The present studies are designed to isolate mastery of the counting procedure from acquisition of the number concept. By isolating non-numerical iterative processes that mirror the algorithmic structure of counting, we test whether procedural competence differs from concept mastery. We present results from a study demonstrating that children can successfully perform simple iterative list tasks before they show full mastery of counting. However, children are only able to perform more complex, theoretically analogous, iterative tasks once they have achieved counting mastery. Together, these findings support a theory in which the extended developmental trajectory of counting arises from the algorithmic demands of iterating through complex lists, as opposed to only conceptual restructuring.

  • Cognitively Grounded Benchmark Generation for Robust Compositional Reasoning in Large Language Models

    Large language models (LLMs) can write fluent mathematics, yet their reasoning collapses when proofs require verifiable composition across many dependent steps. We propose a cognitively grounded protocol for generating robust benchmark datasets in advanced mathematics by operationalizing seven human-inspired dimensions of expert practice: concept formation, dualization, negative knowledge, transfer, invariance control, lemma synthesis, and counterexample search. Rather than toy tasks, we construct meta-prompts from research-grade problem fragments and evaluate complete solution traces, auditing faithfulness and invariant preservation step by step. Under matched conditions, four state-of-the-art systems show a global breaking degree of nore than 90% on stress tests, with failures concentrated in long-horizon planning, premise selection, lemma construction, and counterexample discovery. We conclude with dataset design principles that maximize diagnostic power and reproducibility, and with training levers—rationale SFT, process supervision with reward models, and stepwise preference learning—that directly target step-level correctness.

  • Word recognition selects for morphologically structured representations: The case of English -er

    What kinds of mental representations drive word recognition beyond surface phonetics? We address this question by studying the representations of a speech foundation model at two stages: 1) after pretraining on raw speech audio, and 2) after fine-tuning for word recognition. We focus on English word-final -er, which has three surface-identical morphological sources—comparative (bigger), agentive (diver), and monomorphemic (bitter). We find the model's representational space exhibits a linear geometric relationship between base and derived forms (big_bigger, dive_diver). This structure is specific to particular morphological sources, indicating that abstract morphology can be learned from unlabeled speech alone. However, the model overgeneralizes this pattern to false-friend pairs (bit_bitter). Fine-tuning for word recognition significantly strengthens this structure for true morphological relations, while suppressing relations in monomorphemes and lexically irrelevant information (ex., speaker identity). These results suggest abstract morpho-phonological structure is learnable from raw speech input, and plays a central role in word recognition.

  • Geometric Artifacts or Semantic Bias? The Effect of Embedding Space Anisotropy on Implicit Bias Measurement

    The Word Embedding Association Test (WEAT) was originally developed to assess implicit biases in language models and has since been widely employed as a tool for investigating implicit social biases in human cognition. However, recent research has revealed "anisotropy" in embedding spaces, a phenomenon whereby vectors concentrate in narrow regions, raising concerns that such geometric distortions may confound measurements of implicit social bias. This study examines whether WEAT scores reflect genuine semantic associations or geometric artifacts. We applied a whitening transformation to correct for spatial anisotropy and compared bias scores before and after correction across multiple WEAT categories. Results demonstrated that bias scores decreased substantially following geometric correction, with one racial bias category being eliminated. These findings suggest that standard bias measurements may conflict geometric distortions with semantic content, highlighting the necessity of spatial correction when using embedding-based methods to study human implicit cognition.

  • From Emotions to Virtues: Linking Anger to Justice and Empathic Pain to Humanity

    Emotions serve a host of informative functions, from revealing one's internal states to communicating their goals and intentions. However, little is known about whether emotions signal one's moral character. This work explores this outstanding question in the case of anger and empathic pain. Across two studies, we found that both facial (S1; N= 594) and vocal (S2; N= 572) expressions of anger and empathic pain signaled justice and humanity virtues, respectively. We additionally found perceivers' evaluations of these emotions to be dependent on context. Perceivers were more sympathetic towards empathic pain versus anger expressions in a harm context, but showed no difference in an unjust context. However, in both harm and unjust contexts, perceivers had less negative beliefs about empathic pain expressions. These findings suggest that emotions can serve as multimodal signals of moral virtues and point to the relevance of context in perceivers' evaluations of emotions.

  • Conniving with Continuations: Representing Goals in a Domain-Specific Language of Thought

    Wanting composes flexibly with knowing: we can want to know, want to know what someone wants, and so on. This paper offers a goal representation that allows for this type of rich integration with theory-of-mind concepts. We extend the memo programming language (https://github.com/kach/memo), which is specialized for theory-of-mind reasoning, with a new syntactic construct called "wants" that represents goals. To implement "wants," we borrow an idea from programming language theory: the notion of a "continuation," which allows reasoning about possible future states of a program. We show that this construct can be used to represent a variety of interesting goals, and discuss what it can teach us about the concept of desire more broadly.

  • Automatic morphosyntactic processing in Portuguese as evidenced by electrophysiological responses

    Our event-related brain potentials experiment aimed to test whether we process morphosyntax as a reflex through combinatory mechanisms, based on abstract features. To ensure that brain responses were automatic and independent of attention, we conducted a multifeature paradigm with passive listening experiment. Participants listened to stimuli that were either grammatical ("eu adoto", I adopt), contained 1 violation ("eu *adota", I adopt.3p.sg) or contained 2 violations ("eu *adotam", I adopt.3p.pl). Results showed that, even though there may have been a contribution of the mismatch negativity response on the first interval we analyzed (150-250ms), there was a sustained, directly morphosyntactic effect on both violation conditions (150ms-450ms) and additional, cumulative negativity on the 2 violations condition. This cumulative effect excludes any interpretation of our data based on pure probabilistic processing as verbs in both violation conditions have zero probability of occurring after the pronoun. Henceforth, our data suggests that we process morphosyntax automatically.

  • Ecological Modulation of Facial Preferences and Social Information Use in Human Mate Choice

    Human mate choice is shaped by both individual preferences and social and ecological context. Across three experiments, we test how cues of resource scarcity versus abundance shift (1) preferences for sexually dimorphic facial traits and (2) mate-choice copying (MCC), and whether any effects persist. In Experiment 1, women completed a rapid preference task before and after watching videos depicting harsh (resource-poor) or abundant (resource-rich) environments. Exposure to abundance increased preferences for masculine male faces relative to harsh cues. Experiment 2 asked whether the same environmental cues change MCC, whether this differs for sexually dimorphic versus more neutral faces, and how long any changes last across immediate and delayed tests. MCC was reliable overall, but its longer-lasting effects were observed only in the abundant condition, consistent with reduced reliance on social information under harshness. Experiment 3 repeats Experiment 1 in men using female stimuli to test generality across sexes.

  • Formal Conceptual Blending as a Fundamental Tool for Artificial Mathematical Intelligence and for the Fine-Tuning of Large Language Models

    One of the foundational pillars of Artificial Mathematical Intelligence (AMI) is the taxonomy it has developed of Seminal Cognitive Operations (SMOs) underlying abstract mathematical cognition. Among these, conceptual blending is central. We survey formal approaches to conceptual blending and show how they generate and constrain mathematical reasoning across domains, enabling a quantum jump in building artificial co-creative mathematical agents. The AMI framework fills a longstanding gap in formal systems by integrating principled concept formation into formal reasoning. We additionally propose a regimen for fine-tuning LLMs with mathematical problems ranging from elementary to highly challenging, to bring formal conceptual blending into a constructive chain-of-thought process, resulting in sound final responses.

  • Neural tracking of pre-lexical, lexical and supra-lexical speech features during passive listening

    When humans listen to natural speech, their brain responses become time-locked to certain features of the auditory stimulus. The study of this phenomenon is typically limited to acoustic and linguistic representations of speech. This work examines the contributions of pre-lexical, lexical, and supra-lexical representations of speech as regressors in univariate and multivariate TRF models to predict EEG responses during passive listening of narrative stimuli. Pre-lexical features include acoustic properties and word segmentation. Lexical features incorporate emotive dimensions, sentiment analysis, and sensorimotor dimensions of words derived from both human ratings and transformer-based models. Finally, supra-lexical features include syntactic and semantic characteristics of words. The main finding of our work is that excluding acoustic representations, word surprisal remains the best tracked feature. Overall, our results suggest that LLMs capture multimodal aspects of word meaning and support a view of speech comprehension as an intrinsically predictive process.

  • Morphological Impairment Is Not Unitary: Evidence for Suffix-Selective Vulnerability in Dyslexia

    Morphological processing is central to reading and spelling, yet its impairment in developmental dyslexia remains poorly specified. Dual-route accounts predict that morphological difficulties may arise from distinct mechanisms, including degraded stem–affix representations or impaired morphological assembly/conversion. We examined these possibilities in 33 Turkish-speaking individuals with developmental dyslexia, using error-based analyses of morphologically complex word production in a morphologically rich language. Participant-level measures captured overall morphological error severity and error localization, especially whether errors were concentrated in suffixes or stems. We combined exploratory subgroup discovery using PCA and k-means clustering with confirmatory statistical analyses and phonological control measures. Results revealed heterogeneous morphological profiles. One subgroup showed a clear suffix-biased error pattern above chance, elevated morphological error rates, and no corresponding phonological impairment. This dissociation is most consistent with impaired morphological assembly/conversion rather than generalized phonological difficulty. Findings support subgroup-sensitive, error-based approaches for refining cognitive models of morphological processing in dyslexia.

  • Concrete and Abstract Word Processing in Early Language Acquisition

    Cognitive processes arise from the interaction of biological, psychological, and environmental factors. This research examines how these processes unfold in language acquisition. We hypothesize that access to mental representations is abstract, regardless of conceptual type, and investigate differences in access and processing of concrete and abstract items and lexical access accuracy. Stimulus selection involved (i) a Brazilian Portuguese CDI readaptation completed by 56 caregivers of children aged 24–30 months and (ii) concreteness/abstractness ratings provided by 445 adult native speakers. We conducted an online eye-tracking control experiment using the Lookit platform with 53 adults and a child experiment with 52 children aged 24–30 months. Results show that adults and children differ in how they access and process these words, indicating that conceptual nature modulates lexical access. These findings offer comparative parameters for research on populations at risk for neurodevelopmental disorders.

  • Cognitive drivers of the world's numeral systems: The Numeralbank database

    Numeral systems are uniquely human achievements. They are essential for numerical cognition, present in almost every speech community around the world, invented to reflect and keep track of an important property of things in the world (their quantity) – and yet, there is striking diversity in structure, shape, and function of these systems. Describing and explaining this global diversity is a challenging task. We introduce Numeralbank: an extensible open-access database designed to facilitate the documentation, exploration, and analysis of the world's numeral systems. It provides a large-scale account of their general structural properties and how they relate to usage, communication, and cognition. Numeralbank includes standardized data on numeral systems in over 5,000 languages covering all continents, major language families, and cultural areas. In this talk, we comparatively describe the properties of numeral systems of the world, showing crucial information about the usage, communicative function, and cognitive underpinnings of numeral systems.

  • Resting-state Functional Connectivity Predicts Naturalistic Bilingual Speaking Pattern: Converging Evidence from Hypothesis-driven Tests and Connectome-based Predictive Modeling

    Research on bilingualism and resting-state fMRI (rs-fMRI) connectivity has largely emphasized self-reported static L2 proficiency and use, rarely linking rs-fMRI to real-world speaking patterns. In our study, 79 participants finished rs-fMRI scanning, a detailed language background survey, and a 3-min oral narration. We extracted their L2 speaking proficiency, language entropy, and speech patterns based on hesitations. Linking these behavioral measures to rs-fMRI connectivity with hypothesis-driven and data-driven machine learning approaches (connectome-based prediction modeling; CPM), we found the following patterns: Pauses during speaking emerged as the most reliable, showing consistent associations with Language-Somatomotor networks connectivity and the strongest predictability from CPM. Entropy showed weaker, partially replicable effects consistent with reduced Language-Multiple Demand Networks coupling at higher entropy, whereas L2 speaking proficiency showed minimal and non-replicating connectivity links. These findings highlight the importance of combining both self-reported and objectively derived measures during naturalistic tasks to depict non-native speakers' language profile and functioning.

  • Modality-specific Representations in Multimodal Large Language Models Align with Distinct Neural Systems in the Brain

    The platonic representation hypothesis (Huh et al., 2024, ICML) proposes that as AI models increase in scale and capacity, their internal representations converge toward a shared embedding space of the world. Here, we test whether such convergence extends to alignment between multimodal large language models (LLMs) and the human brain. We collected fMRI and MEG data from two independent cohorts (30 participants each) while participants watched a reality TV show for approximately 30 minutes. The same video segments were input into a multimodal LLM (Qwen-2.5-Omni-7B), from which we extracted text, audio, and visual embeddings using the corresponding modality-specific encoders. We aligned these representations with voxelwise fMRI responses and time-resolved MEG activity across the cortex. Despite being embedded within a unified model, modality-specific representations selectively mapped onto language, auditory, and visual cortices at distinct temporal scales. These results suggest that multimodal LLMs preserve modality-specific representational structure that aligns with dissociable neural systems in the human brain.

  • Predicting Cognitive (In)Efficiency: Heart Rate Variability as a Biomarker for Individualized Responsiveness to Aerobic exercise

    This research investigates the psychophysiological mechanisms of cognitive regulation, specifically how aerobic exercise modulates the autonomic nervous system (ANS) to influence working memory (WM). The study employed an experimental pipeline with 33 university students (mean age = 22.4), comparing a 15-minute moderate aerobic session to a sedentary control group. Using NeuroKit2 for HRV processing, Wilcoxon signed-rank tests revealed significant reductions in RMSSD and MeanNN (p = 0.002), indicating acute sympathetic dominance post-exercise. Furthermore, Mann-Whitney U tests showed a significant statistical gain in WM accuracy for the experimental group (p = 0.013), identifying specific HRV metrics (e.g., RMSSD, SDSD) as predictors of cognitive responsiveness. These findings culminated in the UnnMindFlex mobile prototype, which utilizes real-time HRV biofeedback to personalize cognitive training, advocating for individualized pathways to cognitive optimality

  • Syntactic Comprehension in Healthy Aging: Evidence from Socially Disadvantaged Population

    Even in healthy aging, maintaining autonomy poses cognitive challenges, with language comprehension playing a central role. Socially disadvantaged populations—characterized by lower education, lower-skilled occupations, and limited access to services—are expected to be particularly vulnerable, yet remain underrepresented in cognitive aging research. This study examined linguistic and cognitive performance in two groups of healthy older adults (OA): one socially disadvantaged and one with high educational and socioeconomic attainment. Two control groups of young adults, matched on the same social variables, completed identical tasks. Socially disadvantaged OA performed worse than both their age-matched peers and their young counterparts, with particular difficulties in comprehending syntactically complex sentences. In contrast, higher-attainment OA did not differ from their young controls. This asymmetric pattern suggests that age-related difficulties in syntactic comprehension are not inevitable but strongly modulated by socioenvironmental factors, supporting the view that core syntactic mechanisms can remain largely preserved in healthy aging.

  • Motifs in mental representations of low-level perceptual features

    Prior work has demonstrated structured individual differences in mental representations of colors. Do similar differences exist in other low-level domains? We investigated representations of line orientation using triadic comparisons of Gabor patches varying in orientation angle (0-360–é). When we embedded similarity judgments into personalized 2D spaces, own embeddings predicted judgments better than others' embeddings (p<.001), but no better than orientation (p=0.69), suggesting participants represented similarity by orientation, with some variation. When Gabor patches also varied in contrast (high vs. low), own embeddings predicted judgments better than others' (p<.001) and orientation (p<.001), but not better than contrast (p=.08). When instructed to judge similarity by orientation, participants clustered based on similarity judgments made by orientation angle (n = 13) or by contrast (n = 21). This evidence suggests individual differences in mental representations of line orientation; participants may differ in their ability to make judgments based on one stimulus dimension versus another.

  • Temporal Binding Across Timing Domains: Behavioral Evidence and a Protocol for Causal Manipulation via Transcranial Direct Current Stimulation

    This project investigates temporal binding as a probe of distinct temporal abilities underlying event timing and interval timing. In a behavioral experiment, participants completed four temporal binding tasks under causal and non-causal conditions across multiple action-effect intervals and repeated sessions. Robust and reliable binding effects were consistently observed in event-timing paradigms, particularly the Libet Clock and Response Mapping tasks, whereas interval-timing tasks showed weaker and interval-dependent effects. These results support a functional dissociation between mechanisms supporting temporal localization of events and duration estimation or reproduction. In the second experiment, the causal role of the left angular gyrus was examined using anodal transcranial direct current stimulation in a double-blind within-subject design. Preliminary analyses indicate that stimulation modulates temporal binding in a task- and interval-dependent manner, with clearer effects in interval-based judgments and limited modulation in event-timing tasks. Together, these findings suggest partially dissociable neural substrates for temporal binding across timing domains.

  • Multivariate structural connectivity signatures related to cross-disorder polygenic risk predict negative affectivity

    Genetic 'cross-disorder' risk may influence emotion and behavior by conditioning the structural properties of the brain. Here, we investigated the association between genetic cross-disorder risk and network properties of the structural connectome, and how these relate to negative affectivity in a combined sample of healthy human adults from two large-scale research initiatives. Sparse partial least squares (sPLS) regression was used to define latent components capturing the joint covariance between genetic and imaging features. Polygenic risk scores (PRS) for ten common psychological disorders served as genetic features, while structural connectivity derived from diffusion-weighted imaging served as imaging features. Regression analyses examined the relationship between sPLS-derived component scores and negative affectivity. 135 connectivity edges and PRS of depression, anxiety, and PTSD were most consistently selected to model the joint multimodal relationship. The imaging component scores explained 1% of the variance in negative affectivity.

  • Modality-specific temporal assimilation in a bisection task

    Time perception is fundamental to adaptive behavior, yet its neural mechanisms remain debated. This study tests centralized versus distributed temporal processing models by investigating whether the temporal assimilation effect, in which target intervals are underestimated after short distractors, generalizes across sensory modalities. In Experiment 1 (n = 20), auditory targets showed assimilation only with auditory distractors, not visual ones. Experiment 2 (n = 20) found that auditory frequency variations did not affect assimilation, highlighting modality specificity over stimulus dissimilarity. To further explore this, a tactile-vibratory device was developed and tested. Experiment 3 (n = 20) investigated auditory-tactile pairings but yielded no significant effects, potentially due to the specific interval durations used. These results are consistent with time perception models already reported in the literature, such as state-dependent networks. Taken together, our findings raise questions about centralized models of time perception and suggest the possibility of modality-specific temporal encoding.

  • Modeling Recurrent Neural Networks in Serial Recall Paradigm With Dynamic Self-Excitation

    This study explores recurrent on-center off-surround (OCOS) networks with time-varying self-excitation to model biological neural circuits in working memory. We derive analytical conditions for steady states and global stability, ensuring non-divergent dynamics. Linear stability analysis shows that activity levels and pattern separation depend on the balance between self-excitation (_), lateral inhibition (_), and passive decay (_) through two distinct eigenmodes. Simulations reveal that lateral inhibition regulates co-active nodes, while dynamic self-excitation, specifically the ramp slope and decay time constant (_), controls response amplification and memory persistence. In a serial recall paradigm, systematic variation of _ produces a continuum of serial position curves, transitioning between primacy and recency effects. Our statistical analyses quantify these relationships, linking microscopic network parameters to macroscopic cognitive phenomena. This framework explains how the excitation-inhibition balance manages the trade-off between memory maintenance and interference. These results provide practical guidelines for parameter selection in neural modeling and offer testable predictions for neurophysiological research into working memory mechanisms.

  • How Do Emotional Dynamics in Social Media Language Reflect Depression Risk? An NLP-Based Analysis

    Depression remains undetected in digital contexts, and prior NLP approaches to prediction often lack psychological grounding. This work addresses this gap by investigating how temporal dynamics in emotional language on social media reflect depression risk. Drawing from clinical psychology and affective computing, we analyze longitudinal Reddit data from over 5,000 users (r/depression vs. matched controls). Emotional features, including valence trajectory, agency shifts, and pronoun use, were extracted using lexicon-based (LIWC, VADER) and transformer-based models. Mixed-effects regression and changepoint detection were employed to model fine-grained emotional fluctuations before and after users' self-disclosure of depression. Preliminary findings suggest that increasing emotional volatility, declining agency, and non-linear shifts in affective tone may precede depression expression. This study offers a construct-valid, temporally sensitive framework for detecting mental health risk in language. It advances NLP for mental health by aligning linguistic features with psychological theory and emphasizing the temporal unfolding of emotional signals in naturalistic text.

  • Relative time on the sentence timescale encoded in the brain during naturalistic sentence listening

    Understanding how the brain constructs sentence-level meaning from continuous speech requires identifying neural responses that unfold on the timescale of entire sentences. Using MEG recordings from 71 Dutch participants listening to naturalistic sentences and matched wordlists, we combined temporal response function (TRF) modelling with a group-level canonical correlation analysis (CCA) to isolate neural components that track linguistic structure beyond lexical and acoustic information. We show that the CCA-derived canonical components reveal a shared, sentence-level neural trajectory across participants and across sentences of varying length, consistent with a ramping signal spanning the duration of each sentence. This sentence-long component is dissociable from word-level and envelope-based responses, and is significantly reduced in wordlists lacking syntactic structure. Model comparisons demonstrate that relative time within a sentence—rather than absolute time—best explains the ramping dynamics, suggesting that the brain tracks progress toward the end of the unfolding sentence.

  • The impact of outcome and opportunity inequality on individual creativity

    Economic inequality is a major societal challenge, yet its effects on innovation and creativity remain unclear. In a pre-registered experiment (N = 438), we examined how outcome- and opportunity-inequality affect individual creativity. Participants were assigned to one of three conditions: equality (all received –´5), outcome-inequality (participants received either –´8 or –´2), or opportunity-inequality (participants had either an 80% or 20% chance of receiving –´8). Participants then completed convergent- and divergent-thinking tasks, with performance-based bonuses independent of experimental conditions. Compared to equality, outcome-inequality reduced performance only on the convergent-thinking task. Opportunity-inequality had no effect on either creativity measure. Further, the effect of outcome-inequality on creativity was mediated by negative emotions. These findings highlight the importance of distinguishing between types of inequality and suggest that income inequality could impair constrained forms of creativity, whereas opportunity-inequality may be less impactful on certain behaviors due to greater uncertainty about outcomes.

  • From In-Person Proceedings to Technological Mediation: A Cognitive Reconstruction of Orality in Virtual Justice

    This study examines the transformation of the principle of orality and its corollaries in contemporary criminal procedure, focusing on the impact of digital technologies and, in particular, the use of videoconferencing for evidentiary hearings. It adopts the premise that orality is not merely the performance of spoken acts but an institutional arrangement that ensures adequate cognitive conditions for the production and assessment of oral evidence (Araújo, 2021). Drawing on an analytical and interpretative bibliographic approach, and engaging with empirical research from cognitive psychology and interactional linguistics, the study investigates how technological mediation affects cognitive processes such as shared attention, multimodal perception, discursive processing, and memory, thereby altering the functioning of orality as an evidentiary technique. The research thus seeks to formulate cognitively informed criteria for the normative reconstruction of orality in digital contexts, contributing to forms of procedural innovation that are attentive to the affordances and limits of human cognition.

  • How different text genres affect eye-tracking patterns in fluent Brazilian Portuguese readers: A Work in Progress

    Apesar dos avan–åos na pesquisa sobre leitura em ortografias opacas, os dados oculomotores para o portugu—ís brasileiro (PB) ainda são escassos, particularmente entre adultos com diferentes níveis de escolaridade. Este estudo investiga os efeitos do g—ínero textual nos padr‚Ä∫es de leitura e no processamento cognitivo em PB. Como parte do protocolo internacional MultiplEYE, dados de movimento ocular foram coletados de 100 participantes durante a leitura de sete g—íneros textuais (por exemplo, artigos científicos, blogs) utilizando a tecnologia EyeLink 1000. Métricas, incluindo número de fixa–å‚Ä∫es, dura–åão e sacadas regressivas, foram analisadas para avaliar o impacto da complexidade sintática e da densidade informacional no processamento entre os g—íneros. Os resultados visam preencher uma lacuna na variabilidade linguística e educacional, contribuindo com um corpus oculomotor multigenérico para o PB. Compara–å‚Ä∫es interlinguísticas com outros idiomas no projeto MultiplEYE permitirão análises correspondentes, integrando o portugu—ís brasileiro em modelos globais de leitura e processos cognitivos relacionados ‚Ǩ linguagem.

  • Elevation Affects Demonstrative Choice in English: Implications for Theories of Linguistic Diversity

    Spatial communication systems across languages are thought to exhibit extensive variation (Majid et al., 2004; Evans & Levinson, 2009). Notably, some languages structure spatial communication in relation to absolute features of the environment, even in table-top space (e.g. the cup is uphill/downhill of the teapot) rather than taking an 'egocentric' perspective (e.g. the cup is to the left of the teapot). Here we test if 'elevation' may be important for choice of spatial demonstratives in English – a language usually considered to be 'egocentric-dominant'. Ninetyfive native English-speaking participants rated the acceptability of demonstratives (this/that) to describe images featuring a combination of speaker/hearer/object, while manipulating object distance, position of hearer - and critically the elevation slope of the environment on which all three were situated. Analyses using multinomial multilevel modelling revealed effects of all three variables (i.e. including elevation) on demonstrative rating. Implications for theories of spatial communication diversity are discussed.

  • How Visual Representations Support Metacognitive Monitoring

    Learners often need to compare their own judgments about knowledge with externally generated assessments, yet they may struggle to interpret such information and identify meaningful discrepancies. Although inspectable models are commonly used to support reflection, less is known about how different visual representations shape metacognitive monitoring. To better understand how visual structure influences self-knowledge, we examine how alternative visualizations of self-assessment and external assessment affect interpretation and comparison. We used an interactive environment that presents multiple visual representations of two complementary assessments: individuals' confidence judgments and performance-based feedback. These representations vary in how information is organized, aligned, and contrasted, emphasizing different relationships between the two assessments. We conducted a pilot study with five students. Preliminary results suggest that certain visual structures more effectively support reasoning about discrepancies and reflective judgment than others.

  • Empathy-related differences in emotional valence perception: evidence from the Interpersonal Reactivity Index (IRI)

    Considering empathy as a multidimensional construct with cognitive and affective components, we examined the relationship between individual empathic traits and emotional perception. An online survey was conducted with 37 participants, who completed the Brazilian version of the Interpersonal Reactivity Index (IRI) and evaluated human emotional vocalizations in terms of valence and arousal. Data were analyzed using repeated-measures analyses of variance (ANOVAs), showing that emotional categories significantly influenced valence (F(3, 108) = 99.96, p < .001) and arousal judgments (F(3, 108) = 15.82, p < .001), indicating systematic variation in pleasantness and intensity ratings, respectively. By integrating all stimuli into a general linear model (GLM), we observed a significant interaction between valence and IRI scores (t(292) = 3.28, p = .0012), indicating stronger valence modulation among highly empathetic participants. No significant interaction emerged regarding arousal judgments. Overall, the results indicate that IRI scores are reliably associated with participants' performance in valence recognition.

  • Predicting Second Language Reading Ability from Eye Fixation Patterns During Subtitled Video Viewing

    Reading requires eye movements. This study examines how second-language (L2) reading ability predicts eye fixation patterns on subtitles during L2 video viewing. Thirty-six native Chinese speakers watched English (L2) videos under three conditions: no subtitle, Chinese subtitle, and English subtitle. Eye movements were recorded using WebGazer. L2 reading ability was assessed in terms of reading fluency and reading accuracy. Results showed compared with the no-subtitle condition, there were significantly increased subtitle dwell time proportion and fixation counts in both Chinese and English subtitle conditions. Furthermore, participants with lower L2 reading ability, particularly lower reading accuracy, showed greater reliance on English subtitles, even when eye movements on Chinese subtitles were controlled for. Reading fluency was negatively associated with dwell time and fixation counts during English subtitle viewing. These results suggest that individual differences in L2 reading skills shape visual attention allocation and cognitive load during multimodal language comprehension.

  • How Should Caregivers Speak? The Role of Parentese in Expectations for Caregiver Speech

    Parentese, also known as infant-directed speech, incorporates features such as high-pitched tone, longer duration, and exaggerated vowels. While parentese has been shown to support early language learning in some cultures, little is known about what inferences children make about social relationships from its usage. The current study investigates whether adults and children ages 6-9 infer that a speaker using parentese register occupies a caregiver role to the addressee compared to a speaker using standard speech. Participants listen to auditory interactions in vignettes representing different caregiving relationships (e.g., parent-child, owner-pet) and are asked to identify the relationships. Preliminary results suggest that adults expect the usage of parentese register most with pets and owners, and least expect it with sibling relationships. Our work offers insight into the expectations people hold for caregiver speech, broadening our understanding of how we use language to view caregiving relationships as distinct from other social relationships.

  • Pipeline for SYstematic Curation of Human Experiments

    Large, curated databases have catalyzed key breakthroughs in science—notably the Protein Data Bank's role in shaping AlphaFold. A similar trend is emerging in cognitive science, with the Psych-101 dataset being instrumental in the development of the first foundation model of cognition. However, Psych-101 lacks representative coverage, misses critical metadata, and remains a static dataset with no scope for scaling. We propose that LLM-based agentic workflows can address these concerns. In this work, we first developed a framework to formalize claims within research articles and programmatically validate them against their associated datasets. We then built PSYCHE (Pipeline for SYstematic Curation of Human Experiments), an automated agentic system that iterates based on validator feedback. When applied to studies within Psych-101, PSYCHE produces a dataset with superior metadata coverage and higher claim accuracy compared to the original, bringing us a step closer to the creation of a scalable, 'living' database of psychology tasks.

  • Leveraging Game Structures to Promote Biological Category Learning

    People often assume that biological categories possess a unique essence. This misrepresents within-species variability and the drastic changes that characterize many species' lifecycles. To counter essentialist conceptions, people must not only encounter variable exemplars but also abstract relevant relations across them. We tested a game to promote such learning in young children: a "biologized" version of War using cards showing life stages of two insects, the monarch butterfly and ladybug. Players drew a card on each round, and the player whose card showed the later life stage won. We randomly assigned approximately 150 children (M = 6.5 years) to play War or Memory, a game requiring comparisons based on identity rather than lifecycle relations. Children who played War showed a stronger pretest-posttest shift from essentialist to accurate lifecycle representations for the species in the game. This shows how game structure can be leveraged to promote biological category learning from exemplars.

  • Using Generative AI to Support Retrieval Practice

    Retrieval practice robustly enhances long-term learning relative to rereading, but students struggle to effectively implement retrieval-based activities at scale. The present study examines whether generative artificial intelligence can enable retrieval practice by generating questions with feedback on text passages. Texts were assigned to either rereading or AI-supported retrieval practice conditions. In the retrieval practice condition, students answered AI-generated questions about the readings and received immediate feedback. We assessed learning with multiple-choice or short answer tests either immediately after study or following a delay. Consistent with prior research on retrieval practice, learners in the AI-supported retrieval practice conditions answered more questions correctly than students who reread the material at a delay. These findings suggest that generative AI can effectively instantiate retrieval practice and may offer a scalable tool for promoting durable learning in educational contexts.

  • Operationalizing Inferential Efficiency in Dialogue: A Resource_Rational Framework for Pragmatic Inference

    Pragmatic inference, the recovery of a speaker's intended meaning beyond literal words, is central to everyday conversation, yet it is typically evaluated by accuracy alone. We propose a resource-rational framework that characterizes inferential efficiency along three dimensions: response latency (computational cost), accuracy, and prior reliance (the weight on experiential knowledge vs.\ linguistic evidence). The framework is grounded in a geometric model that decomposes meaning into spatial, temporal, and experiential properties, and employs the Parts of Sense Inference (POSI) tags to quantify experiential priors. Its cross-linguistic validity is demonstrated with examples from English, Malayalam, Tamil, Hindi, and Finnish. An efficiency space reveals trade-offs among speed, accuracy, and prior dependence, captured by a scalar efficiency index. A computational simulation shows that three distinct inferential profiles (efficient, effortful, prior dependent) emerge naturally from bounded optimality. This account reframes variability in pragmatic reasoning as systematic adaptation, offering a bridge between formal semantics, dual process theory, and clinical applications.

  • Motivation for Healthy Lifestyle Behavior

    The adoption of healthy lifestyle behaviors is essential for long-term health and disease prevention, yet engagement varies widely across individuals. Motivational processes play a central role in how health-related behaviors are evaluated, initiated, and maintained. In this study, we examined associations between self-determined motivation, basic psychological need satisfaction, self-efficacy, knowledge of health guidelines, and social network size and multiple health behaviors in a nationally representative sample (N = 1,321). Self-determined motivation and basic psychological need satisfaction were associated with higher levels of physical activity and fruit and vegetable consumption. Self-efficacy was associated with lower levels of tobacco and nicotine consumption, and alcohol intake. Knowledge of health guidelines and social network size did not emerge as consistent significant predictors of health behavior. These findings highlight motivational processes as key cognitive mechanisms underlying healthy lifestyle behaviors and suggest that interventions should focus on supporting motivation in addition to providing relevant health information.

  • What drives intersubject correlation of EEG during the processing of auditory narratives?

    When participants listen to narratives, neural and physiological signals show temporal correlations across individuals. The intersubject correlation (ISC) is increased when attention is directed to the stories suggesting that shared neural and bodily dynamics arise from a similar processing of the narratives. Determining whether ISC primarily reflects low-level acoustic processing or higher-level linguistic processing is important in the clinical context of disorders of consciousness (DoC). Although some DoC patients exhibit evoked responses to auditory narratives similar to healthy controls, ISC is typically reduced, prompting the question of whether this is primarily a difference in processing of the narrative or just low-level auditory processing. In this study, we investigate whether the ISC of the EEG in healthy participants is driven by low-level acoustic (envelope, spectrogram) and/or higher-level linguistic information (word onset, word unpredictability; WU). Combining temporal response functions (TRFs) and correlated component analysis, ISC was computed on the EEG data and on residuals obtained after removing TRF-predicted responses. Significant ISC emerged during both passive listening and attentive story engagement, largely driven by acoustic features. WU uniquely contributed to ISC during active engagement, with temporal and scalp distributions consistent with language processing. However, linear TRF-predicted responses accounted for only a small fraction of total ISC. These findings suggest that combining ISC with linear encoding models can provide meaningful markers of language processing in unresponsive patients. Future work should test whether nonlinear neural dynamics, higher-order narrative features, or arousal-related bodily processes explain the remaining shared variance.

  • Human vs. Large Language Model-Based Sampling: Evidence from Large-Scale Replications

    Large language models are increasingly used to model human behaviour, yet debate remains about whether they can meaningfully approximate both average effects and response diversity in human samples. We revisit replication studies from Many Labs 2 and management-science replications using LLM-based samples, comparing direct prompting with silicon sampling. We extend existing silicon sampling by assigning personas using demographic profiles and Big Five traits. Pilot results from ML2 using GPT-4o-mini show that modelling participant heterogeneity increases response variability: silicon-sampling variance was approximately 2.7 times that of standard prompting, exceeded standard-prompting variance in 81% of focal outcomes, and standard prompting produced zero variance in 41% of focal outcomes. In ML2, 52.9% of primary findings were replicated in human data (using significance in the same direction as the replication criterion). Comparing LLM outcomes to ML2 replications, 38.9% of silicon-sampling outcomes showed consistent signals, compared with 22.2% under standard prompting.

  • Structural principles in the interpretation of blends in Brazilian Portuguese

    This study investigates how morphosyntactic and semantic criteria influence the acceptance and interpretation of blends in Brazilian Portuguese, understood as lexical formations created by merging two base words. Previous research shows that blend formation is shaped by prosodic structure, phonological overlap, base recognizability, and process productivity. Focusing on morphosyntactic and semantic constraints, this study examines whether blends reflect internal structural patterns of Brazilian Portuguese, particularly the tendency for the first constituent to function as the head and the second as a modifier. Adopting the Distributed Morphology (Halle and Marantz, 1997) framework, the hypothesis is that speakers rely on internalized syntactic and semantic structures even in creative word formation. Three experiments were conducted: (i) forced choice between two blends, (ii) forced choice between paraphrases, and (iii) image selection. Results from the tasks support the hypothesis, indicating that blend formation is not random but guided by internalized grammatical principles.

  • Leveling the Playing Field: Environmental Uncertainty Eliminates the Socioeconomic Gap in Working Memory

    Early-life adversity shapes cognitive development, yet its impact on working memory (WM) remains inconclusive. The specialized-sensitization hypothesis posits that adversity-adapted skills are context-dependent: they remain latent in abstract laboratory settings but are recruited when current demands mirror the unpredictability of the developmental context. We tested whether verbal WM varies with socioeconomic status (SES) across contexts that differ in predictability. Ninety-four children completed a within-subjects verbal WM task under two conditions: Uncertainty and Neutral. Accuracy was analyzed using linear mixed-effects models with random intercepts for participants. Results revealed a significant interaction between SES and task condition. While higher SES was associated with better performance in the neutral condition, these performance gap disappeared under uncertainty.These findings support context-sensitive specialization, suggesting that uncertainty does not universally impair cognition but may instead engage adaptive mechanisms tuned to local environmental demands, leveling the playing field.

  • Solving problems and drawing solutions: a cross-linguistic study of ordinality in numerical cognition (Malayalam and Japanese)

    Languages vary in how they describe time. Some use spatial metaphors, others use quantity metaphors. But do languages merely describe time differently, or do they actively shape how we reason about it? Building on prior work on ordinality in numerical cognition and theories of linguistic relativity, this cross-linguistic study investigates whether temporal metaphors influence the perception of ordinality in math problems by using diagrams and strategy choice as externalised representations of thought. We tested 306 participants (children and adults) across two languages: Japanese (distance-metaphors) and Malayalam (quantity-metaphors). Participants solved arithmetic word problems admitting two distinct strategies, and created explanatory drawings. We examined whether the degree of ordinality in participants' drawings predicted their strategy use, and whether this relationship varied across languages and age groups. Results are discussed in light of the linguistic relativism debate and what they say about the role of language in shaping the development of mental representations.

  • Listening Ahead: Phonological Prediction in Down Syndrome

    Phonological prediction allows listeners to anticipate upcoming words based on partial linguistic input. While this ability is well documented in typically developing populations, its status in individuals with Down syndrome remains underexplored. This study examined phonological prediction in individuals with Down syndrome using eye-tracking paradigm. Participants viewed two images while listening to highly constraining sentences (e.g., "The hen laid its–ô"). One image corresponded to the target word (egg; huevo in Spanish), whereas the other was either a phonological competitor sharing the same onset (bone; hueso for huevo in Spanish) or an unrelated distractor. Anticipatory processing was examined during a predefined prediction window prior to target onset. Results revealed reliable visual preference for the target after word onset, indicating successful lexical access. Anticipatory looks to the phonological competitor were observed, but these effects were weaker and temporally delayed. Overall, findings indicate phonological prediction is present in Down syndrome but operates with reduced efficiency.

  • Metacognitive Insight into Event Completion during Everyday Event Perception

    Humans perceive dynamic events as seamless despite incomplete sensory input. This phenomenon, known as event completion, enables perception to appear coherent. However, it remains unclear whether observers have metacognitive insight into this completion when evaluating their own perceptual decisions. To address this question, we replicated the event completion effect using everyday activities (e.g., preparing breakfast) and examined metacognitive insight through metacognitive efficiency. Participants watched short video clips and judged whether a critical contact moment (e.g., grabbing a mug) had been visible and reported their confidence in each decision. Results showed reliable event completion, with participants frequently reporting contact even when it was absent but causally implied. Importantly, metacognitive efficiency was similar for causal and non-causal continuations, indicating that although inferred information can be experienced as perceived, observers' ability to evaluate the accuracy of their judgments remains comparable. These findings suggest that event completion does not impair metacognitive insight during perception.

  • Analogy as reuse in mental model construction

    Humans construct mental models to navigate novel situations, yet how they do so efficiently remains poorly understood. Recent resource-rational accounts have explored what representations should be constructed under computational constraints, but leave open how humans actually discover them – and how prior solutions might be reused in the process. We argue that analogy serves as a core mechanism, enabling people to reuse solution-relevant structure from past experience. Formalising analogies as partial homomorphisms between decision problems, we sketch a framework in which abstract modules, extracted from previous experiences, serve as composable building blocks for new mental models. When a child learns that a password is "like a key," they import not just surface similarity but functional structure: transition dynamics, possible actions, and effective strategies. This modular reuse points toward a process-level account of how humans construct situationally appropriate representations, and connects classic work on analogy with computational approaches to bounded rationality.

  • Typological encoding of agency constraints causal event cognition

    Languages vary in how agency is morphosyntactically encoded, but it remains unclear whether such typological differences relate to non-linguistic representations of causality. We compare Spanish (nominative-accusative) and Basque (ergative-absolutive), two languages that encode intentional vs. accidental causation but differ in agent marking. Thirty-two Spanish speakers and twenty-two Basque speakers completed a verbal event description task and a non-verbal responsibility-rating task using video stimuli from the Causality Across Languages project (USNSF/BCS-1535846). Verbal descriptions revealed convergent strategies for distinguishing intentional from unintentional events across languages. By contrast, non-verbal responsibility judgments differed reliably (Mann-Whitney U = 78, p = .01, r = .68). Spanish speakers organized causal events primarily by agent intentionality, whereas Basque speakers weighted degree of agent involvement more strongly. These findings suggest that typological patterns of agency encoding systematically constrain how causal events are represented, even when linguistic behavior is held constant.

  • How does Semantics Interact with Emotional Prosody in the Bilingual Brain?

    This fMRI study investigated the neural bases underlying semantics and emotional prosody processing in Mandarin Chinese among native and non-native speakers. 15 native Chinese speakers and 15 L1-English L2-Chinese learners completed an emotional prosody judgment task across four 11-minute functional runs. The critical manipulation was the congruence between semantics and prosody in the stimuli. Behaviorally, native speakers outperformed L2 learners, and semantics-prosody congruence facilitated accuracy for both groups. Neuroimaging analyses revealed that emotional prosody recruits a distributed fronto-temporo-parietal network: negative prosody elicited greater activation in temporal regions, whereas positive prosody engaged frontal regions more strongly. Notably, native speakers showed greater activation than L2 learners in semantics-prosody interface regions, such as the superior temporal gyrus, middle temporal gyrus, and inferior frontal gyrus. This increased activation suggests that native speakers effectively engage specialized neural mechanisms for tone categorization, meaning–sound integration, and emotion evaluation in a tonal language.

  • A Principled Framework for Individual Differences in Neural Networks

    Computational modeling is a widely employed method in characterizing individual differences in cognitive mechanisms. However, as cognitive science increasingly turns to complex sequence models, like recurrent neural networks (RNNs), it remains an open question whether and how existing methods for modeling individual differences can be applied to these models. Here, we provide a theoretical framework that generalizes and applies the traditional notion of individual differences to complex sequence models. Using this framework, we design a powerful algorithm that enables the characterization of individual differences using RNN models of cognition. Using synthetic and human data, we show that our theory-guided algorithm robustly outperforms existing heuristics in discovering individual differences from RNNs trained via the vanilla training pipeline. Explicitly accounting for individual differences enables our algorithm to achieve a better behavioral fit than vanilla RNNs. Hence, we establish a principled way to discover individual differences in cognitive mechanisms from behavioral data.

  • Cognitive and Neural Origins of Persistent Mental Content

    Some thoughts and experiences persist in our minds. For example, topics discussed earlier remain fresh in mind and may persist for minutes after the end of a conversation. What cognitive and neural mechanisms underlie this mental persistence? We quantified persistent content by asking participants to generate word chains before and after reading an immersive story and measuring how related generated words were to the story. Across a series of behavioral and one fMRI study, we found that content persisted in mind after reading even if participants had performed an intervening distractor task such as theory of mind reasoning, triangle counting, or reading a new story. Furthermore, instructing participants to suppress story thoughts eliminated semantic biases in generated words, but not conscious thoughts about the story. We propose that persistent content arises from a slowly changing 'deep context' representation that continually biases memory retrieval.

  • Bodily Regulation of Attention and Executive Control under Sustained Cognitive Demand

    Sustained cognitive demands impair attentional efficiency, executive control, and subjective well-being, particularly under environmental distraction. Breathing-based practices have been proposed as low-cost strategies to support cognitive regulation, yet their acute effects under experimentally induced cognitive fatigue remain underexplored. This study examined the immediate effects of a single session of alternate nostril breathing on cognitive performance, perceived stress, and mental fatigue in young female university students exposed to prolonged sustained attention tasks. Participants were assigned to one of three conditions: alternate nostril breathing without auditory distraction, a neutral control activity, or alternate nostril breathing followed by traffic noise exposure. Cognitive fatigue was induced using a Continuous Performance Test, and executive performance was assessed with Stroop and Digit Span tasks. Subjective measures of stress, mental fatigue, perceived cognition, and well-being were collected, alongside exploratory salivary cortisol sampling. Alternate nostril breathing was associated with improved inhibitory control and immediate attention, as well as reduced perceived stress and mental fatigue, even under auditory distraction. These findings suggest that bodily regulatory processes contribute to cognitive control under high mental demand.

  • Phonological Competition Across Languages in Bilingual Spoken Word Recognition

    Spoken word recognition involves competition among phonologically related lexical candidates. In bilingual listeners, this competition may arise both within and across languages. We examined how within- and cross-language phonological competitors are activated during spoken word recognition in monolingual and highly proficient bilingual listeners. Participants heard words in Spanish or English while viewing a within-language competitor, a cross-language competitor, and two unrelated distractors. Eye movements were analyzed in 50-ms bins within a 0–2000 ms window relative to word onset, using log-ratio measures of competitor fixations relative to distractors. Time-course data were modeled using generalized additive mixed models. Both monolinguals and bilinguals showed robust within-language competition in Spanish and English. Cross-language competition, however, emerged only in bilinguals and only when processing their first language (Spanish), but not their second language (English). These findings indicate that bilingual lexical access reflects both language-nonselective activation and input-dependent modulation, highlighting dynamic constraints on cross-language competition during spoken word recognition.

  • The emotional impact not frequency of intrusive thoughts mediate the relation between math anxiety and math performance.

    Math anxiety (MA), feelings of tension during math, is a robust predictor of lower math performance. The Cognitive Disruption Account posits that highly math-anxious individuals underperform because intrusive thoughts consume working memory (WM) resources necessary for math. However, direct evidence linking intrusive thoughts to math anxiety and performance remains limited. Across two studies, we examined how both the frequency and emotional impact of intrusive thoughts relate to MA and math performance. In Study 1, the emotional impact – but not frequency – of intrusive thoughts mediated the relation between MA and poorer math performance. Study 2 manipulated WM demands and replicated this pattern: the emotional impact of intrusive thoughts mediated the relation between MA and performance under both low and high WM demands. These findings extend the Cognitive Disruption Account by highlighting the importance of the emotional dimension of intrusive thoughts and inform broader anxiety–performance frameworks, including Attentional Control Theory.

  • Effects of Poetic Prosody on Metaphor Comprehension During Spoken Sentence Processing

    La comprensión de metáforas puede presentar dificultades en su procesamiento en comparación con el análisis de oraciones literales. Esto se refleja, por ejemplo, en tiempos de reacción más prolongados. Por otro lado, las metáforas son frecuentes en la poesía. Sin embargo, se sabe poco sobre cómo la prosodia modula la comprensión de metáforas. El presente estudio investiga si la prosodia poética influye en la interferencia metafórica durante la comprensión de metáforas orales en espa–ol. Sesenta estudiantes universitarios escucharán oraciones literales y metafóricas producidas con entonación poética o neutra. Los participantes realizarán una tarea de juicio de veracidad literal, seguida de una evaluación de plausibilidad semántica y una tarea de recuerdo libre. Los tiempos de reacción y el rendimiento de la memoria se utilizarán como indicadores de la carga cognitiva y la profundidad del procesamiento. Planteamos la hipótesis de que la prosodia poética amplifica el procesamiento de metáforas. Este estudio tiene como objetivo extender los modelos cognitivos de comprensión de metáforas al dominio auditivo y destaca el papel de la prosodia en el procesamiento del lenguaje figurado. Preferencia de sesión principal

  • Item and binding retention as mechanisms mediating the relationship between attentional control and working memory capacity

    The binding hypothesis as a mechanism for the known relationship between working memory (WM) capacity and attentional control was tested. Three hundred forty-nine participants completed tasks assessing attentional control (Stroop, Simon, Flanker), working memory capacity (Symmetry Span), and two tasks specifically designed to differentiate item and binding errors in visual working memory. Structural equation models showed that attentional control predicted both item memory (_ = .301, 95% CI [.164, .438]) and binding precision (_ = .344, 95% CI [.228, .460]), which predicted WM performance (item: _ = .312, 95% CI [.206, .418]; binding: _ = .260, 95% CI [.140, .380]). Moreover, both pathways partially mediated the relationship between attentional control and WM capacity (item: _ = .094, 95% CI [.043, .145]; binding: _ = .089, 95% CI [.040, .138]). These results suggest that the link between attentional control and WM capacity reflects partially overlapping item and binding processes.

  • Sequential Coordination: How Language Enables Complex Innovation Discovery

    Cultural evolution research demonstrates that population size and network structure affect cumulative innovation (Moser & Smaldino, 2023). However, few models examine how linguistic categories enable coordination in multi-step innovation tasks on dynamic social networks. We extend the Potions Task framework (Derex & Boyd, 2016) by integrating naming game dynamics (Steels, 1995; Puglisi et al., 2008), where agents must coordinate sequential potion selections to discover innovations. Bayesian agents develop shared 2D Gaussian mixture categories through communication, learning potion-label associations while tracking combination effectiveness. Connection weights increase with communicative success and partner innovation performance, driving co-evolution of linguistic categories, innovation trajectories, and network structure. We varied population size, initial network structures, and critical period constraints. Our findings indicate that pressures for linguistic coordination shape the emergence of network structures that facilitate cumulative innovation, and that the equilibrium architecture depends on the balance between categorical diversity and information transmission.

  • Understanding does not guarantee connection in folk theories of conversation

    Do people believe that interpersonal understanding sparks connection? Folk theories of conversation drive social behavior, and a growing body of research suggests that overly-simplified folk theories systematically undermine social connection. Prior work demonstrates that many people underestimate the value of conversation for information gain and relationship building. Less is known, however, about how these dimensions are linked. One could believe that understanding drives connection, or that the two are unrelated. We present a study of 50-minute lightly-structured conversations (N=74) across identity differences (e.g., religion), specifically advertised as an opportunity to build mutual understanding. We find that these participants accurately predict how well they will understand their partner post-conversation; however, they still systematically underestimate increases in closeness, connection, and liking. This indicates that even people who value bridge-building maintain folk theories of conversation that do not inherently link epistemic and relational outcomes.

  • Lay Sensitivity to Sampling and Selection Errors in Scientific Studies

    Laypeople and experts struggle to reason accurately about scientific evidence. Here, we tested whether laypeople noticed methodological flaws related to participant sampling and selection. Across four studies (n = 1,031), we investigated whether lay readers could identify sampling and selection errors related to non-random assignment and limited generalizability in news-style summaries of scientific research, and whether individual differences in thinking dispositions predicted success. Although participants were sensitive to errors when explicitly highlighted in texts, they did not notice most flaws (the exception being biased samples) when comparing articles or judging individual articles. Even when participants identified an article with a methodological flaw, few could articulate the specific issue, suggesting a lack of conceptual understanding. Across tasks, actively open-minded thinking consistently predicted better error detection as reflected by article ratings and free responses, while the effects of cognitive reflection and numeracy were less robust.

  • Rhythm and Mind: Bridges in Early Language Childhood Development

    Rhythm processing has been proposed as a key neurocognitive mechanism for language acquisition, supporting syllabic segmentation and temporal coordination of auditory attention. This study examined the relationship between rhythmic sensitivity and early linguistic and cognitive development in monolingual Spanish-speaking infants in Mexico. Eighty developing children aged 36 months participated in a rhythm discrimination task adapted from previous work, in which they judged whether melodic sequences shared the same rhythmic pattern. Children performed above chance, indicating reliable rhythm discrimination. Linguistic and cognitive abilities were assessed using subtests from the WPPSI-III, including expressive and receptive vocabulary, information, object assembly, and block design. Correlational analyses revealed significant associations between rhythm discrimination performance and expressive vocabulary, as well as visuospatial abilities measured by object assembly and block design. No significant relations were found with receptive vocabulary. These findings suggest rhythm sensitivity is linked to verbal production and visuospatial planning abilities during early development.

  • Validating an experimental paradigm for inducing and measuring habits

    Habits are difficult to induce and measure in human laboratory tasks, and putative behavioral indices of habits often fail to correlate with self-reported habitual tendencies. We recently developed a behavioral paradigm that reliably reveals overtraining-induced inflexibility within a single ~1-hour session, using a Habit Index (HI) that contrasts slips of action between extensively and minimally practiced contexts (Oh & Collins, 2025). Here, we test whether individual differences in HI reflect the same habitization process underlying real-life habits, by examining correlations between HI and self-report measures: the COHS Automaticity subscale, the Habitual Tendencies Questionnaire, and the Self-Report Behavioural Automaticity Index administered with respect to task behavior. Pilot data (N=106) showed positive associations between overall HI and these measures (_–ï.25–.35). We preregistered this correlational study (target N=200) and will present confirmatory results.

  • Sleep Enhances the Consolidation and Integration of Newly Learned Words

    NREM sleep has been associated with the stabilization and strengthening of newly acquired memory traces, whereas REM sleep supports the integration of new information into pre-existing memory networks and semantic structures. To examine these processes, participants (N=50, 18–40 years) learned rare Spanish words paired with images-definitions, and then either took a 90-minute nap monitored with polysomnography (PSG) or remained awake. Memory performance was subsequently assessed analyzing word–image–definition associations and measures of semantic integration. Participants who slept preserved word memory across the retention interval, whereas those who remained awake showed a significant decline, indicating a beneficial effect of sleep on word-form consolidation. Memory for definitions showed a similar but weaker pattern, with no significant group differences. No group differences were observed in semantic integration measures, suggesting a ceiling effect or that short naps may be insufficient to support integration. Exploratory analyses examined associations between PSG-derived sleep measures and memory performance.

  • Bayesian Inference over Data Distributions: A Gaussian Process Approach

    We present BI*, a Bayesian inference framework that places beliefs over data and data patterns rather than model parameters. Unlike conventional Bayesian approaches that require specifying parametric models, BI* operates directly on the space of possible data-generating distributions—one of which represents the true state of the world from which observed data are sampled. The framework infers the probability that each candidate distribution is the true one, given the observed sample. BI* is coherent, broadly applicable across scientific settings, and provides general methods for model selection. In addition, it can incorporate aspects of science often left to human judgment, such as the effects of experimenter bias and measurement error. We make this general theory computationally feasible through fully Bayesian Gaussian Processes. We illustrate the framework's workings with an artificial dataset, demonstrating how beliefs update from prior to posterior, and how data priors influence model selection.

  • Human Preferences of Sycophantic Behavior in Language Models

    Sycophantic behavior—such as excessive flattery or agreement, even at the expense of truth—is a growing concern in language models. Despite the urgency and pervasiveness of this issue, little is known about how humans perceive and evaluate such behavior. Across a series of studies, we examine human preferences for sycophantic versus non-sycophantic model outputs. In Study 1, we measured the gap between participants' considered preferences (what users say they prefer when given time and information) and their unconsidered preferences of sycophantic behavior. In Study 2, participants evaluated responses across normative contexts and situational stakes, finding greater preference for sycophantic responses in low-stakes and reassurance-seeking contexts. In Study 3, participants interacted with sycophantic and non-sycophantic models and reported their preferences, trust, and willingness to use each model in future interactions. Together, these studies characterize sycophancy as a context-sensitive interactional behavior with implications for trust, reliance, and decision making in human-AI interactions.

  • Emojis and inferential processes: an analysis of bridging in multimodal settings

    Situated within Neurolinguistics, this work explores inferential bridging triggered between text and emoji. We draw primarily on Grosz et al. (2021), who posit that emojis are independent discourse units that can have anaphoric effect. Proposing that this observed effect is frequently mediated through an inferential bridge, we investigate the processing of new and given information using ERP methodology. Mirroring the experimental design of Burkhardt (2006), we constructed a corpus of sentence-emoji pairs across three conditions: simple retrieval, inferential bridging, and new information. Two experiments were conducted: an offline norming test to validate the stimuli, followed by a main experiment recording neural activity using EEG (30 participants). The analysis focused on the N400 and P600 components. Following Burkhardt (2006; 2007), we expected distinct electrophysiological signatures for dependency formation and integration effort. We propose that examining neural responses in this multimodal setting will shed light on bridging phenomena and online comprehension processes.

  • What you see is what you guess: Explanations for how a choice was made are paradoxically driven by the choice (regardless of accuracy) in a human vs. AI sorting task

    We present evidence that explicitly explaining holistic classification choices may promote inverted mental models. That is, people believe their choices are based on observed properties but, paradoxically, their choices drive illusory observations. Participants viewed images, some generated by a human artist and others by Google Gemini. They answered questions about each, sorted them into human- or AI-generated categories, then described how they made their choices. The vast majority provided explicit explanations for how they sorted the images. Critically, many gave similar explanations (e.g., perfection indicates AI), even when incorrectly identifying human-created work as AI and vice-versa. In addition, dimensions offered as diagnostic tended to be more abstract when associated with AI (e.g., "artificialness") and more concrete (e.g., "attention to detail") for human-generated attributions. We propose a psychological mechanism for how these mental models may become inverted. Implications for inductive reasoning, naïve theories, human-computer interactions, and (mis)information effects are discussed.

  • Influences of voluntary actions on temporal preparation in multiple temporal contexts

    It has recently been shown that self-initiation of foreperiods (FPs) influences temporal preparation to visual stimuli in choice reaction time (RT) tasks. Here, we investigated whether this effect extends to a simple-RT task using variable FPs. Participants performed the task with 20 FP durations, across three different FP distributions, under two conditions: in the action condition, participants self-initiated FPs with a voluntary action; in the external condition, FPs were initiated externally, after a random interval. Curves of RT as a function of FP duration showed faster RTs in the action condition, irrespective of distribution. Actions also reliably led to more negative slopes of RT curves. Moreover, we found that participants made a larger number of premature responses in the action condition, particularly at longer FPs. This suggests that voluntary actions increase motor readiness, positioning participants closer to response threshold, speeding up responses but hindering the ability to inhibit early responses.

  • Architecture- and data-constrained LLMs don't produce childlike language

    Large language models (LLMs) demonstrate impressive, adult-level language capabilities. State-of-the-art models produce syntactically correct, coherent, and fluent text that is often indistinguishable from human writing. However, their validity as models of language acquisition is less clear. A compelling model of the acquisition process should not only capture mature language use but also reproduce the characteristic syntactic and semantic errors children make during learning. To test this, we simulate children's cognitive and experiential constraints by limiting GPT-2 architecture and training data, and then compare the language output to child language. Variants of GPT-2 constrained in number of transformer blocks, number of attention heads, embedding dimensionality, context length, or number of training steps all fail to exhibit childlike linguistic patterns. These findings suggest that current LLMs are poor theories of the human language acquisition process.

  • Exploring self-biases in reference games

    Classic tasks of pragmatic reasoning involve inferring a possible referent given a speaker's utterance (Frank & Goodman, 2012). Given prior work demonstrating self-favoring biases in perception and causal attribution, we investigated whether listeners show biased pragmatic reasoning when the referent could be themselves. Participants in a fake multi-player reference game were assigned an avatar with two features that were shared or not shared with two other avatars. After completing a drawing task, they were asked to guess who performed the best (prebet), self-rated their performance, and then given a single feature of the player who did best before guessing the winner again (postbet). Preliminary results suggest participants were more likely to choose themselves as best in the prebet, while the postbet reflects a complex relationship between participants' own avatar, their self-rated performance, and the described feature. Planned work will investigate the possible effects of outcome valence and stimulus set.

  • Optimal and heuristic teaching in vast concept spaces

    Humans are remarkably adaptive instructors who adjust advice based on their estimations about a learner's prior knowledge and current goals. Inspired by prior work in rational pedagogy, we model teachers that reason about how learners will update their beliefs when given different examples, and thereby select examples that minimize expected learner error. We demonstrate that Bayesian non-parametric approaches can characterize teaching strategies in continuous domains, where traditional rational teaching models are intractable. We compare human teaching choices against our model and a variety of heuristics. Our model explains significant variance in human choices beyond a mixture of heuristics, and provides insight insight into how teachers formulate pedagogical guidance in computationally tractable ways, even in vast spaces.

  • The Grunt Work of Conversational Grounding: Minimal Responses in Multimodal Dialogue

    Conversational grounding, the establishment of mutual understanding between participants, is fundamental to studies of dialogue. However, methods of grounding multimodal information remain understudied. We investigate how minimal vocalizations (conversational grunts like hm and ooh) encode grounding of verbal information versus observed events in collaborative tasks. Acoustic analysis of 497 grunts from natural dialogues revealed that token selection is strongly specialized by information source. Backchannels like mm-hm predominantly acknowledge speech, while reactive tokens like oh preferentially responded to events. Event-following grunts also exhibit longer duration and more open vowel production at the corpus level, though these differences are largely a consequence of which tokens speakers select rather than independent prosodic modulation. These findings establish observable events as informational contributions requiring explicit acknowledgment, and identify lexical selection as the primary mechanism by which speakers differentiate grounding sources in co-situated task interaction.

  • Classifying Between Congenitally Blind and Sighted Adults with Natural Language Processing Features Using Support Vector Machine

    Semantic memory, the capacity to store and retrieve conceptual knowledge, is central to cognitive science. A key debate concerns how sensory experience shapes conceptual processing and whether semantic representations require sensorimotor simulation. To test this, we trained a Support Vector Machine classifier on natural language features from a Property Listing Task. Recursive Feature Elimination with Cross-Validation selected six optimal features. Univariate analysis revealed no significant differences between congenitally blind and sighted groups for any feature (p > 0.05). Classifier F1-scores did not differ significantly from chance (p = 0.26). However, prediction errors for the six-feature classifier were asymmetric. Accuracy for predicting blind participants (0.31 ± 0.31) did not differ from chance (p = 0.508), whereas accuracy for sighted participants (0.80 ± 0.18) was significantly above chance (p = 0.029). This pattern supports greater semantic heterogeneity in congenitally blind individuals and shows how classification errors can reveal structure in conceptual knowledge.

  • From Finding to Feed: Translating Empirical Research for Social Media

    Cognitive scientists spend much of their professional training learning how to explain empirical research to one another: in talks, papers, posters, grant proposals, and conference discussions. An integral part of our training is the communication and dissemination of scientific ideas at an expert level. Far less attention is given to translating that same work for communities beyond academia. This gap matters because cognitive science produces findings about learning, reasoning, attention, memory, communication, development, decision making, and human-technology interaction—topics that many members of the public already encounter in everyday life. Cognitive research is fascinating, relevant, and often meets urgent needs. Yet when research findings circulate without context, they can be misunderstood, oversimplified, or mistrusted. People may not trust what they cannot understand, and many misunderstandings begin at a basic level: what empirical research is or looks like, what evidence can show, and crucially, why scientific claims are often conditional and open to revision.

Papers with Poster Presentation

  • How Literacy-Related Experience Shapes Spatial–Numerical Associations: A Developmental Shift in Vertical Number–Space Mapping

    Spatial-numerical associations (SNAs) reflect systematic mappings between numerical magnitude and spatial directionality and are shaped by both biological predispositions and cultural experiences. The present study examined the developmental emergence of horizontal and vertical SNAs in Japanese-speaking children exposed to both left-to-right and top-to-bottom writing conventions. In Study 1, children aged 2–7 years and adults completed card-ordering, counting, and sticker-placing tasks in which stimuli were arranged horizontally or vertically. Horizontal SNAs showed a left-to-right bias in children aged 5 and older. In contrast, vertical SNAs changed with age, shifting from a bottom-to-top bias in early childhood to a top-to-bottom bias by around age 7. In Study 2, we examined whether this developmental shift was related to literacy-related experience. Logistic regression analyses showed that hiragana character-sound knowledge significantly predicted top-to-bottom vertical SNAs in 5- to 7-year-olds, controlling for age and nonverbal intelligence. These results demonstrate that SNAs are not fixed mappings but rather flexible representations partially shaped by literacy-related experiences.

  • Real-Time Estimation of Listener Comprehension Using Blink Synchronization

    Blink synchronization is closely associated with cognitive state and comprehension during communication. This study developed a real-time model to estimate listener comprehension using blink synchronization. Data were collected from 22 participants who watched a 3-minute explanatory video. An estimation model was trained on the full dataset, with blink synchronization as a key feature. The model discriminated between comprehension levels, with relatively balanced performance (accuracy = 0.727, F1-score = 0.750, and ROC-AUC = 0.726). Real-time comprehension probabilities were then calculated using a 60-second sliding window updated every second. These estimates generally aligned with overall comprehension labels and suggest group-level stability, with limited late-stage deviations among some low-comprehension participants. Overall, the findings highlight the potential feasibility of blink synchronization as a dynamic indicator of listener understanding and contribute to the development of tools for real-time communication support.

  • Examination of the Relationship Between Metaphor Generation for Abstract Images and Creativity Using Electroencephalography (EEG)

    This study examined electroencephalographic (EEG) activity in the upper alpha band during the generation of descriptions of abstract images to clarify the relationship between metaphor generation and creativity. Shape-based descriptions predominantly involved metaphorical expressions, indicating engagement of metaphor generation processes, whereas color-based descriptions were largely literal. Upper alpha activity during metaphor generation was compared with that during literal expression. Results showed reduced upper alpha power, particularly in posterior regions, during metaphor generation relative to literal expression. These findings suggest that metaphor generation for abstract images recruits enhanced visual and figural processing, reflecting activation of visual cognitive mechanisms underlying creative thinking.

  • KANformer: Personalized Vigilance Estimation with Transformer Features and Kolmogorov–Arnold Sequence Modeling

    Driver vigilance estimation is critical for preventing fatigue-induced traffic accidents, yet existing multimodal EEG–EOG methods often suffer from limited personalization and poor generalization. We propose KANformer, a personalized vigilance estimation framework that integrates subject-specific priors, Transformer-based feature encoding, and Kolmogorov–Arnold Networks (KAN) for adaptive temporal modeling. Raw EEG and EOG signals are first encoded by a Transformer to capture long-range dependencies and cross-modal interactions. A personalized channel attention module then reweights multimodal features using demographic and task-related metadata, enabling subject-aware representation learning. These representations are further modeled by a KAN-based temporal module, which replaces Mamba-style state-space modeling with more expressive functional compositions while retaining low computational complexity. Experiments on the SEED-VIG and SADT datasets under cross-subject and zero-shot settings demonstrate that KANformer consistently outperforms Mamba-based baselines in RMSE, MAE, and PCC. The results indicate that KANformer provides a scalable and generalizable solution for real-time fatigue monitoring and personalized neurocognitive modeling.

  • Some Averse, Some Not: Individual Differences in Algorithm Aversion among Railway Planners

    Human–AI collaboration is common, yet AI-assisted decision-making processes among experts in cognitively demanding tasks remain understudied. We investigated how railway planners evaluate advice attributed to either an AI algorithm or a human colleague, and whether witnessing poor advice influences advice acceptance. Seventy-nine railway planners completed an abstract scheduling task. Bayesian hierarchical models showed similar acceptance of AI and human advice, both before and after witnessing poor advice, at the group level. Despite similar average acceptance across groups, models that explicitly capture individual differences indicate heterogeneity only in the AI condition. Some participants discounted the advice after witnessing poor advice, whereas others did not show this discounting. No comparable pattern was evident in the human advice condition. Our results suggest that inconsistent findings in the literature may reflect individual differences in how people respond to algorithmic advice. Future research should therefore focus on identifying and explaining these individual differences.

  • Immersive Virtual Reality Alters Behavioral and Neural Markers of Spatial Attention

    Traditionally, cognitive neuroscience experimentation has relied on tightly controlled laboratory settings that often lack ecological validity. This study examined whether immersive stereoscopic virtual reality (VR) alters attentional orienting and its neural correlates by embedding a classic Posner cueing task within a three-dimensional environment while recording electroencephalography (EEG). Thirty-five participants completed 2D and VR versions of the task. Behaviorally, the classic cueing effect was observed in both conditions: faster reaction time and higher accuracy for valid trials. However, despite overall slower reaction times in VR, attentional reorienting costs were significantly reduced. Event-related potential analyses revealed no early validity effects on P1 or N1 components, potentially reflecting inhibition of return or increased visual complexity. The P3 component showed larger amplitudes for invalid trials and a posterior topographical shift in VR. These findings suggest that immersive stereoscopic contexts preserve core attentional effects while biasing processing toward more perceptually grounded, bottom-up mechanisms.

  • How do people watch AI-generated videos of physical scenes?

    The growing prevalence of realistic AI-generated videos on media platforms increasingly blurs the line between fact and fiction, eroding public trust. Understanding how people watch AI-generated videos offers a human-centered perspective for improving AI detection and guiding advancements in video generation. However, existing studies have not investigated human gaze behavior in response to AI-generated videos of physical scenes. Here, we collect and analyze the eye movements from 40 participants during video understanding and AI detection tasks involving a mix of real-world and AI-generated videos. We find that given the high realism of AI-generated videos, gaze behavior is driven less by the video's actual authenticity and more by the viewer's perception of its authenticity. Our results demonstrate that the mere awareness of potential AI generation may alter media consumption from passive viewing into an active search for anomalies.

  • Moral Explainability Across Domains: Justification Norms and the Communicability of Decision Processes

    Moral decisions vary in the extent to which they are expected to be justified. We investigate whether normative expectations for justification influence how effectively people articulate decision strategies that others can use. In a two-phase study (N = 573), participants made moral decisions in either a kidney allocation domain or a personal forgiveness domain (high vs. low justification expectations), providing written explanations of their reasoning. New participants either used these explanations to predict the original decisions or made independent decisions, establishing baseline agreement. Participants perceived greater reliance on structured processes in kidney allocation, baseline agreement was higher in kidney allocation, and explanations increased prediction accuracy in both domains, with a larger effect in forgiveness. Despite these differences, matching accuracy converged across domains when explanations were provided, suggesting that moral explainability is fundamentally about communicability rather than introspective accuracy.

  • Cross-Cultural Analysis of Infant Facial Expression: Automated Emotion Detection

    The current study aims to investigate cross-cultural and gender-related differences in infants' facial expression using a novel, infant emotion detection model – the Hierarchical Multimodal Attention Neural (HMAN) Network – for automated Action Unit (AU) prediction. Analyses investigate how facial AUs (1) are distributed across positive and negative emotion classes, (2) co-occur within positive and negative emotion classes, and how (3) their temporal dynamics unfold over time. Compared to Singaporean infants (n=48), Brazilian infants (n=59) were more likely to lower brows (AU04) during negative emotion, and took longer to fully develop a facial expression. Male infants across countries were more likely than female infants to pull lip corners (AU12), and less likely to lower brows (AU04) during positive emotion. These findings advance theoretical understanding of early emotional development and highlight the need for culturally and contextually sensitive tools for interpreting infant emotion, particularly for the design of diagnostic and intervention tools.

  • Analyzing Human Heuristics and Strategies in Everyday Decision-Making Conversations for Conversational AI Design

    Conversational AI increasingly supports everyday decision-making, yet most systems rely on data-centric reasoning rather than the heuristic and interactional strategies people use in natural conversation. To ground design in actual human practice, we analyze 955 real-world Korean conversations (15,476 utterances) involving food and travel decisions, applying a decision-making codebook through an LLM-assisted coding pipeline. Our findings reveal that people prioritize satisficing over optimization, relying heavily on internal knowledge and interactional strategies to manage cognitive load. Critically, we identify a frequency-efficiency mismatch: the most prevalent heuristics sustain conversational flow during exploration, whereas infrequent, rule-based strategies are highly effective at driving resolution during exploitation. By mapping how these patterns transfer across the spectrum of human-AI interaction, this work provides empirical grounding consistent with cognitive theories of decision-making and offers design implications that align AI systems with human heuristic processes.

  • Inferring Joint Commitment from Observed Coordination

    Successful social life requires agents to accurately perceive social bonds and infer whether interacting partners are committed to a shared goal. Yet observable behavior is often ambiguous regarding whether a joint commitment is present. How do observers infer joint commitment from observed behavior alone? We propose a computational model of commitment inference based on Bayesian Theory of Mind. The model predicts that observers update beliefs about latent commitments based on two cues: whether coordination requires mutual participation to avoid a suboptimal outcome (interdependence) and the history of successful coordination (repetition). We tested these predictions in two experiments (N=785) in which participants observed two farmers harvesting berries in a grid world farm and evaluated the level of commitment between them. Results show that both interdependence and repetition strengthen inferences of commitment. These findings demonstrate how observers track the emergence of collective agency from observed coordination.

  • CoSALT: A Resource-Rational Model of Context-Adaptive Thresholds for Detecting Sparse Network Super-Spreaders

    Detecting rare, meaningful signals in non-stationary, heavy-tailed environments challenges biological agents and artificial monitors alike. In network security, analysts must identify sparse super-spreaders within vast benign traffic; as baselines drift, fixed thresholds either overload attention with false alarms or miss emerging threats. We present CoSALT, a resource-rational anomaly detector that frames criterion setting as a Minimum Description Length (MDL) problem. CoSALT adaptively partitions observations into a compressible bulk and a salient tail, while an explicit cognitive selection cost penalizes attending to too many candidates. It further uses spline-based quantile regression to learn context-dependent sparsity expectations. On real backbone-traffic benchmarks, CoSALT improves F1 relative to strong baselines. In a behavioral study, its selection boundaries closely match human visual judgments, especially under explicit budget constraints. These findings are consistent with a resource-rational account of adaptive detection in dynamic, resource-constrained environments, while leaving room for alternative cognitive mechanisms.

  • Mamba-GazeNet: Emotion-Guided Graph–Mamba Architecture for Social Gaze Interaction Understanding

    Gaze interaction behavior recognition (GIBR) is essential for understanding social cognition, yet existing multimodal methods often rely on implicit affect inference and are limited to short-term temporal modeling. We propose Mamba-GazeNet, a multimodal spatiotemporal framework that integrates explicit emotion-calibrated gaze graph construction with efficient long-sequence modeling via Mamba. Facial affect vectors extracted from a pre-trained emotion recognizer are combined with interpersonal geometry to build emotion-aware gaze graphs, while a gated fusion strategy suppresses unreliable emotional cues to ensure structural robustness. A graph embedding layer aggregates intra-frame spatial relations, and a Mamba-based temporal module models long-range interaction dynamics with linear complexity. This design jointly captures affective cues, dynamic gaze patterns, and evolving social relations. Experiments on VACATION and the GIBR benchmark suite demonstrate that Mamba-GazeNet achieves state-of-the-art performance in both full-category recognition and single-category generalization, highlighting the importance of explicit emotion supervision and scalable temporal reasoning for reliable social behavior understanding.

  • Grounding Before Generalizing: How AI Differs from Humans in Causal Transfer

    Extracting abstract causal structures and applying them to novel situations is a hallmark of human intelligence (Griffiths & Tenenbaum, 2005; Holyoak & Cheng, 2011; Lake et al., 2017). While Large Language Models (LLMs) and Vision Language Models (VLMs) have shown strong performance on a wide range of reasoning tasks (Brown et al., 2020; Xu et al., 2025), their capacity for interactive causal learning—inducing latent structures through sequential exploration and transferring them across contexts—remains uncharacterized. Human learners accomplish such transfer after minimal exposure, whereas classical Reinforcement Learning (RL) agents fail catastrophically (Edmonds et al., 2018). Whether state-of-the-art Artificial Intelligence (AI) models possess human-like mechanisms for abstract causal structure transfer is an open question. Using the OpenLock paradigm (Edmonds et al., 2018) requiring sequential discovery of Common Cause (CC) and Common Effect (CE) structures, here we show that models exhibit fundamentally delayed or absent transfer: even successful models require initial environmental-specific mapping---what we term environmental grounding---before efficiency gains emerge, whereas humans leverage prior structural knowledge from the very first solution attempt. In the text-only condition, models matched or exceeded human discovery efficiency. In contrast, visual information---in both the image-only and text-and-image conditions---overall degraded rather than enhanced performance, revealing a broad reliance on symbolic processing rather than integrated multimodal reasoning. Models further exhibited systematic CC / CE asymmetries absent in humans, suggesting heuristic biases rather than direction-neutral causal abstraction. These findings reveal that large-scale statistical learning does not produce the decontextualized causal schemas underpinning human analogical reasoning, establishing grounding-dependent transfer as a fundamental limitation of current LLMs and VLMs.

  • Multi-Level Narrative Evaluation Outperforms Lexical Features for Mental Health

    How people narrate their experiences offers a window into how the mind organizes them. Computational approaches to therapeutic writing have evolved from lexical counting to neural methods, yet remain fragmented: dictionary tools miss discourse structure, while embeddings conflate local coherence with global organization. No existing framework maps these techniques onto the hierarchical processes through which narratives are constructed. Here we introduce a three-level framework - micro-level lexical features, meso-level semantic embeddings, and macro-level LLM narrative evaluation - and show, across 830 Chinese therapeutic texts spanning depression, anxiety, and trauma, that macro-level evaluation substantially outperforms lexical and embedding features for mental health prediction. This challenges the field's emphasis on word-counting: formal structural features (Labov's story grammar, RST coherence, propositional composition) demonstrate that narrative organization per se carries predictive signal, while clinically-grounded narrative dimensions capture how psychological states are expressed through discourse. Semantic embeddings add minimal independent value but yield incremental gains in multi-level classification. By grounding computational levels in discourse processing theory, this framework identifies macro-structural organization as the primary locus of clinical signal and generates testable hypotheses for intervention design and longitudinal research.

  • Rational Communication Shapes Morphological Composition

    Human languages expand vocabularies by combining existing morphemes rather than inventing arbitrary forms. Communicative efficiency shapes lexical systems at multiple levels (Gibson et al., 2019), yet morphological composition---combining morphemes through compounding or affixation---has rarely been modeled as a historically situated speaker choice among competing morpheme sequences, leaving unanswered why a language settles on one morpheme combination over other plausible alternatives. We ask whether a trade-off between listener recoverability and speaker production cost can predict attested compositions over contemporaneously available alternatives. Here we show, within the Rational Speech Act (RSA) framework (Frank & Goodman, 2012; Goodman & Frank, 2016) using a time-indexed lexicon constructed from Corpus of Historical American English (COHA) and Corpus of Contemporary American English (COCA), that across 4323 naturally occurring English compounds and derivations spanning 1820--2019, attested compositions are systematically ranked above unattested alternatives generated from contemporaneously available morphemes. Models integrating semantic informativeness with production cost outperform semantic-only and cost-only baselines on Mean Reciprocal Rank (MRR) and top-k accuracy (Acc@k), with the advantage of the Pragmatic Speaker model (__1) over the semantic-only baseline growing as the candidate set expands, where meaning alone leaves morphological choice underdetermined. These findings suggest that lexicalization reflects a communicative trade-off between expressiveness and efficiency, extending rational accounts of communication from utterance-level choice to the internal structure of words.

  • Lumos: A Cognitive Control-Inspired Agent for Long-Horizon Threat Investigation

    Large language model (LLM) agents have shown strong capabilities in long-horizon tasks such as cyber threat analysis. However, despite competent step-level reasoning, existing agents often fail to maintain coherent behavior over extended investigation trajectories, leading to accumulated process-level errors. Prior approaches mainly improve individual reasoning steps through prompting or reflection, but provide limited support for sustained goal maintenance, strategy regulation, and execution monitoring. We propose Lumos, a cognitively inspired agent framework that realizes an explicit executive-control loop via process-level control mechanisms, including goal-centric memory for long-term goal maintenance, failure-driven strategy adaptation for inhibitory regulation, and action and output control for explicit monitoring and verification. Experiments on real-world security investigation scenarios demonstrate that Lumos consistently outperforms strong baselines, improving long-horizon stability and overall task success.

  • Rethinking SMT Defect Inspection from Human Visual Cognition: Modules that Boost Defect Performance

    Human visual recognition is shaped by constraints such as sensitivity to peripheral evidence, integration of complementary visual cues, and tolerance for ambiguous category boundaries. We ask whether these principles can improve automated inspection in Surface Mount Technology (SMT), where defects are often subtle, off-center, and difficult to distinguish categorically. We implement these ideas in a standard classifier using lightweight modules for position-aware normalization, Fourier and color/position cue integration, and ambiguity-aware learning over confusable defect classes. Experiments on industrial and public PCB datasets show that these cognitively inspired modifications improve defect recall under low false-alarm conditions and encourage more consistent focus on defect-relevant regions. The results suggest that human-inspired perceptual constraints offer a useful framework for designing robust machine vision systems in applied settings.

  • Do cross-linguistic animacy-based constraints on plural marking reflect learning biases?

    Nominal number marking is widespread across languages, yet number distinctions are often restricted to specific categories. Typological research suggests that such restrictions follow the Animacy Hierarchy Constraint (AHC), whereby higher-ranked categories in the animacy scale (human > animate > inanimate) are more likely to exhibit number contrasts than lower-ranked ones. However, large-scale cross-linguistic evidence for the AHC is limited, and a mechanistic explanation for its typological prevalence is missing. This study combines typological and experimental studies to investigate the impact of the AHC on number neutralisation in nominal paradigms, and assess whether the AHC mirrors cognitive biases at play during language learning, which could in turn explain cross-linguistic regularities. We analyse data from 509 diverse languages using Bayesian hierarchical models, demonstrating a robust monotonic increase in number neutralisation down the animacy hierarchy. To examine whether this cross-linguistic regularity reflects learning biases, we conduct an artificial language learning experiment manipulating animacy-conditioned number neutralisation patterns. Results show that systems conforming to the AHC are more learnable than AHC-violating systems. Together, these findings suggest that the AHC is supported by learning biases that contribute to its typological prevalence.

  • Children's reasoning, visual attention, and social cognition in Mexico and the US

    Cross-cultural variation in cognition is increasingly well-documented, but less research has explored the developmental origins of this variation. In addition, Central & South American cultures are underrepresented in research with both adults mand children, but particularly the latter. This study helps to fill this empirical gap by comparing Mexican and U.S. preschoolers (34–62 months; Mexico N=153; U.S. N=125) across six tasks spanning relational and similarity reasoning (causal Relational Match-To-Sample, Taxonomic-Thematic Triads), visual attention (Picture Free Description; Feature Task), and social cognition (Uniqueness Preference; Symbolic Self-Inflation). We find cross-cultural differences in four of the six tasks, with Mexican children showing greater contextual and relational sensitivity than U.S. children on several measures. In contrast, cRMTS and feature- based recognition do not show reliable country differences, underscoring that cultural contrasts depend on the specific task and developmental process being measured.

  • When Emotion Slows You Down: Valence-Arousal Conflicts Across L1 and L2

    Across two experiments, we tested whether valence and arousal interact during speeded affective judgements, and whether this interaction differed between pictures and words, or between L1 and L2 words. German–English sequential bilinguals immersed in an L1 environment completed an affective judgement task. Experiment 1 focused on valence judgements (e.g., positive vs. negative) whereas Experiment 2 focused on arousal judgements (e.g., arousing vs. calming). In Experiment 1, a Valence–Arousal interaction pattern was present only within picture stimuli. In Experiment 2, this pattern was evident for all stimulus types showing that task demands shaped when the interaction was observable in words. Across both experiments, the interaction did not differ between L1 and L2 word blocks, suggesting that both languages were processed similarly. These findings suggest that speeded affective appraisals were not strongly language specific in highly proficient bilinguals. L2-only analyses showed that bilingual experience variables moderated response times.

  • The Geometry of Macro-Syntax: Geometric Resonance and Hierarchical Derivation in the Left Angular Gyrus

    A central challenge in cognitive neuroscience is identifying the algorithmic principles that allow the human brain to transform linear linguistic input into hierarchical macro-syntactic structures. We propose the Geometric Resonance Hypothesis, postulating that the Left Angular Gyrus (L-AG) operates on a representational geometry that is topologically isomorphic to the deep derivational manifolds emergent in Large Language Models (LLMs). Utilizing fMRI data from the Natural Stories Corpus and the BERT transformer architecture, we track the alignment between neural activity and the model's computational hierarchy. Residual Representational Similarity Analysis (RSA) reveals a distinct phase transition: neural alignment remains minimal in lexical layers but peaks significantly in the deepest integration layers. Quantitative manifold analysis demonstrates that this resonance is driven by the intrinsic structural complexity of the representational space. Furthermore, a causal mediation analysis establishes that the relationship between syntactic complexity by Dependency Locality Theory scores and L-AG BOLD responses is significantly mediated by the model's macro-syntactic manifold. These results suggest a biological convergence, where disparate substrates, biological circuits and artificial networks, converge upon shared geometric solutions to resolve the complexity of human language.

  • Dissecting the Neuro-Cognitive Dynamics of Vision-Language Models

    Vision-Language Models (VLMs) have achieved remarkable success in multimodal reasoning, yet the extent to which their internal representations mirror human neural dynamics remains an open scientific question. In this study, we propose the Component-Specific Residual Disentanglement (CSRD) framework to systematically evaluate the alignment between 17 state-of-the-art VLMs and human fMRI activity. Unlike traditional layer-wise approaches, CSRD disentangles the contributions of Attention and Feed-Forward Networks (FFNs) by selectively retaining or stripping residual connections, thereby isolating information integration from functional transformation. Leveraging the Natural Scenes Dataset (NSD), our analysis reveals a hierarchical "ascent" in brain-model alignment, but with a critical functional divergence in deep layers: Attention modules act as "cognitive hubs" maintaining high alignment with the human visual cortex, whereas FFNs progressively decouple from biological signals to prioritize task-specific predictions. Furthermore, we find that instruction-tuned generative architectures significantly outperform contrastive baselines (e.g., CLIP) and scaling-law-based expectations. These findings suggest that the generative objective fosters more biologically plausible representational structures, providing new theoretical guidance for designing brain-inspired multimodal systems.

  • Strategy Discovery for Long-Horizon Physical Manipulation Puzzles

    To support the systematic study of long-horizon human strategies in physical problem solving, we introduce a set of 28 virtual physics puzzles that require sequential manipulation and involve diverse interaction mechanics such as pick-and-place, pushing, and tool use and creation. The paradigm provides rich interaction logs and object-state trajectories, enabling computational analyses of strategies. We analyze data from 38 participants and estimate puzzle difficulty and participant performance. Using trajectory-based clustering, we identify recurring solution strategies and quantify solution diversity and repeatability between attempts. The results show that puzzles requiring both tool creation and pushing are the most challenging to solve with the longest time-to-solution and the highest failure rates, as they combine the large search spaces of tool use with the challenges of precise control and less predictable dynamics. Notably, participants demonstrate strong improvement after solving the puzzles once, suggesting learning and reuse of discovered strategies.

  • Automatic Identification and Classification of Individual Alpha Frequency

    Individual alpha frequency (IAF) is a key marker of stable differences in neural processing speed and attention, and a critical parameter for frequency-tuned interventions such as transcranial alternating current stimulation (tACS), yet it is still often estimated with fragile, peak-picking heuristics. We propose a machine-learning algorithm that automates IAF detection in EEG data by (i) attenuating the 1/f background via exponential detrending, (ii) localizing candidate alpha peaks, and (iii) classifying them with a support vector machine using morphological features. On a longitudinal dataset (N = 204 young adults; frontal and parietal ROIs, two sessions), we benchmarked the method against consensus human ratings and the Philistine algorithm. The new approach halved localization error (MSE = 0.18 vs. 0.44) and reached 94% accuracy in detecting the presence of an IAF peak, while providing a continuous certainty measure. This enables expert-level, confidence-weighted IAF estimates for individualized cognitive and clinical protocols.

  • Cognitive Offloading to AI Distorts Confidence Calibration: Effects of Memory and Judgement Support in AI-Assisted Decision Making

    Appropriate AI reliance requires users to accurately evaluate their own performance in order to discern whether to retain or defer responsibility. While underconfidence is a known driver of automation bias, little is known about how collaboration with AI itself influences users' metacognition. We investigated how cognitive offloading different stages of the decision-making process affects confidence calibration. Participants completed diagnostic decision-making tasks with varying levels of memory and judgement support. Overall, cognitive offloading impaired confidence calibration, with unaided decision makers showing the greatest alignment between confidence and accuracy. Importantly, offloading different cognitive processes produced distinct metacognitive biases: judgement offloading led to overconfidence, whereas combined offloading of memory and judgement processes led to underconfidence. These findings demonstrate that AI support can disrupt user confidence calibration in systematic ways, depending on the type and extent of cognitive delegation. The results highlight metacognitive miscalibration as a critical and underexplored consequence of human-AI collaboration.

  • Looped Manifold Trajectories of Dynamic Functional Connectivity Reveal Continuous Task-related Brain Reconfiguration

    Dynamic functional connectivity (dFC) matrices capture time-varying interactions among large-scale brain networks, but their high spatiotemporal dimensionality complicates the interpretation of how whole-brain functional organization is configured and reconfigured across cognitive tasks. In this work, we introduce a geometric framework that embeds sliding-window dFC from resting-state and task fMRI into a shared low-dimensional manifold. Within this space, resting-state dFC embeddings cluster within a compact intrinsic resting-state region of the manifold that serves as a functional baseline, whereas task engagement yields smooth, looped trajectories that depart from and return to this baseline. Distinct tasks trace separable loop structures that remain stable across parcellation resolutions. Phase-dependent sampling of dFC matrices along these loops reveals structured, time-ordered reconfigurations of functional networks during task performance. Together, these findings depict brain dynamics as continuous, recurrent trajectories rather than discrete state transitions, providing a unified geometric representation of task-evoked dynamic functional reconfiguration.

  • Treating Distillation as Pedagogy: Simplicity and Scaffolding in Large Language Models

    Human pedagogy relies on scaffolding to tailor information complexity to a learner's capacity, yet current AI knowledge distillation often subjects resource-constrained student models to computational cognitive overload via unstructured, verbose reasoning. To investigate the alignment between human learning principles and machine learning dynamics, we propose a unified computational framework that treats distillation as a controlled cognitive experiment. We manipulated the granularity and sequence of the pedagogical signal across two distinct tasks. Our results demonstrate two key phenomena paralleling human cognition: a simplicity effect, where concise elucidations significantly outperform complex expert rationales by minimizing extraneous cognitive load, and a scaffolding effect, where a simple-to-complex curriculum proves essential for convergence in reasoning tasks. These findings provide computational evidence that less is more in the context of knowledge distillation for resource-constrained models, suggesting that effective alignment in student-teacher paradigms requires a shift from maximizing information quantity to optimizing pedagogical quality.

  • Contribution of Cognitive and Semantic Control to Performance on Stroop Tasks

    The Stroop effect is perhaps the most recognized cognitive phenomenon in psychology. People are slowed at naming ink colors, when the ink is used to represent a word denoting a color. Some theories suggest early perceptual conflict, whereas others suggest that later response conflict is the cause of the effect. We examined whether adding a divergent thinking aspect to a Stroop experiment would enhance the effect, as this would occur only at the response selection stage. In a controlled experiment, we found that making a choice at the response stage produced a Stroop effect of identical magnitude to the traditional incongruent condition of the Stroop task. This may suggest that Stroop effects can be produced at the response generation stage. However, correlation and regression analyses hinted that this was linked to the processing limits of semantic control, rather than domain-general cognitive control, which may underlie the traditional Stroop effect.

  • BrainTune: An Artifact-Aware Meta-Contrastive Framework for Fast and Lasting Personalization of EEG Emotion Models

    The practical application of EEG emotion recognition is hindered by two critical challenges: reliance on offline processing pipelines that assume access to complete session data, and high inter-subject variability that complicates personalization. To address these, we propose BrainTune, an Artifact-Aware Meta-Contrastive Framework. BrainTune features a causal stream-compatible pipeline with an artifact-aware graph neural network to bypass offline dependencies, as well as a meta-contrastive strategy using supervised multi-level contrastive learning for the backbone to handle session/subject/task domains and meta-learning for the classifier to enable rapid personalization from brief calibration. Experiments on SEED, DEAP, and DREAMER demonstrate that BrainTune achieves state-of-the-art performance and exhibits "lasting personalization", where early calibration progressively benefits future sessions. Code is available at https://github.com/iewug/BrainTune.

  • Free or Systematic? Structured AI Guidance Improves Tutor Knowledge-State Estimation Quality and Self-Regulated Learning in Programming Education

    Generative AI can accelerate programming learning, but friction-free, answer-oriented help may weaken learners' self-regulation. We compare unstructured answer-oriented assistance with systematic scaffolding that provides step-by-step hints, knowledge-state feedback from a Graph-enhanced Interactive Knowledge Tracing (GIKT) model, and difficulty-matched practice. We conduct two within-subject crossover studies across Java and Python programming courses. Students alternate between conditions by concept unit with counterbalanced starting conditions. The systematic scaffolding condition delivers interactive guidance through a Learn-Practice-Evaluate-Support loop, providing Socratic prompts and step-by-step hints without revealing final solutions, along with personalized practice recommendations and diagnostic views. We analyze primary outcomes of self-reported transfer (Likert 1--6) and tutor knowledge-state estimation quality, quantified by Expected Calibration Error (ECE) under time-split evaluation, using linear mixed-effects models. We demonstrate that systematic scaffolding yields higher self-reported transfer with convergent trends in exam subscores, and substantially better prediction calibration (ECE: 0.13 vs. 0.24, effect size d=1.64). Process analyses reveal increased planning and monitoring engagement alongside reduced time pressure and frustration, transforming engagement behaviors into productive learning processes while enhancing the diagnostic value of interaction traces. Our findings demonstrate that systematic scaffolding outperforms unstructured answer-oriented assistance, establishing it as the superior approach for AI-supported programming education.

  • Band-Gated Identity-Disentangled Training for Cross-Subject Auditory Attention Decoding

    EEG-based auditory attention decoding (AAD) seeks to identify which speaker a listener attends to in multi-talker "cocktail party" settings. A central challenge is cross-subject generalization: neural responses vary substantially across individuals, inducing distribution shifts across subjects; consequently, models trained in subject-dependent or mixed-subject regimes may latch onto subject-specific cues that hinder transfer and interpretation. To address this, we propose a band-gated multi-band framework that decomposes EEG into low- and high-frequency pathways and adaptively fuses them at the sample level to learn attention-discriminative representations while accommodating inter-individual spectral variability. We further introduce an identity-disentangled objective that leverages confidencefiltered pseudo-labels to perform alignment in an auxiliary bottleneck space, encouraging a more subject-invariant bottleneck representation while mitigating subject-specific variability. Evaluated on KUL, DTU, and AVED under leave-one-subject-out protocols and two decision-window settings, our approach achieves the best or highly competitive performance against strong baselines. Analyses of the learned gate and representation geometry provide qualitative support for the roles of adaptive band reweighting and identity suppression in improving robustness in cross-subject AAD. Code is available at https://github.com/siyingtao/BDGI_for_AAD.

  • Bridging Acoustics and Semantics: Native Language Experience and Hierarchical Temporal Integration in the Human Brain

    Deciphering the real-time transformation from acoustic signals to semantic meaning remains a challenge in cognitive neuroscience. Leveraging high-temporal-resolution magnetoencephalography data from native Chinese speakers and the Whisper computational framework, we investigated the neural mechanisms underlying speech processing. We demonstrated a robust scaling law, whereby larger models increasingly align with neural activity. Crucially, a model fine-tuned on native language experience outperformed generic multilingual models, underscoring the brain's sensitivity to language-specific statistics. Furthermore, variable-context analyses revealed a hierarchical temporal architecture: cortical integration of low-level acoustic features occurred within a rapid 40 ms window, whereas high-level speech representations required an integration window of approximately 400 ms. These findings provide a quantitative account of how native language experience and hierarchical temporal integration shape the neural encoding of human speech comprehension.

  • Comparison classes and alternative salience in deriving adjectival upper bounds

    Listeners often strengthen the interpretation of expressions by deriving upper-bounded readings (URs), as when warm is understood as warm but not hot—an inference typically assumed to arise due to reasoning about a pragmatic scale 'warm, hot'. But, UR rates are also shown to vary substantially across and within scales. We investigate this diversity for gradable adjectives, aiming to explain why less URs arise for relative adjectives with relative scalemates ('warm, hot') than absolute ones ('likely, certain'). We tease apart a view that attributes this effect to differences in salience of the stronger term from one attributing it to the way standards are set differently for these adjective types. Two experiments manipulated alternative salience and comparison classes while controlling for adjective type, showing that comparison classes modulate UR rates 'rel, rel', but not 'rel, abs' scales, while alternative salience affects both. This pattern challenges salience-only accounts but is compatible with an account where URs depend on the way adjectival standards are derived from the comparison class.

  • Do infants track multiple altercentric perspectives in an active behavioral task?

    Studies suggested that infants may track others' perspectives (Baillargeon et al., 2010; but see Dörrenberg, et al., 2018), showing altercentric effects (Kovács et al., 2010; Manea et al., 2023). However, it remains unknown whether infants can track multiple perspectives distinctively, a hallmark of real-world social interactions, and how this relates to their developing self-concept. We presented a search task to 15-to-18-month-olds (N=96), involving two agents with different perspectives (true/false-beliefs). We varied which agent prompted infants' search, and asked whether search is modulated by the prompter's perspective. Infants' displayed interference from the other's perspective when the false-belief agent made the prompt in the false-belief, but not the true-belief trials, and not when the true-belief agent made the prompts. Furthermore, such altercentric influences were unrelated to self-awareness, as measured by the Mirror Self-Recognition task. Thus, infants may sustain multiple other perspectives besides their own and rely on them differentially.

  • Why Produce Garden-Path Sentences? Evidence from Corpus Data across Registers

    Garden-path sentences induce processing difficulty by requiring readers to revise an initial syntactic analysis (e.g., The policeman saw the lights were off). While garden-path comprehension has been widely studied, far less is known about why such sentences are produced in natural language. We thus address this gap by analyzing large-scale written-text corpora in English and Japanese across five registers. English Noun/Sentence garden-path constructions, arising from optional complementizer omission, were frequent in blogs and newspaper, suggesting that their emphasis on communicative efficiency promotes their use. Rather, Japanese Main Clause/Relative Clause garden-path constructions, which arise from grammatical constraints, were most frequent in legal texts, reflecting the register's tendency to favor complexity within a clause. Legal texts also exhibit the longest ambiguous regions across constructions, which can increase reanalysis cost. These findings suggest that grammatical properties and register-specific writing styles jointly shape the production of garden-path sentences in real-world language.

  • From "Few" to "Many": Explicit Spatial Associations for Language vs. Automaticity for Numbers? Evidence from Card-Sorting, Lexical Decisions, and Numerical Comparison

    Do linguistic representations of magnitude (e.g., few, many, cold, hot) share the same spatial architecture as numerical quantities? While the Spatial–Numerical Association of Response Codes (SNARC) effect shows that numbers are automatically mapped onto a mental number line, it remains unclear whether linguistic magnitudes elicit similar automatic spatial associations. We conducted three experiments to address this issue. Experiment 1 showed that Italian quantifiers and adjectives are explicitly mapped onto a horizontally compressed scale. Experiment 2 found no evidence for an implicit Spatial–Linguistic Association of Response Codes (SLARC) effect, with Bayesian analysis providing strong support for the null hypothesis (BF01=87). Experiment 3 confirmed robust SNARC and ratio effects for Arabic digits in the same participant pool. Together, these findings indicate that although linguistic and numerical magnitudes share a similar psychophysical structure, spatial mapping in language may be task-dependent and lack the automaticity observed in numerical cognition.

  • Fear and Anxiety Differentially Reweight Control Objectives: Evidence from Inverse Optimal Control

    Exposure to height induces pronounced changes in human postural behavior, yet the computational mechanisms underlying these adaptations remain unclear. Previous studies report both reduced and increased postural sway under height exposure leading to seemingly contradictory interpretations. Here, we apply inverse optimal control (IOC) to infer how control objectives governing postural behavior are reweighted under threat. Participants performed quiet standing in a virtual reality environment under ground and height conditions while joint-level kinematics were recorded. Postural control was modeled as a multi-joint optimal control problem with fixed dynamics, and condition-specific cost weights on joint velocity and control effort were inferred. Height exposure was associated with increased penalization of joint angular velocity and altered effort weighting, consistent with more conservative control strategies. Substantial inter-individual variability was observed. Moreover, anxiety-related measures were linked to increased control variability, whereas fear of heights and heart-rate change were associated with motion-suppressive control. We thus provide computational evidence for a functional dissociation between fear of heights and anxiety in human balance control.

  • iSense: Passive Inference of Student Self-Regulated Learning and Emotional Regulation Using Smartphone

    Early detection of self-regulation challenges in university students is critical for academic success, yet conventional assessment methods struggle to capture these behavioral processes' multidimensional and dynamic nature. To address this, we develop iSense - a smartphone-based system that enables privacyaware, passive monitoring of self-regulated learning (SRL) and academic emotional regulation (AER) through multi-modal behavioral sensing (social interactions, mobility, app usage, and sleep). In our 12-month longitudinal study involving 211 college students, the system successfully identified significant behavioral correlates of SRL/AER states. Our hybrid populationpersonalized prediction model achieved mean absolute errors of 7.9% (SRL) and 9.3% (AER), representing a 22% improvement over conventional baselines. Importantly, the results demonstrate how knowledge transfer from student populations can facilitate the development of accurate, personalized models that require minimal individual-specific data for rapid adaptation to new students. This scalable solution overcomes limitations of traditional self-report methods for diverse educational settings.

  • Neuro-Composer: Bridging Intention and Articulation via Dual-Stream Predictive Coding and Metacognitive Feedback

    Direct speech synthesis from neural activity offers a lifeline for severe paralysis, yet decoding high-dimensional, noisy signals remains a formidable challenge. Current approaches rely on passive regression, overlooking the hierarchical and generative nature of speech production. We propose Neuro-Composer, a biologically grounded framework reframing decoding as the dynamic integration of top-down semantic planning and bottom-up acoustic execution. Grounded in Dual-Stream Theory, our architecture decomposes decoding into a ventral stream extracting abstract intent via foundation models (BERT/Wav2Vec 2.0), and a dorsal stream capturing fine-grained articulatory dynamics. These streams are synthesized via a latency-aware Predictive Refiner. A bio-mimetic auditory feedback loop enforces semantic consistency, endowing the system with self-monitoring capabilities. Experiments demonstrate Neuro-Composer significantly outperforms baselines in reconstruction fidelity. Analyses reveal emergent semantic manifolds and brain-like gating patterns, providing computational validation for neurocognitive theories of language production.

  • The Content of Common Ground Shapes Coordination

    Common ground is widely understood to facilitate coordination, but what happens when the shared expectation is competitive rather than cooperative? In a repeated coordination game, pairs received advice that was either cooperative or selfish with common ground manipulated by whether pairs knew they received identical advice. We found that advice type shaped the pattern of coordination: pairs with selfish advice converged toward asymmetric equilibria, while pairs with cooperative advice maintained symmetry. Common ground amplified this pattern: pairs who knew they shared selfish advice showed the steepest increase in round-level payoff asymmetry as rounds progressed. We interpret this as recursive reasoning shaping equilibrium selection: common ground on selfish advice, increases expectations of selfishness, making acceptance of asymmetry rational—anticipation reinforcing the pattern. Linguistic analysis complements this, where participants' use of "we" tracked fairness outcomes. These findings demonstrate that common ground shapes not just whether coordination succeeds, but the kind of coordination.

  • Modeling Decision-Making with Will for Cooperation in Social Dilemmas

    Standard rational actor models often attribute cooperation failures in social dilemmas to insufficient incentives, overlooking the destabilizing effects of continuous utility maximization. To address this, we propose a framework of "will" defined as a mechanism that persistently pursues goals while ignoring local cost-benefit fluctuations. We formalize the Willed Agents as potential minimizers, distinguishing them from cumulative utility maximization. Dynamical analysis of infinite population demonstrates that willed agents shrink the feasible state space, acting as boundary constraints that accelerate convergence in canonical social dilemmas. Through multi-agent simulations in a spatiotemporal Stag Hunt Game, we show that willed agents function as "cooperation catalysts", enabling groups to surmount high-risk thresholds where purely utility maximization fails. We find that heterogeneous will strength promotes cooperation, and that agents who autonomously suspend rational re-evaluation can significantly outperform continuous optimizers. These findings suggest that successful cooperation relies on the cognitive capacity to strategically constrain calculation.

  • Children Seek Help from Technological and Human Agents Depending on Domain

    Seeking help from knowledgeable others is a crucial part of human learning. Prior work shows that children consider informants' social and epistemic traits when deciding whom to trust, but has primarily focused on human sources. As artificial intelligence becomes increasingly integrated into children's lives, they face new choices about whom to seek help from. The present study examines young children's help-seeking preferences between human and technological agents across various domains. Children chose which agent to consult when faced with problems involving encyclopedic information, moral harm, or conventional norms. Results show a clear pattern of help-seeking based on domain. Children preferentially seek help from technological agents for encyclopedic questions, but favor human informants when faced with problems about moral transgressions or social norms. These findings suggest that children's help-seeking behavior is not driven by a general preference for humans or technology, but instead reflects domain-specific understanding about the kind of knowledge different sources can provide.

  • A neural process model of compensation and adaptation as independent mechanisms in speech motor control

    Experimental paradigms that perturb auditory feedback induce two distinct but related behavioral responses: (1) compensation, a real-time, within-token, adjustment to speech, and (2) adaptation, an adjustment to speech that persists even after the removal of perturbation, evidencing learning (e.g., Houde & Jordan, 1998; Purcell & Munhall, 2006a, 2006b; Tourville et al., 2008). In this paper, we present a fully dynamic model of one-shot adaptation, i.e., learning after one trial of perturbation (Hantzsch et al., 2022), in the framework of Dynamic Field Theory (DFT; Schöner & Spencer, 2016). The model triggers compensation and adaptation via prediction-perception mismatch driven by lexical-item-based predictions. Compensation derives from resolving multiple inputs into a neural field and via update to the mapping between phoneme representations and production targets, while adaptation relies only on the latter. We discuss how the model, which integrates speech motor control with linguistic representations, accounts for established results while generating novel predictions.

  • Working Memory and Syntactic Processing in ADHD: Evidence from Relative Clauses

    Individual differences in executive control, including working memory capacity (WMC), are known to modulate sentence processing, but their role in neurodivergent populations remains unclear. This study examined how WMC relates to the processing of subject- and object-relative clauses in Brazilian Portuguese speakers with and without ADHD, using a WMC task and a self-paced reading paradigm with comprehension questions. Online measures showed that WMC modulated reading times differently across groups at the onset of the relative clause in object-extracted structures. This pattern suggests that both expectation-based and memory-based factors contribute to processing difficulty, consistent with locality accounts. Offline measures showed that accuracy was reduced in the OR condition and modulated differently by WMC across groups, with WMC predicting higher accuracy in the control group but not in the ADHD group. These findings suggest that reading times may reflect different processing strategies across cognitive profiles.

  • Multi-Agent Debate under Reasoning Chain Attacks: Exploiting Cognitive Vulnerabilities

    With the advancement in large language model (LLM) capabilities, multi-agent debate has emerged as a crucial approach for tackling complex problems. However, security research on this front remains largely focused on traditional adversarial attacks that manipulate surface-level semantics, neglecting the deep reasoning pathways behind model decisions. To bridge this gap, we propose a novel reasoning chain debate attack that targets cognitive vulnerabilities within LLMs' internal reasoning chains during debates. This framework dynamically identifies cognitive vulnerabilities in target agents by deploying virtual agents with real-time evolutionary capabilities. Subsequently, it executes precise cognitive attacks using a weighted fusion attack strategy. We also design an adaptive termination mechanism that automatically halts debates when attack effectiveness stabilizes, thereby reducing computational overhead. Experimental results validate the framework's effectiveness: across mainstream LLMs, the attacker achieves over 56% persuasion success rates while reducing token consumption by approximately 27%, providing novel foundational research for multi-agent security.

  • How does language bias memory for gender stereotypes?

    Combating stereotypes ("girls do ballet") may not be as simple as conveying counter-stereotypical messages ("boys do ballet"). For example, hearing generic language ("girls/boys do X") predicted greater stereotyping among young children than hearing specific language (that girl/boy does X"), even when parents explicitly conveyed counter-stereotypical messages (Benitez et al., 2025). Why might counter-stereotypical messages delivered generically backfire, reinforcing the very stereotypes they aim to combat? We propose that generic language activates category-level representations stored in memory, bringing to mind stereotypical information about the implicated categories. 183 adults and 104 4- to 6-year-olds recalled the language (generic vs. specific) and message (stereotypical vs. counter-stereotypical) heard in 12 statements. Adults, but not children, systematically misremembered counter-stereotypical messages as stereotypical ones when language was generic. The task was difficult for children, perhaps explaining their recall patterns and high exclusion rate (48% of 200 children were excluded). Follow-up work probing recognition memory is underway.

  • Imagining internal mental representations for planning

    Computational approaches to human planning propose that people construct simplified mental representations of their environment, known as value-guided construals (VGC), for computational efficiency. The VGC model successfully predicts patterns of external attention when people are required to navigate mazes externally. However, it remains unknown whether internal selection processes untethered to external stimuli (i.e., offline planning and imagination) engage similar VGCs. We developed a novel maze navigation task in which the start and goal locations are revealed after stimulus offset. Despite the physical absence of the maze stimulus, we find that participants remain capable of constructing efficient mental representations aligned with the VGC model. These internal plans demonstrate significant but weaker influences of spatial grouping, similar to the pattern demonstrated for external attention. Together, our work reveals how VGC extend to purely internal, imagined task representations. These findings will help guide future research on how human-like planning may be implemented in artificial systems.

  • Was the Ishango Bone—'The Oldest Numerical Device'—a Mathematical Tool? An Empirical Approach

    The Ishango bone—a 10cm, approximately 20,000 year old notched bone—has been frequently claimed as evidence of human numerical competence in the Upper Paleolithic, on the basis of patterns in the collections of notches' quantities. We test whether the notch collections are suitable for university-level numerate adults to perform numerical tasks precisely. We find that performance on a range of simple numerical tasks is significantly below the precision required for exact calculation. Moreover, in an explicit counting task, we find that increasing the size of the bone to about hand-length substantially improves performance. These results are at odds with numerical interpretations of the Ishango bone, and raises the question of why the notches were produced at such a small scale given these performance difficulties and their amelioration after moderate size increases.

  • Separable contributions of value and choice to policy learning

    Reinforcement learning (RL) models have been highly successful in describing human behavior. However, in most experiments, reward history is highly correlated with choice history. Consequently, behavior could be partially driven by a value-free habitual learning (HL) mechanism but be misinterpreted as RL. Do people select actions they have chosen frequently in the past (HL), or actions that have been rewarding in the past (RL)? To disentangle the influence of value-based RL and value-free HL, we designed two experiments that unconfound reward and choice histories. Behavioral and modeling results suggest that both value-based and value-free mechanisms contribute to test phase choice. Their relative contributions differ in the two experiments, suggesting that different environments may recruit the two processes to different degrees. Our results highlight the existence of partially redundant processes that are often difficult to disambiguate, and call for careful consideration of underlying processes when interpreting experimental results.

  • Do First Impressions Matter? How Problem Representation Shapes Performance in Complex Math Problem Solving

    Initial perceptions of math problems are critical for performance; novices tend to fixate on surface structures, while experts look deeper towards strategies. However, it remains unclear how these early representations change across problem-solving phases and explain performance. In this study, 126 undergraduates solved 20 SAT-level math problems: 10 timed (90sec time-limit) and 10 untimed. For each problem, participants rated perceived calculation—a surface feature of problem representation—after brief (3sec) exposure and again after attempting the problem. Participants substantially overestimated calculation demand at first exposure and revised these judgments downward after solving (Cohen's d=0.84). Participants with lower self-efficacy showed even larger revisions (Cohen's d=0.61). Critically, only pre-solving perceived calculation demand predicted performance (partial r=-0.45), an effect primarily mediated by task-specific self-efficacy (49.65% of the total effect). Math-anxiety also contributed explanatory power, particularly under time-pressure. Overall, effective problem-solving appears to rely on rapid, structure-based representations that guide strategy selection.

  • More than One Agreement Planning Mechanism: Reflexive Number Source Sensitive to Syntactic Structures

    Prior studies on agreement attraction in sentence production have investigated the differences in verb and reflexive agreement planning, as evidenced by a distinction in their error rates. This study investigated two possible accounts: Planning Order, where verbs and reflexives share the same underlying mechanism that is applied at different stages, and Number Source, where verbs and reflexives retrieve features from different sources. To further adjudicate between the two, we extended the empirical ground to reflexives in complement and adjunct positions. Results replicated previous findings on the distinction between verbs and reflexives, while no significant difference was found in attraction effects among reflexive conditions, thereby supporting Number Source. Predicted by neither account, we novelly observed a significant increase in baseline error rate for adjunct but an insignificant increase for complement, suggesting a finer-grained Number Source model sensitive to syntactic structural differences.

  • Cognitive Correlates of Team Performance: Variations Across Dimensions of Skill and Degrees of Expertise

    While individual cognitive metrics have been shown to predict real-world task performance, their scalability to team behavior remains poorly understood. This study addresses the critical disconnect between individual cognitive abilities and team-based outcomes. We present an approach for studying this relationship using rich behavioral data collected from a virtual team-based task. Then, we further extend the analysis to show that these relationships can evolve with changes in team expertise. Regression models were used to predict performance across two dimensions of team skill, derived from an individual's ability to (1) perform task-specific actions (taskwork), and (2) coordinate their actions with others (teamwork). Individual performance across seven cognitive tasks were used as predictor variables. Results indicated significant associations between efficient information processing abilities and taskwork skill, in the game. Teamwork skill initially varied with visuospatial working memory. However, attention control predicted teamwork skill after sufficient practice in the task. Finally, independent measures of analytic intelligence reliably predicted both skills; however, task knowledge may have mediated that effect. Our results demonstrate that team performance is multifaceted, with distinct cognitive abilities contributing to various dimensions of team skill, at various stages of expertise.

  • Computational Evidence for the Functional Necessity of Marr's 2.5D Sketch and Gestalt Principles in 3D Scene Perception

    The human visual system robustly reconstructs 3D scenes from fleeting glances and sparse 2D inputs, effectively solving an ill-posed problem. In contrast, machine vision systems like 3D Gaussian Splatting often fail under such conditions, suffering from severe overfitting. We hypothesize this gap stems from the absence of biologically plausible inductive biases: specifically, Marr's 2.5D sketch and Gestalt organizational principles. To test this, we operationalized these theories within a computational framework, incorporating surface orientation cues (Marr) and the Gestalt principle of Pr–âgnanz to enforce structural simplicity and suppress redundancy. Experiments demonstrate that these cognitive constraints are computationally necessary to resolve geometric ambiguities in sparse-view environments. Crucially, our biologically constrained model yields representations judged by humans as significantly more realistic and structurally coherent than unconstrained baselines. These findings provide compelling computational evidence supporting the functional necessity of intermediate geometric representations and global organization principles in facilitating robust human 3D perception.

  • Listeners store category identity and uncertainty in memory during spoken word recognition, but not acoustic detail

    During spoken language processing, listeners maintain gradient information about past speech input in memory, but it is unclear how detailed these representations are. We test this question in an experiment using the speech discrimination paradigm. We focus on two perceptual decision-making effects: (1) more distant stimuli are more accurately discriminated; and (2) people are more confident about accurate choices, and this relationship is stronger as stimulus distance increases. We consider three types of stimulus 'distance' as a proxy for level of representational detail: (a) distance in acoustic cue space; (b) distance in category certainty space; and (c) difference in category identity. We assess which of these levels drive decision-making effects, finding that only category identity and category certainty are significant; difference in acoustic space was not a driver of perceptual decisions. This suggests that during real-time processing, listeners can maintain uncertainty about past speech, but cannot access its acoustic details.

  • Group-Aware Cognitive Diagnosis via Variational Item Response Theory

    Group-level cognitive diagnosis is fundamental in educational assessment, where examinees are organized into groups like schools or classes. However, most methods rely on mean-field assumptions that treat students independently, ignoring dependencies induced by shared resources. This oversight yields unstable estimates, particularly under sparse observations. We propose Group-aware Latent Ability Diagnosis (GLAD), a hierarchical variational framework that models within-group dependence via a latent group context while retaining scalable inference. We decompose student ability into a global mean, a group effect, and an individual residual, yielding three variants: a group-conditioned prior, a deterministic effect, and a stochastic effect. For efficient inference, we develop a permutation-invariant encoder to aggregate student representations. Experiments on real-world Eedi data demonstrate consistent improvements over baselines across multiple CDMs.

  • Investigating the Role of Task Representation Switch Costs in Goal Persistence

    Goal pursuit profoundly shapes human cognition. This is often beneficial – for example, by focusing information processing. However, goal-dependent computations occasionally lead to seemingly maladaptive behavior, such as a bias toward goal persistence – the tendency to continue pursuing the current goal even when suboptimal. While various psychological explanations have been proposed for goal persistence, the underlying cognitive mechanisms remain unclear. Here, we explore one potential hypothesis. We posit that switching goals incurs the computational cost of reconfiguring internal representations, e.g., to new stimulus-action mappings, and that this cost motivates persisting in a goal even when another is more valuable. Thus, we predict that 1) higher reconfiguration costs will increase persistence bias, and 2) given the choice, participants will prefer switching to lower-cost goals. To test these predictions, we developed a task where participants chose between competing goals, and sudden changes in the task dynamics encouraged goal abandonment toward more valuable options. Crucially, some goals shared the same action selection rule, while others required different rules, allowing us to test how similar vs. different internal representations influenced goal persistence. Across two experiments (N = 125), we replicated goal persistence biases and found that stimulus-action representation switch costs influence both goal persistence and selection preferences. However, they only accounted for a small part of the total goal persistence, suggesting that such biases are predominantly driven by other factors.

  • Learning Domain-Invariant Representations for EEG Emotion Generalization via Representational Similarity Alignment

    Cross-subject electroencephalography (EEG) emotion recognition is challenged by pronounced inter-subject variability, which hinders generalization to unseen subjects in the domain generalization setting. Most existing approaches focus on aligning feature distributions, whereas the invariance of relational structure in latent representations has received far less attention. Inspired by representational similarity analysis (RSA), we revisit cross-subject EEG generalization from a structural perspective and propose an RSA-based structural consistency constraint (RSASC). This regularizer aligns representational dissimilarity matrices (RDMs) computed from different subjects' latent representations, encouraging the model to learn emotion-relevant representations that are more consistent across subjects. Experiments on the SEED and SEED-IV datasets show that the proposed regularizer yields stable improvements in cross-subject generalization: it achieves the best accuracy on SEED, and on SEED-IV it attains competitive performance with the lowest variance. These results suggest that structural invariance offers a novel perspective that complements distribution alignment for learning subject-invariant EEG representations.

  • Learning to Look before Learning to Like: Incorporating Human Visual Cognition into Aesthetic Quality Assessment

    Automated Aesthetic Quality Assessment (AQA) treats images primarily as static pixel vectors, aligning predictions with human-rating scores largely through semantic perception. However, this paradigm diverges from human aesthetic cognition, which arises from dynamic visual exploration shaped by scanning paths, processing fluency, and the interplay between bottom-up salience and top-down intention. We introduce AestheticNet, a novel cognitive-inspired AQA paradigm that integrates human-like visual cognition and semantic perception with a two-pathway architecture. The visual attention pathway, implemented as a gaze-aligned visual encoder (GAVE) pre-trained offline on eye-tracking data using resource-efficient contrast gaze alignment, models attention from human vision system. This pathway augments the semantic pathway, which uses a fixed semantic encoder such as CLIP, through cross-attention fusion. Visual attention provides a cognitive prior reflecting foreground/background structure, color cascade, brightness, and lighting, all of which are determinants of aesthetic perception beyond semantics. Experiments validated by hypothesis testing show a consistent improvement over the semantic-alone baselines, and demonstrate the gaze module as a model-agnostic corrector compatible with diverse AQA backbones, supporting the necessity and modularity of human-like visual cognition for AQA. Our code is available at https://github.com/keepgallop/AestheticNet

  • The Geometric Structure of Shared Neural Semantics: Evidence from Cross-Subject EEG-Text Alignment

    A fundamental challenge in cognitive science is whether the human brain employs a structured, relational representational scheme for semantics that is shared across individuals. While existing electroencephalography (EEG) studies have primarily focused on decoding or correlational analyses, it remains unclear whether sentence-level EEG activity encodes semantic structure at the level of representational geometry, beyond stimulus-specific decoding. This study investigates whether EEG signals recorded during natural reading can be aligned with representations capturing shared semantic structure and systematic correspondence with text-based distributional semantic models. Using the ZuCo 2.0 dataset, we learn an alignment between sentence-level EEG activity and a shared semantic embedding space and evaluate cross-subject generalization under a strict leave-one-subject-out protocol. Representational alignment is quantified using cross-subject semantic retrieval and Representational Similarity Analysis (RSA), linking neural activity to abstract semantic relations. Our results suggest that human EEG activity during language processing is consistent with a stable, relational organization of semantic information, consistent with structural properties of semantic organization captured in computational models, without implying representational equivalence.

  • Integrating Cognitive Strategies and Motor Dynamics to Analyze VR Learning Efficiency

    Learner performance in Virtual Reality (VR) often varies substantially across individuals, while conventional assess ments provide limited insight into the links between cogni tive strategies and motor execution. Focusing on the wire connection phase of a VR-based Wheatstone bridge experi ment, this study proposes a task-driven dual-path framework grounded in Perception-Action Coupling theory. The frame work combines PAC-informed visual sampling markers with implicit spatiotemporal dynamics from raw controller trajec tories. Using data from 98 participants, the hybrid model achieved 71.03% accuracy, showing a modest advantage over single-path baselines. Error-pattern analysis further suggests a possible cognitive-motor decoupling pattern, in which strate gic planning and motor execution may contribute differently to learning efficiency. This study provides a task-specific compu tational perspective for analyzing learner states in immersive procedural tasks.

  • Dynamics of creative exploration: A study of signal space constraints in music and communication

    Creative exploration is shaped by the functional pressures placed on signals. Yet we know little about how those pressures structure exploration, or how they connect to individual creativity and task constraints. Using a rhythm production task, we investigate how signal space exploration is affected by constraints in a communicative or musical act, as well as the presence of feedback. We quantify exploration using three measures: distance, novelty and syncopation, following our preregistered analysis. Rhythms produced in the music task were more novel and more syncopated than those produced in the communicative task, suggesting a communicative pressure towards more predictable structures. Feedback reduced novelty and syncopation, while increasing distance. These findings reveal how functional pressures shape signals and call for more research on how collective creativity and environmental constraints shape cultural acts across domains.

  • Cognitive Computation Beyond Optimality

    Cognitive science has often modeled cognition as an efficiency-driven process, assuming that mental computation aims at optimizing performance under resource constraints. While this view has been successful in many domains, it struggles to explain cognitive phenomena in which inefficiency, ambiguity, and non-convergence play a functional role, such as creative insight, exploratory problem solving, and conceptual restructuring. In this paper, we argue that these forms of cognitive inefficiency are not mere by-products of limited resources, but reflect a deeper computational organization of cognition. We propose a quantum-algorithmic perspective in which cognitive processes are understood as inherently incomplete computations, where superposition, interference, and delayed resolution contribute to flexibility and context sensitivity. From this viewpoint, inefficiency emerges as a structural feature rather than a deviation from optimality. By rethinking cognitive computation beyond output-driven optimization, we introduce a process-based notion of cognitive optimality and discuss its implications for cognitive modeling and computer science.

  • Hidden in Plain Sight: Expectation Congruency and Encoding Context Influence Object-Feature Memory

    An important component of episodic memory is the ability to bind together objects and their constituent features. Past research suggests that this binding process might be influenced by prior expectations about the relationship between objects and their features, as well as the contexts in which they are encoded. Yet it remains unclear how robust an effect these factors have in influencing memory for task-relevant and irrelevant features during recognition and recall. The current study explored how encoding context and congruency interact to shape memory. Recognition and recall were assessed following memorization and visual search for objects embedded in expectation-congruent or incongruent locations. While overall congruent locations were better remembered, incongruent locations were better recalled following visual search compared to memorization. Further, memory confidence may play an important role in individuals' ability to recall incongruent information. These findings highlight the contributions prior expectations and task demands make in shaping object-feature memory.

  • Inferring Arithmetic Skill from Speed and Accuracy

    People routinely infer others' competence under uncertainty, often relying on cues such as task difficulty and past accuracy. An emerging body of research suggests that people approximate Bayesian inference when doing so. We extend these results by testing whether people can infer others' numerical ability in a way that is consistent with a rational Bayesian model. In Study 1, we find that participants accurately predict the arithmetic performance of another individual from information about their past performance. Computational modeling shows that participants' inferences are better described by Bayesian processes than by plausible heuristics. Study 2 introduces a modified paradigm, in which participants are told about both past performance and time taken to solve problems. We find that, although participants are quite accurate in their predictions, they do not seem to take into account information about speed.

  • Does collaboration constrain insight problem-solving?

    This study examined whether collaboration helps or constrains insight problem solving using the Mutilated Checkerboard Problem, a classic task requiring representational change. One hundred participants solved the problem either individually, using think-aloud protocols (n = 50), or in dyads, using free conversation (n = 25 dyads). Dyads did not outperform individuals in solution rate or solution time, and solved the problem significantly less often than nominal dyads generated from individual performance. In unsolved cases, dyads also produced substantially more covering attempts than individuals, consistent with the possibility that partners became jointly fixated on a misleading strategy. Verbal-protocol analyses showed that successful solving was associated with producing more solution-relevant invariants and referring less often to invariant dimensions that were not relevant to the solution. Overall, collaboration appeared to constrain rather than enhance insight problem solving, suggesting that interaction may sometimes stabilize a shared but unproductive problem representation.

  • Modeling Bounded Rationality in Drug Shortage Pharmacists Using Attention-Guided Dynamic Decomposition

    Hospital pharmacists make high-stakes decisions to mitigate drug shortages under uncertainty, time pressure, and patient risk. Interviews revealed that pharmacists focus attention on a small subset of drugs, limiting cognitive effort to the most urgent cases. Motivated by these findings, we formalize a bounded-rational, attention-guided decision framework that dynamically decomposes drugs into a subset for high-cost reasoning and a complementary subset for low-cost monitoring. We develop two agents: an Expert Agent that applies attention weights derived from pharmacist interviews, and a Learner Agent that adapts attention allocation over time through experience. Across simulated scenarios spanning short to long horizons, we show that attention-guided planning supports stable decision-making without complete state reasoning. These results suggest that a primary decision is not what action to take, but where to allocate cognitive effort, and that attention-guided, satisficing strategies can reduce problem complexity while maintaining stable performance.

  • Local–Global Association Between Pitch and Chord Identification Strategies

    Strategies for identifying pitch and chords can be understood in relation to a shared local–global processing dimension. Absolute pitch (AP) and pitch-element identification tend to reflect more local strategies, whereas relative pitch (RP) and harmonic-function identification tend to rely more strongly on global strategies. The present study examined whether individual internal dominance in pitch identification strategies is associated with corresponding dominance in chord identification strategies. Twenty-one professionally trained musicians completed pitch identification tasks (AP and RP) and chord identification tasks requiring either pitch-element or harmonic-function judgments. Individual dominance was indexed using residual-based measures to dissociate identification bias from overall ability. Dominance in local versus global pitch identification was significantly associated with corresponding dominance in chord identification across accuracy and reaction time. Analyses based on raw scores showed consistent directional trends, though most did not reach statistical significance. Overall, the findings suggest that pitch- and chord-level identification strategies interact along a shared local–global dimension, highlighting the role of individual processing dominance in musical identification.

  • Modeling Human-Like Color Naming Behavior in Context

    Modeling the emergence of human-like lexicons in computational systems has advanced through the use of interacting neural agents, which simulate both learning and communicative pressures. The NeLLCom-Lex framework (Zhang et al., 2025) allows neural agents to develop pragmatic color naming behavior and human-like lexicons through supervised learning (SL) from human data and reinforcement learning (RL) in referential games. Despite these successes, the lexicons that emerge diverge systematically from human color categories, producing highly non-convex regions in color space, which contrast with the convexity typical of human categories. To address this, we introduce two factors, upsampling rare color terms during SL and multi-listener RL interactions, and adopt a convexity measure to quantify geometric coherence. We find that upsampling improves lexical diversity and system-level informativeness of the color lexicon, while many-listener setups promote more convex color categories. The combination of moderate upsampling and multiple listeners produces lexicons most similar to human systems.

  • Mouse tracking for syntactic and semantic interference

    This study investigates the effectiveness of 'mouse tracking for reading' (MoTR), a novel incremental processing paradigm developed by Wilcox et al. (2024). MoTR adapts the eye-tracking procedure for online implementation. The present experiment is a replication of Mertzen et al.'s (2023) English-language eye-tracking task. Compared to Wilcox et al.'s original study, this work replicated not just syntactic but also semantic effects on reading, tested longer and more complex sentences, and directly compared the behaviour of trackpad and regular mouse users. Evidence of semantic interference was observed, with limited evidence of syntactic interference. The replication of these results adds to the evidence that MoTR can usefully supplement or even replace eye tracking.

  • Making room for new words

    When words are borrowed into a language, their meanings often narrow. For example, 'salsa' can mean any sauce in Spanish but refers only to tomato-pepper sauces in English. What principles of individuals' learning shape this collective semantic narrowing? We propose that semantic narrowing reflects an accumulation of pragmatic inferences about word meaning. When listeners encounter a novel word for an already-nameable concept, they infer that it has a more specific and atypical meaning than existing terms. Across four experiments, we tested these inferences among adults and preschoolers. Participants saw arrays of familiar category exemplars (e.g., leaves of varying subtypes) and were prompted with familiar words ("leaves") or novel words ("nefts"). Both adults and children selected more specific subtypes and more atypical exemplars when prompted with novel words. Learners carve out distinctive meanings for new words, a pragmatic inference that may shape meaning from early childhood to long-range language change.

  • Naka-SAM: A Cognition-Inspired Framework with Nakagami Prior for Ultrasound Segmentation

    Radiologists often implicitly rely on tissue backscattering characteristics to interpret noisy ultrasound (US) images under challenging imaging conditions. However, existing deep learning-based segmentation models rarely incorporate such intrinsic physical statistics into representation learning. To address this limitation, we propose Naka-SAM, a physics-informed ultrasound segmentation framework built upon the Segment Anything Model (SAM). Specifically, we design a Physical Prior Generator (PPG) to model ultrasound backscattering statistics through learnable Nakagami distribution estimation, where the shape parameter \(m\) characterizes local scattering concentration and the scale parameter \(\Omega\) reflects backscattered energy variations. These physics-informed priors provide domain-invariant tissue representations associated with underlying microstructural properties. Furthermore, a dual-stage gated fusion strategy is introduced to progressively integrate physical priors with both low-level structural features and high-level semantic representations, thereby improving robustness under noisy and low-contrast imaging conditions. Extensive experiments across seven ultrasound datasets demonstrate that Naka-SAM consistently outperforms both conventional segmentation methods and recent SAM-adapted frameworks, while exhibiting strong cross-domain generalization capability across heterogeneous ultrasound imaging scenarios.

  • Environmental Statistics Shape Learning Dynamics in Meta-Reinforcement Learning

    Adverse environments have been proposed to influence multiple aspects of cognition, including learning and adaptation. However, causal tests of long-term effects of adversity on human cognition are extremely difficult. Moreover, different types of adversity often co-occur in real life, making it challenging to disentangle their individual effects. Computational models provide a way to independently manipulate specific environmental statistics and examine their causal impact on learning. Here, we integrate meta-learning with reinforcement learning to investigate the mechanisms by which environmental statistics shape learning processes. Specifically, we examine how two core environmental dimensions of adversity, volatility and controllability, shape trajectories of learning rates in deep neural networks. We find that both high volatility and low controllability reduce the maximum learning rate, but only controllability affects the initial slope of learning, leading to a flattening of early learning dynamics. Notably, deeper layers of the model are more strongly influenced by environmental factors than earlier layers, leading to reduced hierarchical differentiation in learning rates under conditions of high volatility and low controllability. These results suggest that a meta–reinforcement learning framework may provide a simplified yet useful approach for gaining mechanistic insight into how distinct types of adverse environments affect learning.

  • Filtering the Evidence: How Communicative Goals Shape the Balance of Opinion About Health Claims on Social Media

    Increasingly, people encounter health-related claims made by others on social media. When faced with conflicting information, an important question is how people decide which information to share. Selective sharing (e.g., passing on only posts that support a particular claim) may distort others' perceptions of the balance of opinion. Across two experiments, we examined how communicative goals shape information sharing by asking participants to choose sets of posts to share with different goals in mind. In Experiment 1, we found distinct sharing behaviours across goals, and that when given no explicit goal, people tended to share as if advocating for their preferred stance rather than conveying the balance of opinion in a more nuanced way. Experiment 2 replicated these results even when the balance of opinion opposed people's own beliefs. This research highlights how biases in information sampling may systematically distort social learning.

  • CausalMoE: Enhancing Expert Specialization and Interpretability through Cognitively-Inspired Causal Routing

    Mixture-of-Experts (MoE) architectures enable scalable language modeling via sparse, input-dependent computation, but suffer from weak expert specialization and limited routing interpretability. From a cognitive perspective, this contrasts with theories of modular cognition that emphasize selective recruitment of functionally specialized systems. We propose CausalMoE, a causality-aware MoE framework that aligns expert activation with causal structure in language. CausalMoE introduces a specialized Causal Expert and a Causal Routing mechanism that identifies tokens along causally salient dependency structures inferred from attention patterns. The Causal Expert is specialized in processing sequences with strong causal and logical dependencies. To further encourage functional differentiation, we introduce a Causal Auxiliary Loss that reduces representational redundancy between experts. Integrated into canonical MoE architectures, including Switch Transformer and GLaM, CausalMoE achieves an average relative perplexity reduction of 18.42% on logic-intensive code datasets and a mean improvement of 6.83% on causal reasoning benchmarks, demonstrating structure-sensitive and interpretable computation allocation.

  • TP-DID: Temporal Prediction of Driver Interpretation Distributions

    Driving assistance systems require not only predicting pedestrian-side crossing intent, but also estimating how drivers interpret that intent from ego-view evidence. Existing pedestrian intent prediction methods often rely on one-hot labels or binary crossing probabilities, which cannot capture observer variation or intermediate interpretation states. We introduce \emph{Temporal Prediction of Driver-side Interpretation Distributions} (TP-DID), a causal frame-level task for predicting empirical distributions of driver-side interpretations towards pedestrian intent. To capture temporal distribution changes, we propose \textsc{Input-Modulated Mamba} (IM-Mamba), a causal state-space model that reweights current visual evidence before temporal state prediction. Experiments on PSI show that IM-Mamba improves distribution prediction, especially for observer variation, awareness-related states, and temporal changes.

  • Modeling and Applying Tree-Like Path Planning of Multiple Ant Colonies

    Ant colonies exemplify rigorous and efficient distributed cognition, providing profound inspirations for understanding collective intelligence. Specifically, inter-colony competitions and foraging strategies of ants offer insights into how biological systems search resources. Although it is well-known that ants naturally create tree-like structures in foraging paths, little is investigated regarding how multiple colonies optimize these structures simultaneously under competitions. Hence, we propose MT-ACO, a Dual-Colony Competition Model driven by a Pheromone Repulsion Mechanism. In this model, agents treat the pheromone trails of competitors as dynamic environmental constraints, simulating the biological principle of Competitive Exclusion. Simulations reveal that this simple rule leads to near-optimal Edge-Disjoint Directed Trees. This emergent topology maximizes resource coverage while minimizing conflict, providing a decentralized heuristic for the NP-hard Minimal Directed Edge-Disjoint Double Tree problem. Our findings demonstrate that avoidance functions offering a computational framework to understand how swarms resolve the tragedy of the commons through adaptive niche differentiation.

  • CogMAME: A Cognitive-inspired Meta-learning Adaptive Model for Multi-modal Entity-relation Extraction

    Multi-modal Named Entity Recognition (MNER) and Multi-modal Relation Extraction (MRE) are critical for knowledge extraction. Existing methods often suffer from cross-modal confusion due to semantic misalignment and data imbalance. Inspired by predictive processing theory, i.e., the brain continuously refines cognition through active prediction and error minimization, we propose a Cognitive-inspired Meta-learning Adaptive model for Multi-modal Entity-relation extraction named CogMAME. It consists of five modules: (i) multi-grained representation learning, establishing a unified semantic space; (ii) class-adaptive feature augmentation, reinforcing signals for rare categories; (iii) meta weight hypernet, dynamically calibrating modality contributions; (iv) self-adversarial perturbation, improving stability via controlled noise; (v) multi-modal information extraction, refining outputs in a joint feature space. Extensive experimental results show that our model can achieve human-like extraction capabilities and obtain state-of-the-art performance on both MNER and MRE tasks, verifying its superiority and effectiveness.

  • Names as Social Signals: Quantifying Name-Based Appearance and Socioeconomic Biases in Text-to-Image Models

    Human perception often relies on names when forming first impressions of unfamiliar people. As text-to-image (T2I) models are increasingly used in everyday content creation, names may also influence what these models generate under otherwise neutral prompts. We study name-based bias in mainstream T2I models along two dimensions: appearance patterns and socioeconomic status patterns. We curate name sets from China, America, Britain and France, generate over 6,280 images, and use automated VLM-based scoring to assess age appearance and styling, as well as occupation and living-environment attributes in work and home scenes. Across models, changing only the name yields reliable shifts in these attributes. The strength and consistency of the pattern vary by model, with Nano Banana and ChatGPT-5 showing more uniform name-linked defaults. We propose a quantitative framework to identify and compare name-based biases, helping advance fairer AI systems.

  • How Real-Time Signaling Improves Ad Hoc Collaboration

    Many collaborative tasks require people to coordinate interdependent actions under time constraints, often without explicit communication. In such ad hoc settings, observing others' emerging intentions may help groups rapidly align their actions. While information cascades are often studied as sources of bias in individual decision-making, they may play a constructive role in coordination by facilitating fast convergence on shared plans. We investigated this possibility in a multi-agent coordination task in which teams of three repeatedly solved a time-constrained block-moving problem (N = 471). We manipulated whether players could observe others' intentions in real time or only observe the outcome of each round. Teams with real-time signaling were more likely to complete the task. Within rounds, early votes were frequently followed by teammates, and these actions typically advanced task progress along efficient paths. Together, these findings suggest that access to real-time intention signals supports efficient ad hoc coordination by reducing decision uncertainty and stabilizing shared action plans.

  • Online Games for Perceptual Learning: Assessing Acquired Affective Connotations of Sounds

    Understanding how perceptual judgements develop requires methods that capture both existing biases and their susceptibility to short-term experience. We present an online game-based paradigm designed to investigate how affective responses to sounds can be shaped by exposure. 48 participants first completed an affective priming task to establish baseline associations. They then played a simple game in which specific sounds were systematically paired with positive or negative in-game events, providing a controlled but engaging learning environment. A second priming task assessed post-exposure changes. Initial findings indicate that explicit self-reports show consistent shifts in valence judgements aligned with the in-game contingencies, whereas the results on priming are weaker and reach significance only partially. Improvements to the design and implementation are discussed. Overall, the results suggest that the paradigm can capture the experience-dependent modulation of affective perception of sounds while also highlighting the distinction between explicit and implicit measures. The approach offers a general framework for studying short-term plasticity in perceptual cognition.

  • Effects of surprising rewards on pattern separation

    Surprising feedback alters memory of events. Reward prediction errors (RPEs) signal when feedback violates expectations. Although RPEs are known to influence memory, their role in pattern separation---a fundamental computation supporting discriminated episodic representations---remains unclear. We ask whether surprising outcomes strengthen memory encoding and if welcome (positive RPE) versus unwelcome (negative RPE) surprises differentially affect memory discrimination. We designed a variant of an established task developed to probe pattern separation, the Mnemonic Discrimination Task, that introduces post-trial reward and punishment feedback during encoding. We find that while surprising rewards enhance memory discrimination overall, trials with positive RPEs lead to better discrimination of similar lure items, but trials with negative RPEs lead to better recognition of repeated target items. Generalizing expectations to semantically similar stimuli further benefits discrimination. These findings suggest that surprising feedback enhances memory discrimination, with effects depending on the type of feedback and similarity structure of experiences.

  • A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities

    Imbuing Large Language Models (LLMs) with specific personas is prevalent for tailoring interaction styles, yet the impact on underlying cognitive capabilities remains unexplored. We employ the Neuron-based Personality Trait Induction (NPTI) framework to induce Big Five personality traits in LLMs and evaluate performance across six cognitive benchmarks. Our findings reveal that persona induction produces stable, reproducible shifts in cognitive task performance beyond surface-level stylistic changes. These effects exhibit strong task dependence: certain personalities yield consistent gains on instruction-following, while others impair complex reasoning. Effect magnitude varies systematically by trait dimension, with Openness and Extraversion exerting the most robust influence. Furthermore, LLM effects show 73.68% directional consistency with human personality-cognition relationships. Capitalizing on these regularities, we propose Dynamic Persona Routing (DPR), a lightweight query-adaptive strategy that outperforms the best static persona without additional training. Code and data are available at: https://github.com/cjia7/DPR.

  • Framing a Biogenic and Predictive Perspective on Attention

    The study of attention faces significant conceptual ambiguities rooted in persistent reductive and anthropocentric assumptions. To address this, we advocate for a non-reductive, biogenic framework, reconceptualizing attention as an embodied, evolutionarily conserved process essential for maintaining thermodynamic equilibrium. We operationalize this perspective using Active Inference, defining attention as the selective allocation of precision (gain) to prediction errors. This mechanism tunes the organism's sensitivity to salient information, serving to minimize variational free energy (uncertainty). Precision modulation offers a compelling interpretation of selective spatial attention across psychology and neuroscience while aligning with established philosophical theories. Ultimately, this biogenic, predictive account may provide the necessary formalisms to resolve current conceptual impasses and invigorate the understanding of attention within the context of embodied cognition.

  • Training Students to Recognize Comparisons in Science Texts

    Science texts frequently use comparisons to communicate concepts, based on research that has established comparison's benefits in learning. However, applying these findings to student learning relies on a critical premise yet to be examined—namely, that students recognize comparisons in their readings. The current work aims to examine (1) students' ability to recognize comparisons in science texts, and (2) whether this ability can be trained. In two experiments across two different institutions, college students were asked to mark all the comparisons in expository science texts, with some receiving instructions and practice on identifying comparisons. Without training, students missed many of the intended comparisons. However, targeted training reliably improved comparison detection. The findings showed that detecting comparisons is a bottleneck in learning from science texts, but that this ability can be improved with brief training.

  • D-Cog: Distilling Human Cognitive Representations into Visual Prompts for Few-Shot SAR Recognition

    Synthetic Aperture Radar (SAR) recognition is frequently impeded by severe speckle noise and data scarcity. In contrast, human vision excels in these challenging conditions due to robust top-down cognitive modulation. This work investigates whether human cognitive representations can be transferred to SAR recognition models. We introduce the first paired fMRI–SAR dataset designed to capture the cognitive neural responses of human viewing SAR images. Building upon this dataset, we propose a cognitive fusion–distillation framework. This framework integrates fMRI-derived cognitive representations into a teacher network and distills the cognitive capability into a lightweight student model via prompt tuning. Extensive experiments demonstrate consistent performance improvements in few-shot, cross-domain, and out-of-distribution conditions. Furthermore, Representational similarity analysis reveals that cognition-guided models learn representation patterns distinct from purely data-driven models. These results suggest that neuro-cognitive signals can serve as transferable inductive biases to enhance robustness and generalization in specialized vision tasks.

  • Learning Vowel Harmony from Speech: What Emerges, What Requires Additional Structure

    Vowel harmony is a non-local phonological process in which vowels within a word share a common feature value. We examine which aspects of this process are learnable from speech alone, using Assamese Advanced Tongue Root (ATR) harmony as a test case. We use a generative (fiwGAN) and a predictive model (wav2vec2) as comparative probes of what different speech-learning objectives make accessible. Both models learn ATR-related phonetic distinctions: fiwGAN latent variables control F1 differences, and wav2vec2 separates ATR categories with _81% accuracy. Wav2vec2 also learns global feature agreement (98% accuracy) and recognizes the blocker /A/ (27.9% agreement drop in blocked contexts). However, neither model shows strong evidence for regressive spreading (d = 0.12) or iterative harmony across longer words. These results suggest that raw speech learning exposes the acoustic and co-occurrence structures, while directional phonological generalization may require morphology, lexical alternations, or temporal and architectural biases.

  • A computational operationalisation of competing maturational theories of syntactic development via statistical grammar induction

    This paper is concerned with what intermediate syntactic categories children acquire during first language development, and in what order. Maturational theories make different predictions. Bottom-up accounts (Growing) propose that lexical and inflectional structure emerges first, while inward accounts (Inward) predict early access to discourse-related categories. We computationally operationalise these hypotheses of staged syntactic emergence using statistical grammar induction, asking what each proposed ordering makes learnable when input and learning algorithm are held constant. Our framework makes category acquisition explicit and allows us to explore how different maturational orderings shape the structure that can be learned under identical conditions. Based on this operationalisation, the Growing account significantly outperforms the Inward account across three evaluation metrics.

  • Deep Inverse Reinforcement Learning for Semantic Fluency: Uncovering the Reward Structure of Cognitive Search

    Semantic fluency tasks reveal the temporal structure of memory retrieval through "clustering and switching'' behavior. Existing models rely on static similarity (e.g., cosine similarity in Word2Vec) or heuristic switching rules, failing to capture this sequential decision-making process. Recently, fine-tuned semantic vectors to better fit transitions, but assumed a single context-independent weighting throughout the task. We propose Deep Inverse Reinforcement Learning for Semantic Fluency (DIRL-SF), modeling the participant as an agent navigating a semantic graph to maximize an unknown reward function. A Transformer encoder captures retrieval history, while a dynamic attention mechanism re-weights semantic dimensions per step, accounting for similarity, novelty, and effort in a state-dependent manner. On a 27-category dataset, DIRL-SF significantly outperforms static and LSTM baselines, providing a theoretically grounded tool for comparing search strategies across populations.

  • Human Hippocampal Replay as Search

    Human-like systematic generalization remains a challenge for artificial systems. Recent work suggests hippocampal replay as a mechanism for online planning that facilitates generalization in novel, dynamic, and complex environments. Recent models, based on spatial tasks and rodents, present a serial notion of planning. In contrast, recent work on humans finds replay sequences traversing multiple trajectories at the same time on a compositional generalization task. Such a possibility for parallel execution does not interface well with serial planning. Here, we present a novel model of replay as planning with parallel search capabilities, extending an existing serial model of replay. Our model displays human-level generalization faster than the serial baselines. Furthermore, we find that models actively using parallel search produce internal dynamics qualitatively resembling neural replay sequences in humans performing the task. Our findings present search as a powerful general computation leveraged during replay for quickly learning good policies from little data.

  • Investigating the Manifestation of naïve Physics at the Embodied Level Through Motor and Judgment Tasks in Virtual Reality

    People often hold the intuition that heavier objects fall faster, yet can still accurately catch objects of different masses. Such coexistence of naïve physics and the internal model of Earth's gravity is well documented, but how these are weighted when they conflict remains unclear. We designed a Virtual Reality ball-catching experiment to examine how acceleration manipulations and the presence of a second, visible ball that enabled visual comparison modulate motor catching and perceptual judgment. We found that, when participants attempted to catch a single occluded ball, both motor catching and naturalness judgments showed a partial reliance on the internal model and were especially successful in distinguishing slow accelerations. Introducing a second ball with manipulated acceleration shifted judgments towards naïve physics expectations, while motor catching accuracy remained relatively robust. Our findings highlight how naïve physics can manifest at the perceptual level and inform motor-level embodied learning designs to initiate conceptual changes.

  • Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents

    Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an extended 7-point ICAP framework based on the Interactive, Constructive, Active, and Passive modes to characterize variation in cognitive engagement during collaborative dialogue. Engagement was coded by trained human annotators and compared with large language model (LLM)–based labeling approaches, including in-context learning (ICL), zero-shot prompting, and self-reflective agents. Interrater reliability among human annotators was robust across framework refinement stages (_ = 0.906–0.998), higher than the moderate agreement observed for ICL-based annotation (_ = 0.541–0.609). The human-refined framework improved agreement among human annotators (–ñ_ = 0.10), but produced only modest gains for ICL-based LLMs (–ñ_ < 0.04). Agent-refined frameworks improved cross-model agreement but remained below the human-refined framework. These findings highlight the promise of agent-based approaches and the importance of continued interaction between theory-guided human annotation and LLM-based methods in future work.

  • The Cost of Coordination: Why Multi-Agent LLMs Explore More But Discover Less

    Large language models (LLMs) are increasingly deployed in multi-agent teams (MAS), yet recent works have shown that unstructured multi-agent collaboration often degrades performance in noninsight problems. We examined whether previous findings can be generalized to an insight problem solving paradigm and identified process-level signatures that may explain the performance gap. We compared performance between solo versus two-agent ChatGPT-4o teams to solve situation puzzles in a 2 (Agent Configuration) _ 2 (Solving time) experiment (N = 82). Teams were significantly less accurate than solo agents despite asking questions across broad categories. However, question quality was equivalent between conditions. A multilevel mediation analysis controlling for time decomposed the coordination cost into two components: a protocol failure component (47%), in which inter-agent discussion reduced host-directed questioning, and a residual coordination cost (53%) consistent with impaired convergent integration that persisted after controlling for question volume and quality. Doubling the time budget proportionally increased question volume in both conditions but preserved the team to solo ratio, indicating that the cost is structural rather than driven by limited time. These findings extend multi-agent coordination costs to insight problem solving and suggest that unstructured multi-agent interaction supports the divergent phase but disrupts the convergent phase.

  • Spotlight and handover: Register socialization in Korean child-directed speech

    How do children learn to use language as a social tool flexibly calibrated to the identities and relationships of their interlocutors? While developmental research has benchmarked when children recognize social registers (e.g., matching casual vs. formal registers to specific conversational partners), the mechanisms by which they acquire these context-dependent mappings remain unclear. We leverage the high-resolution morphological signal of Korean honorifics to investigate the dynamic processes underlying early register socialization. Through analysis of cross-sectional and longitudinal corpora from CHILDES, we demonstrate that caregivers spotlight the socio-pragmatic functions of honorifics in child-directed speech (CDS) by selectively clustering them within specific communicative acts. Furthermore, we reveal a pragmatic handover in which caregiver scaffolding systematically fades as the child's independent sociolinguistic competence emerges. Together, these findings suggest that register learning is supported by an adaptive tutoring process—one that facilitates the child's transition from a passive linguistic recipient to a register-flexible social agent.

  • CogEvolution: A Human-like Generative Educational Agent to Simulate Student's Cognitive Evolution

    Generative Agents, owing to their precise modeling and simulation capabilities of human behavior, have become a pivotal tool in the field of Artificial Intelligence in Education (AIEd) for uncovering complex cognitive processes of learners. However, existing educational agents predominantly rely on static personas to simulate student learning behaviors, neglecting the decisive role of deep cognitive capabilities in learning outcomes during practice interactions. Furthermore, they struggle to characterize the dynamic fluidity of knowledge internalization, transfer, and cognitive state transitions. To overcome this bottleneck, this paper proposes a human-like educational agent capable of simulating student cognitive evolution—CogEvolution. Specifically, we first construct a cognitive depth perceptron based on the Interactive, Constructive, Active, Passive (ICAP) taxonomy from cognitive psychology, achieving precise quantification of learner cognitive engagement. Subsequently, we propose a memory retrieval method based on Item Response Theory (IRT) to simulate the connection and assimilation of new and prior knowledge. Finally, we design a dynamic cognitive update mechanism based on evolutionary algorithms to simulate the real-time integration of student learning behaviors and cognitive evolution processes. Comprehensive evaluations demonstrate that CogEvolution not only significantly outperforms baseline models in behavioral fidelity and learning curve fitting but also uniquely reproduces plausible and robust cognitive evolutionary paths consistent with educational psychology expectations, providing a novel paradigm for constructing highly interpretable educational agents.

  • Pupillary Indices of Cognitive Workload in (In-)Efficient Reference Production

    Speakers' use of redundant information in referential communication has been the subject of much debate, as it seemingly defies Grice's Maxim of Quantity (Grice, 1975). While some research argues for an egocentric view on referential redundancy (i.e., that it results from a concern to ease production effort), other work supports an audience-design view (i.e., that speakers strive for communicative efficiency). This research has mainly leveraged listeners' comprehension as an indirect index of production strategy, while it is yet unclear whether audience-design indeed incurs increased cognitive effort for the speaker. In this study, we used pupillometry to directly assess the cognitive load associated with different referential production strategies, finding evidence for modulated effort based on speakers' engagement in reasoning about the efficiency of their utterances.

  • Breaking Control - How confusion of control can help understand the Sense of Agency

    The Sense of Agency (SoA) is a crucial component of robust human motor control and learning. To study the SoA, existing work has largely focussed on retrospective ratings about the level of perceived control over a single visible object. Here we propose a new design, using queries during a trial, for assessing participants' attribution of control to one of multiple acting objects performing the same obstacle avoidance task. Correlation analyses between the controlled object and the object to which Agency was (wrongly) attributed showed that Agency confusion is largely driven by similarities in the dynamics of participants' control input and the confounding object's behaviour. Visual attention was also more attracted to the object believed to be under the participants' control. These results are consistent with the comparator and cue-integration models of Agency and the presented approach opens up novel avenues to explore the dynamic processes underlying the SoA.

  • From Prediction to Justification: Aligning Sentiment Reasoning with Human Rationale via Reinforcement Learning

    Although Aspect-based Sentiment Analysis (ABSA) systems attain high sentiment polarity detection accuracy, they act as "black boxes" without the explicit reasoning of human affective cognition, where humans form causal explanations for sentiment judgments. To address this, we propose ABSA-R1, a large language model framework mimicking the "reason-before-predict" cognitive process. Using reinforcement learning (RL), it generates natural language justifications to underpin sentiment predictions. We design a Cognition-Aligned Reward Model to ensure consistency between reasoning paths and emotional labels, and a performance-driven rejection sampling strategy (inspired by metacognitive monitoring) targeting hard cases with uncertain/inconsistent internal reasoning. Experiments on four benchmarks show this explicit reasoning boosts model interpretability and outperforms non-reasoning baselines in sentiment classification and triplet extraction.

  • When Emotional Intensity Is High, Feelings Start to Look Like Facts

    Moral conflicts frequently involve ambiguous evidence, creating opportunities for emotion to shape what people find credible. Across three studies, we examine how emotional expressions influence the perceived credibility and usability of moral evidence. In Study 1, we analyze posts from the Reddit forum r/AmITheAsshole, finding that readers expressed more trust in posts that were emotionally intense. In Study 2, we use a social media-style post to manipulate emotional intensity to test perceived credibility of subjective and objective evidence. High emotional intensity of expressions narrowed the perceived difference in credibility between subjective and objective evidence. In Study 3, we find that emotional intensity also enhances evidence usability—higher perceived causality and counterfactual necessity are attributed to a person who delivers evidence with high emotional intensity (vs. low). This research highlights how the intensity of emotional expressions influences both the evaluation and use of moral evidence.

  • The Integrated Information Bottleneck: A theoretically-grounded model unifying category- and system-level structure in communicative efficiency

    Much research has shown that lexical semantic systems are shaped by a drive for efficient communication, which yields desirable structure at the level of individual lexical categories (i.e., the semantic coherence of words). Less attention has been given to what we call semantic systematicity: structural regularity that exists across lexical categories, where different words make the same semantic distinctions. We develop the first model of the efficiency trade-off that captures both category- and system-level regularity using the same domain-general information-theoretic underpinning. We also propose a stricter evaluation method for comparing alternative formalizations of the efficiency trade-off. Our model that integrates the pressures for category-level structure and semantic systematicity explains crosslinguistic variation better than previous approaches in two semantic domains, shedding new light on how a pressure for efficiency shapes the structure of language.

  • Towards a Cognitive Pattern Language for Tonal Music Representation

    We present a symbolic, rule-based framework for modeling tonal music composition as the construction and transformation of structured cognitive programs. Building on the Language of Thought hypothesis, our approach treats musical scores as externally observable artifacts that encode hierarchical relations, melodic patterns, and harmonic scaffolds in a reusable symbolic system. Using Minimum Description Length (MDL) principles, we show that tonal abstractions support compression: themes function as shared models, allowing variations to be represented with fewer parameters and transformations, reflecting cognitively plausible reuse and systematicity. We demonstrate the framework on Paganini's Caprice No. 24 and Steve Reich's Clapping Music, showing how surface diversity emerges from a compact set of symbolic operations applied over inherited structures. Unlike data-driven neural models, our method preserves interpretable, inspectable representations aligned with musicological analysis. This work provides a proof-of-concept that tonal music can be formalized as an executable symbolic system, offering both a computational model of composition and a window into the cognitive structures underlying musical creativity.

  • Processing and Acceptability of Singular They: Evidence from Eye Tracking

    Singular they has a long history in English. Generic uses of singular they refer to nonspecific individuals, whereas referential uses refer to specific persons. This study examined acceptability and processing of singular they and the role of individual differences. Fifty-one native English speakers read sentences containing gendered pronouns or singular they in generic and definite contexts while their eye movements were recorded. Participants also completed measures of gender role attitudes, prescriptive grammar attitudes, attitudes toward singular they, and the Bem Sex Role Inventory. Singular they was rated as highly acceptable in both contexts, with generic uses preferred over gendered pronouns. Eye-tracking data revealed no overall processing cost for singular they. These findings indicate that singular they is largely standardized in Canadian English, and any processing cost is brief and influenced by evaluative monitoring rather than comprehension difficulty.

  • An Experimental Study of the Evolution of Hierarchical Category Systems

    Hierarchical category systems evolve in response to a changing environment encountered over time. To better understand how such systems develop, we designed a task in which participants sequentially incorporated novel stimuli into an existing taxonomy. We manipulated presentation order, similarity to existing categories, and category system depth. All three factors shaped the hierarchical systems that emerged, and the likelihood of creating a new high-level category decreased as system depth increased. Our findings show that hierarchical category systems depend on the sequence in which items are encountered, and that sequence effects can result in hierarchical category systems that are skewed representations of the world.

  • Narrative Skills in Bilingual Children: The Impact of Shared and Language-specific Skills

    Despite increasing research on bilingual narratives, it remains unclear whether macrostructure is shared across bilinguals' two languages, how microstructure supports macrostructure within each language, and whether language dominance influences these relations. This study examined cross-language macrostructure, microstructure–macrostructure associations, and the moderating role of language dominance. Participants were 64 Grade 1 (Mage = 7;2 years) Greek–German bilingual children who completed a picture-based narrative task in both languages. Macrostructure was assessed using story structure, while microstructure was measured through total number of words (TNW), lexical diversity (VOC-D), mean length of utterance (MLU), and clausal density (CD). Language dominance was operationalized as relative vocabulary proficiency in Greek and German. Results showed a partial cross-language macrostructure association and significant contributions of microstructure to macrostructure in both languages, with no moderation effects of dominance. TNW emerged as the strongest and most consistent predictor. Findings highlight shared cognitive and language-specific factors supporting narrative macrostructure.

  • Inducing an Incremental Grammar for Production in First Language Acquisition

    This paper investigates how children may acquire an incremental grammar that supports word-by-word speech production via computational modeling. We formalize language acquisition as the learning of a probabilistic, bidirectional mapping between utterances and intended meanings. Our first contribution is the design of a precise incremental grammar grounded in Synchronous Hyperedge Replacement Grammar. The grammar constrains syntactico-semantic analyses to be strictly left-branching, reflecting the incremental nature of human language comprehension. We then conduct statistical production-oriented grammar induction over naturalist data. For the meaning-to-text generation task, the incremental grammar yields slightly lower semantic constituent recognition accuracy but unexpectedly higher language generation quality, compared to a hierarchical grammar. It also exhibits a small and consistent accuracy gap of semantic constituent recognition relative to a supervised oracle, an effect not observed for the hierarchical grammar. Together, these results suggest a new method for comparing alternative grammatical architectures in silico.

  • The Agency Gap: Testing Human–Human and Human–AI Collaboration Using a Minimal Coordination Game

    Humans coordinate by forming and committing to shared intentions—a form of shared agency that is largely absent in standard reinforcement learning agents and remains unclear in large language models, which function more like tools than autonomous agents. We compared human–human and human–AI collaborations using a non-linguistic sequential coordination game designed to disentangle successful collaboration outcomes from the underlying coordination mechanisms, such as commitment to an established joint goal despite the emergence of conflicting alternatives. Although adults and 7-year-old children were often able to adapt to different AI partners and achieve successful outcomes, human–human collaboration was marked by uniquely high levels of commitment and coordination efficiency, both of which were reduced in human–AI interactions. Five-year-old children showed greater difficulty collaborating with AI partners that functioned merely as asocial tools. Our findings suggest that human-compatible AI require mechanisms beyond outcome optimization to support efficient and stable collaboration.

  • Three-year-olds Ability to Reason using Disjunctive Syllogism

    Logical reasoning is fundamental for human cognition, yet its developmental and evolutionary origins remain unknown. In two experiments, we investigated whether 3-year-old children can reason using disjunctive syllogism. Children (N = 73) participated in a gumball game where they helped an experimenter obtain a desired gumball. To succeed, they needed to use disjunctive syllogism to determine which of two gumball machines must or might produce the desired gumball. Results suggest a merging ability in 3-year-olds.

  • Context-dependence of dispositional ascriptions

    Dispositional concepts like fragile or vulnerable are crucial for prediction and action planning, yet no established cognitive theory explains their mental representation. We propose a relevance theory of dispositions, positing that the function of dispositional concepts is to encode an object's disposition-characteristic behavior across relevant situations. We formalize this theory as an idealized cognitive mechanism, in which dispositional strength is calculated as a sum of response intensities across situations weighted by relevance. We test this theory in three experiments examining judgments of vulnerability, fragility, flammability, and solubility. Their results show that a disposition is judged stronger when it would manifest in actual rather than merely possible situations (Experiment 1), in situations an agent intends to realize (Experiment 2), and in high-stakes situations (Experiment 3).

  • The Role of Category Confusability in Classification Versus Observation Learning

    Concept learning can be studied through training that involves either classification or observation. In classification, learners actively make a judgment about category membership of presented stimuli and receive feedback; in observation, learners study a labeled stimulus. Extant work shows contradictory findings, as some studies show a benefit of classification over observation, whereas others show the opposite pattern. Recent work posits that these findings might be explained by category confusability. We present an experiment that investigates this question by directly manipulating category confusability (within-subjects), wherein subjects either learned through classification or observation. All subjects learned three different pairs of natural categories that shared low, moderate, or high feature similarity. Subjects were presented a single stimulus on each trial; every fourth trial, they made an endorsement judgment. Our findings suggest that training both with (classification) and without (observation) error are equally beneficial for learning natural categories, challenging several prominent category learning models.

  • Vocabulary knowledge and social contact in L2-accented speech perception: Evidence from low-exposure listeners

    Laboratory training improves L2-accented speech recognition, but it is unclear whether everyday exposure produces similar benefits. Usage-based theories propose that lifetime exposure to diverse accents facilitates recognition, whereas resource-dependent models argue that recognition depends primarily on linguistic processing resources. To evaluate these competing hypotheses, we tested 78 young monolingual English adults on their word recognition of Italian-accented English in noise. A multidimensional questionnaire was developed to measure the frequency, duration, amount, and diversity of participants' real-world accent exposure. Receptive vocabulary strongly predicted recognition accuracy (p < .001), which remained significant after controlling for social exposure (p = .003) in a sample with low exposure to L2-accented speech. Composite measures of social interaction did not independently predict performance. Our findings suggest that receptive vocabulary is a primary predictor of accented speech recognition in noise among monolingual listeners with limited exposure to L2-accented speech, whereas direct effects of real-life interactions were not observed.

  • Free-Energy-Driven LLM Agent for Modeling Information Exploration in Web Media

    Modeling cognitive dynamics in web media exploration can contribute to healthier information environments, for example by enabling inference of users' false beliefs and estimating how media designs may affect belief formation. We propose Free-Energy-driven Web Explorer (F-Explorer), a cognitive LLM agent that instantiates active inference for web exploration. In contrast to data-driven LLM approaches that learn behavior end-to-end, F-Explorer takes a theory-driven approach: it embeds an LLM as the belief module within a cognitive model grounded in the Free-Energy Principle (FEP), extending prior work that models human web exploration under FEP. We evaluate F-Explorer on a virtual SNS exploration dataset from a previous study and compare it with a reproduction of WEB-FEP with aligned initial-belief distributions. Using maximum likelihood estimation over 840 candidate settings per model (initial belief x learning rate), F-Explorer achieves higher action-sequence log-likelihood and improved top-k choice accuracy. Moreover, a proxy belief measure shows larger correspondence with participants' self-reported initial opinion scores, while correspondence with final opinion scores remains comparable. These results suggest that integrating LLM-based belief representations can improve FE–based models of web exploration while preserving interpretability through explicit latent-state and parameter inference.

  • Comprehending Novel And Conventional Metaphors In Marathi: An EEG Study

    This study investigates processing of conventional and novel metaphors in Marathi, an SOV language, with behavioural and electrophysiological measures, to extend graded models of metaphor comprehension established with SVO languages. Participants performed literal truth value judgment tasks with reaction time (RT), accuracy, and event-related potentials (ERPs) recorded. Behavioural results showed that literal sentences elicited the fastest RTs and the highest accuracy, contrasted by conventional metaphors' longer RTs and lower accuracy, reflecting decision-level conflict. Novel metaphors showed intermediate performance, indicating increased semantic processing demands with reduced conflict. ERP analyses of the N400 revealed greater centro-parietal negativity for novel metaphors relative to literal and conventional sentences, consistent with increased semantic integration difficulty, although ROI-level effects were not statistically significant. Overall, the findings support cross-linguistic robustness of graded model of metaphor comprehension, and how semantic processing varies as a function of familiarity by extending N400-based accounts of metaphor processing to Marathi.

  • The Mask of Consensus: Modeling Pluralistic Ignorance with Generative Agents

    Individuals frequently endorse views they privately reject to align with a perceived majority. This dynamic, known as pluralistic ignorance, is challenging to model because traditional scalar agents lack the semantic capacity to distinguish between private belief and public performance. To address this, we propose the Generative Cognitive-Social Simulation (GCSS) framework. We introduce the Grounding-Evaluation-Masking (GEM) architecture, which utilizes Large Language Models to simulate the strategic suppression of dissent based on Theory of Mind (ToM) risk assessment. Virtual classroom simulations suggest that ToM-guided risk evaluation is important for the emergence of the "spiral of silence." Crucially, we identify a computational mechanism of belief internalization: under prolonged masking, agents gradually adjust their private beliefs toward their publicly maintained stances. These findings quantify how social pressure can transform strategic conformity into more durable, collectively reinforced belief.

  • Relations are slower to encode than objects, explaining the same-different paradox

    Relations are key to everyday cognition, but research has suggested that reasoning about relations is more challenging than reasoning about objects. We propose that one source of the difficulty resides in encoding: relations are slower and more resource-intensive to encode than objects. Two experiments provide evidence for this claim. Participants completed two straightforward matching tasks: given a standard, they had to identify either a relational match from a nonmatch, or an object match from a nonmatch. The critical manipulation was how long the standard was presented (i.e., the time allowed for encoding). The results showed that as presentation duration diminished, accuracy for relational matching deteriorated steeply, while that for object matching was relatively unaffected. This supports the idea that relational encoding is more resource-intensive. This pattern held for both similarity (Experiment 1) and difference judgments (Experiment 2), suggesting that the underlying comparison process is the same. We discuss the implications of these findings.

  • Cognitive Behavior Trees: Integrating Somatic Marker Hypothesis into LLM-based Agent Planning

    Enabling embodied agents to complete long-horizon tasks in complex environments remains a fundamental challenge. Current LLM-based agents primarily focus on logical reasoning, yet they lack an intuitive perception of latent environmental risks. This often leads to failure during execution as they inadvertently overlook potential hazards. Inspired by the Somatic Marker Hypothesis in neuroscience, we propose a novel architecture called Cognitive Behavior Tree (CBT). By integrating LLM-based symbolic reasoning within behavior trees with activation steering risk premonition, we empower agents with the dual capacity for both deliberate planning and a premonition of potential risks. Extensive experiments in Minecraft demonstrate that CBT excels at handling high-difficulty tasks, achieving an average success rate of 26.60%—significantly surpassing existing state-of-the-art methods.

  • Who said that? Voice Recognition Across Changes in Language and Emotional Affect

    Voice recognition depends on listeners' ability to retain and identify voices despite variability in the speech signal. Here, we examined the relative impact of changes in the speaker's language and changes in their emotional state on voice recognition. Participants were exposed to a speaker and later asked to identify that speaker from a lineup of voices. Across trials, exposure and test were either matched or mismatched in language and emotion, introducing variability in the speaker's phonological and acoustic characteristics. This design yielded four conditions: no-mismatch, language-mismatch, emotion-mismatch, and complete-mismatch. Overall, switching emotion impaired voice recognition more than switching languages, especially when the exposure took place in an unfamiliar language and a neutral tone of voice, demonstrating that the magnitude of these effects is context dependent. A tentative explanation of these results is proposed in the discussion.

  • Possible Sex Differences in the Wakeful Rest Effect

    Prior work shows that a brief period of minimal interference, otherwise known as wakeful rest, has memory enhancing effects. A recent meta-analysis, Parra et. al (2026), observed large improvements in memory for older adults and patients, and, disappointingly, a smaller effect for young adults. Building off of the findings of this work and using existing data from our lab, we did a random-effects meta-regression to explore why. Specifically, we pooled 38 standardized mean differences, derived from 14 studies, to investigate the potential moderation effect of sex on the wakeful rest effect. The results show that rest was significantly beneficial for men, but not for women. Future work is needed to clarify why this pattern was observed.

  • From semantic memory to collective creativity: A generative cognitive foundation for social creativity models

    Simulation-based theory development has yielded powerful insights into collective performance by linking social structure to emergent outcomes, yet it has struggled to extend to collective creativity. Creativity is hard to capture purely at the social level, as novel ideas are generated through cognitive mechanisms. To address this gap, we introduce a multi-level socio-cognitive agent-based framework in which agents share a common semantic vocabulary and substrate but differ in semantic network topology. A single generative parameter tunes semantic modularity, yielding emergent individual differences in ideational breadth. When agents exchange ideation traces, two canonical social-creativity phenomena arise without being imposed: lower pre-interaction ideation overlap predicts larger stimulation gains, and shared inspiration sources induce network-level redundancy. The framework enables mechanistic theory-building about cognition and social structure in collective creativity.

  • Corn on the What? Development of Phrase Frequency Effects in Background Noise

    Children recognize and produce more frequent phrases (e.g., "corn on the cob") more accurately than less frequent phrases (e.g., "corn on the tray"). Here, we ask how phrase frequency affects language processing in background noise and whether this effect changes across childhood. We presented high- and low-frequency four-word phrases ("corn on the cob" vs. "corn on the tray") in quiet or in speech-shaped background noise to 40 children with typical language development aged 3-12 years. Children's repetition accuracy, response time, and phrase duration were measured. Across measures, results showed advantages for high- vs. low-frequency phrases, quiet vs. noise, and older vs. younger children. Interactions between age, frequency, and noise condition were near-zero for most measures, showing limited change across development in the effect of frequency and listening condition. Across typical development, phrase frequency aids language comprehension and production but is not preferentially used to overcome background noise.

  • The range of perceptual content in action-oriented perception

    In this paper we investigate the contents of object representations in action-oriented perception. We propose a view whereby perceivers represent the egocentric relationship that they have with potential perceptual objects, rendering those objects actionable for the perceiver. Further, we argue that the perceptual content associated with this actionable object representation stands apart from other standard forms of perceptual representation or experience. The actionability of perceptual objects is neither a kind of thin content (relating to bare sensory properties), nor is the content necessarily high-level (relating to personal-level cognitive conceptualization). Accordingly, we argue that the action-oriented content of perceptual experience is rich, but not so rich that it necessitates high-level representation. To draw the distinction between rich content and high-level content in object perception, we argue that high-level content is neither necessary nor sufficient for action-oriented perception. We support this claim by surveying empirical cases of optic aphasia and ideomotor apraxia.

  • Predictive Temporal Alignment in Motor Skill Acquisition: A Bio-inspired Computational Model of Quadrupedal Locomotion

    This study explores how agents achieve stable locomotion in complex terrain. They do this relying only on proprioception, without visual guidance. We propose the Asymmetric Reinforcement Learning Contrastive Optimization (ARLC) framework. This framework simulates predictive coding in biological neural systems through a contrastive learning mechanism. The core of this model is to establish temporal consistency between historical action sequences and future expected actions. Experimental results show that this model has excellent generalization ability in the physical world. This includes zero-shot transfer from simulation to reality. More importantly, its learned latent representation space shows a self-organizing mapping of terrain categories. This provides a new computational perspective for understanding the evolution of internal representations in biological motor control.

  • Rebound Dynamics in Illusory Contrast-Driven Motion Perception

    Motion perception is a fundamental component of biological vision, and contrast-driven motion illusions offer a particularly revealing window into its underlying mechanisms. These phenomena include reverse-phi motion, in which contrast polarity inversion leads to the perception of motion opposite to the physical displacement, and Mario-type apparent motion, where coherent motion is perceived in the complete absence of spatial displacement. Existing computational accounts often rely on explicit interactions between contrast pathways, yet provide limited insight into the temporal structure of such percepts. Here, we propose a computational model in which rebound-driven temporal dynamics play a central role in generating illusory motion. The model reproduces the characteristic non-monotonic dependence of perceived motion direction on temporal offset observed in psychophysical data. Our results suggest that contrast-driven motion illusions can arise from timing-dependent perceptual dynamics rather than explicit pathway coupling, highlighting the importance of temporal computation in motion perception. The code is available at: \url{https://github.com/NBELab/Cogsci2026}

  • Semantic-Aware Multimodal Fusion with Cross-Scale Features for Alzheimer's Disease Detection

    With the increasing prevalence of Alzheimer's disease (AD), automatic detection has emerged as a practical auxiliary diagnostic tool. Communication-related signals, being non-invasive, have been widely used in unimodal and multimodal approaches for AD detection with promising results. However, existing methods face two limitations: (1) relying heavily on final-layer encoder representations prevents comprehensive characterization of modality-specific features, and (2) treating all modalities equally introduces noise and reduces discriminative capability. Therefore, we propose SCMF-Net, a multimodal fusion framework. Specifically, it incorporates a cross-scale feature alignment module to exploit complementary representations across encoder layers, producing comprehensive modality-specific representations. In addition, SCMF-Net integrates a semantic-aware attention module that uses textual semantics to guide cross-modal fusion, thereby promoting semantic alignment. Experiments on the ADReSS and ADReSSo benchmark datasets demonstrate the effectiveness of SCMF-Net, achieving accuracies of 91.67% and 90.14%, respectively, with ablation studies validating the contribution of each component.

  • When Images Feel Wrong: Stress Testing Text-to-Image Safety via Cognitive Threats

    As text-to-image (T2I) models rapidly advance, their safety and ethical limits are increasingly scrutinized. Existing safeguards focus on filtering overt harms (e.g., violence, sexual content) but overlook subtler, cognition-based psychological attacks. We propose SPECTRE (Synthetic Perceptual Exploitation via Cognitive Threat Replication), a vulnerability framework that integrates seven visual-fear cognitive effects, such as the uncanny valley, trypophobia, and schema-violation, to synthesize images that provoke unease, disgust, or fear. Tested on four mainstream T2I models (SDXL, Qwen-Image, Doubao, and FLUX.1 Schnell), SPECTRE consistently bypassed built-in safety checks and produced numerous images that induced strong psychological discomfort. Our findings reveal a critical blind spot in current AI defenses: vulnerability to cognition-targeted attacks, and call for more robust safety strategies that understand deep semantics and human psychological effects. Warning: This paper may contain AI-generated content that could evoke disgust, fear, or discomfort. It is created solely for research. Readers are advised to proceed with caution based on personal judgment.

  • From Evidence to Action: Auditable Prompts for Calibrating Reliance in Cooperative Driving

    Cooperative driving depends on time-critical vehicle-to-everything (V2X) alerts whose evidence is distributed, incomplete, and sometimes adversarial. A key cognitive failure is \emph{miscalibrated reliance}: decision-makers may act on invalid alerts (overtrust) or dismiss valid ones (undertrust) under time pressure. We introduce \emph{auditable trust prompts}, a decision-oriented interface that compresses verifiable cues into compact badges and an action recommendation (\textsc{rely}/\textsc{verify}/\textsc{ignore}). Prompts integrate cross-source corroboration (peer vehicles and roadside units (RSUs), plus on-board sensing such as on-board diagnostics (OBD) and global navigation satellite system (GNSS) cues), recency-weighted sender reliability, and provenance availability/verification status, while highlighting uncertainty triggers that warrant selective checking. In a human-in-the-loop alert assessment task manipulating evidence uncertainty, adversarial contamination, and congestion-driven availability stress, auditable prompts reduce unsafe overreliance on invalid alerts without a commensurate increase in missed reliance on valid ones. Behavioral logs show that verification increases mainly in hard regimes, suggesting a targeted verification pathway rather than blanket distrust. These findings suggest that auditability can serve as a usable cue for calibrated reliance, beyond backend accountability, in dynamic cooperative driving.

  • Ascribing Leadership for Solving Coordination Problems with Asymmetric Information

    Coordination problems can be solved by designating a leader whose decisions others follow. However, this solution creates a second-order coordination problem: who should lead and who should follow? We investigate the decision to follow in coordination games where one player has supplementary information about which strategy yields which payoff. We varied whether this information was relevant for maximizing benefits and whether players' interests were fully aligned. Participants chose to adopt the follower role when their partner had relevant information and their interests were aligned. Asymmetric information further served as a focal point for coordinating on who should lead even when the supplementary knowledge was irrelevant to payoff maximization, with knowledgeable individuals emerging as leaders. Surprisingly, participants chose to follow even when their informed partner could potentially use their private information to the follower's detriment.

  • Pragmatic Explanation of Causal Inference from Correlational Statements

    People sometimes infer causal relationships from correlational evidence. For example, upon hearing cancer is associated with smoking, listeners might conclude smoking causes cancer. Without background information to inform priors, readers tend to interpret subjects of associated with as effects, rather than causes. We replicate this preference across multiple correlational predicates, and then derive it within a rational speech acts model of pragmatic inference. Central to our analysis is the observation that English asymmetrically constrains speakers: a speaker who plans to utter an effect in subject position has fewer, higher-cost alternatives compared to a speaker who plans to utter a cause in subject position. This asymmetry drives the pragmatic listener to infer that the use of a symmetric predicate signals that the subject is an effect. Our analysis additionally accounts for some of the variability in preferences for the subject-as-effect interpretation across different correlational predicates.

  • Forgetting Our Way to Shared Meaning: Effects of Forgetting on Conceptual Alignment in a Non-Partnership Coordination Game

    Shared meaning in language requires people to learn and agree on categories. We ask how characteristics of agents' memories change the emergence and evolution of shared meaning. Without a coordination game, models of conceptual semantics cannot explain how shared meaning emerges and changes in groups of people; however, existing games assume that players share payoffs in a partnership setting. We model conceptual alignment as a non-partnership game and illustrate differences in actual and perceived conceptual convergence from counterfactual simulations using agents with varying levels of adaptiveness and memory degradation. We found that adaptive players achieved actual convergence faster and had closer final conceptual regions than non-adaptive players, while non-adaptive players perceived convergence earlier. Weighing novel information less over time resulted in more stable agreements than fixing the weight of novel information. Memory features are critical to the emergence and evolution of actual and perceived convergence.

  • Markedness as Surprisal: Information Theory and Language Inefficiency

    Information-theoretic approaches to language predict pressure toward efficient, low-surprisal encoding, yet languages systematically retain forms that violate these pressures. Such departures are traditionally described as marked. In this study, we model markedness as localized phonological inefficiency, quantified using phonemic bigram surprisal. Focusing on iconic words in English (e.g., vroom) and ideophones in Japanese (e.g., fuyafuya), we examine how surprisal is distributed within words relative to language-specific phonotactic baselines. Across both languages, iconic forms exhibit reliably higher surprisal than non-iconic controls, indicating increased phonological unpredictability. This inefficiency is not uniform: surprisal is concentrated at specific word-internal positions, particularly at word onsets. In Japanese, baseline asymmetries in CV structure strongly shape surprisal, but ideophones retain distinct information-structural profiles once these constraints are controlled. The results presented here are limited in that they are correlational and do not by themselves establish that phonological inefficiency is functionally deployed, but nonetheless demonstrate a systematic association between iconicity and elevated surprisal.

  • Trial-Consistent Learning Improves Reliability in EEG Emotion Classification

    In stimulus-driven EEG emotion classification, labels are defined at the trial level. Each trial corresponds to a fixed stimulus, and windowing yields multiple dependent repeated measurements within the same trial. However, standard pipelines treat windows as independent samples, which ignores the trial as the true unit and encourages models to latch onto transient noise, leading to inconsistent within-trial predictions and unstable generalization. We turn the trial-as-unit prior in stimulus-driven emotion EEG into a trial-defined structural constraint. Specifically, windows are grouped by their trial IDs and a loss is added that makes predictions for windows within the same trial agree with each other. This penalizes reliance on window-specific transient noise and encourages the model to encode stable, trial-level representations. Across SEED and SEED-IV and three backbones including EEGNet, EEGConformer, and CTNet, the proposed constraint increases within-trial prediction consistency and stabilizes validation behavior. It also improves cross-session classification accuracy.

  • CogWave-KT: Multiscale Cognitive Volatility Modeling for Knowledge Tracing

    Knowledge Tracing (KT) aims to predict students' future performance based on their historical interaction data. With the advancement of attention mechanisms, attention-based KT models have achieved significant improvements in predictive performance. However, most existing attention-based KT models typically employ deterministic cognitive embeddings to represent students' knowledge states. Although such representations effectively capture the stability of cognitive processes, they fail to account for the inherent volatility in students' learning behaviors, thereby limiting the capacity for modeling authentic learning dynamics. To address this limitation, we propose a novel knowledge tracing model with enhanced capability for modeling cognitive volatilities—CWKT. Specifically, we model students' interaction representations as Gaussian distributions to capture the intrinsic long-term volatility in the learning process. We further apply a frequency decomposition to the mean of the Gaussian distribution, enabling us to extract local volatility information. To model the distributional transitions of students' knowledge states during learning, we introduce a Wasserstein Self-Attention mechanism. Moreover, an attention penalty module is incorporated to mitigate the model's potential overemphasis on long-term volatility, thereby improving the overall stability of predictions. Extensive experiments conducted on four public educational datasets demonstrate that the proposed model exhibits significant advantages in capturing cognitive volatilities.

  • Psychological Flexibility and AI Intimacy Addiction: A Paradoxical Mediation by Anthropomorphism

    While conventional addiction arises from disengagement with reality and can be prevented through increased psychological flexibility, intimacy with AI fosters emotional closeness and a sense of reciprocity, particularly when users attribute both agency and experiential capacity to their AI partners. It remains unclear whether psychological flexibility likewise mitigates addictive AI intimacy. The study recruited 116 Chinese college students with prior romantic interactions with AI partners and assessed their psychological flexibility, AI intimacy addiction level, and implicit attitudes toward these AI entities. Our results revealed a paradoxical role for psychological flexibility: among college students, it was associated with a stronger tendency to anthropomorphize AI, which in turn served as a significant indirect pathway to addictive behaviors. This suggests that the very ability to adaptively engage with internal experiences may, in this specific context, become a potential cognitive vulnerability, fueling anthropomorphism of AI and subsequent dependence.

  • Two-Phase Tempo Increase in Collective Clapping: Acceleration–Plateau Dynamics and Modality Effects

    In group rhythmic behavior, the tempo increase phenomenon can emerge even when participants appear well synchronized. We analyzed collective clapping from 49 participants while manipulating modality (auditory/visual/audiovisual) and group size (pair/trio). We tested a sequential effect in which beat-wise asynchrony A(k) predicts the subsequent tempo change _(k). For each trial, we estimated accel and plateau phases by fitting a one-breakpoint piecewise linear regression to the smoothed tempo series and selecting the breakpoint kp via an AIC-based grid search, using only the tempo series to avoid circular inference. In the confirmatory analysis w=7, an AR(1) mixed-effects model revealed a robust sequential effect: larger asynchrony predicted a more negative immediate tempo change (shorter subsequent ICI). The coupling was stronger during accel and attenuated toward zero during plateau, and it was strongest in the visual condition during accel. An exploratory check across w=1...15 yielded consistent patterns.

  • Tell-tale Signs of Implicit Bias: Language Abstraction for Automated Bias Analysis

    Our implicit biases are often reflected in our utterances. Existing AI research works are limited by their narrow focus on predefined, explicit biases. Thus, we propose a more generalized approach to bias analysis by implementing Linguistic Intergroup Bias (LIB) theory, which suggests people tend to use abstract terms to describe in-group positive and out-group negative behaviors, and concrete terms for the opposite. To leverage LIB, we propose a novel task named language abstraction span extraction, and finetuned a large language model for this task. With the model, we analyze the intergroup bias between the political left and right on Twitter and news datasets. Results show that our method can be used for bias analysis both quantitatively and qualitatively. More importantly, our paper provides the first large-sample-based empirical evidence to support LIB, affirming the correlation between abstraction and intergroup bias in social media and news domains.

  • Overhang Tower: Resource-Rational Adaptation in Sequential Physical Planning

    Humans effortlessly navigate the physical world by predicting how objects behave under gravity and contact forces, yet how such judgments support sequential physical planning under resource constraints remains poorly understood. Research on intuitive physics debates whether prediction relies on the Intuitive Physics Engine (IPE) or fast, cue-based heuristics; separately, decision-making research debates deliberative lookahead versus myopic strategies. These debates have proceeded in isolation, leaving the cognitive architecture of sequential physical planning underspecified. How physical prediction mechanisms and planning strategies jointly adapt under limited cognitive resources remains an open question. Here we show that humans exhibit a dual transition under resource pressure, simultaneously shifting both physical prediction mechanism and planning strategy to match cognitive budget. Using Overhang Tower, a construction task requiring participants to maximize horizontal overhang while maintaining stability, we find that IPE-based simulation dominates early stages while CNN-based visual heuristics prevail as complexity grows; concurrently, time pressure truncates deliberative lookahead, shifting planning toward shallower horizons: a dual transition unpredicted by prior single-mechanism accounts. These findings reveal a hierarchical, resource-rational architecture that flexibly trades computational cost against predictive fidelity. Our results unify two long-standing debates (simulation vs. heuristics and myopic vs. deliberative planning) as a dynamic repertoire reconfigured by cognitive budget.

  • The Temporal Efficiency Paradox: Fast Brains React, Slow Brains Understand

    The cortical temporal hierarchy—sensory regions process rapidly while association regions operate slowly—is well-documented, yet whether slow processing enables deeper integration remains debated. We analyzed fMRI data from four participants viewing ~65 hours of naturalistic movies, computing Temporal Dynamics Speed (TDS) and encoding performance across seven functional networks. Networks differed significantly in TDS (F(6, 993) = 18.42, p < .001): Frontoparietal fastest (1.20), Ventral Attention slowest (0.64), Default Mode intermediate (0.78). We hypothesized slower networks would benefit more from extended temporal integration; however, this was not supported: temporal averaging impaired encoding across all networks (negative TWG, p < .001), with no relationship between TDS and integration benefit. Critically, LSTM models outperformed ridge regression—even with concatenated temporal features—by 47–57% uniformly across all networks (p < .001). These findings resolve the paradox: all networks benefit from temporal integration, but only when sequential dependencies are explicitly modeled. Code: https://github.com/xiaoyanLi629/temporal-efficiency-paradox

  • Voice Preferences Across Yami, Atayal and Paiwan Languages: A Hierarchical Bayesian analysis

    Philippine-type Austronesian languages are well known for their rich voice systems, which include up to four voices that differ in which argument is the subject, rather than a simple active-passive contrast. These languages present a challenge to language acquisition theories, because they lack the robust relationship between causal agents and sentential subjects, a foundational assumption of both Nativist and Empiricist theories. We focus on one key issue: verbs differ in the voices they allow or prefer, raising questions about how children may acquire these patterns. We examine voice–semantics associations in three Austronesian languages spoken in Taiwan: Atayal, Paiwan, and Yami, using hierarchical Bayesian modeling to estimate probabilistic voice preferences for individual verbs while accounting for data sparsity. Hierarchical clustering of posterior distributions reveals overall biases in voice usage and partial alignment between semantic themes and voice-marking patterns, suggesting that voice-semantics distributions are systematic in ways learners may exploit during acquisition.

  • Episodic Counterfactual Generation Prioritizes Plausible and Vivid Imagined Alternatives

    Episodic counterfactual thinking refers to the capacity to imagine different ways in which the personal past might have unfolded. However, when generating imagined alternative versions of personal past events, people face a potentially unlimited number of ways to alter them. Current research on modal cognition has posited that when generating imagined alternative versions of events, people tend to sample more plausible alternatives first. In the current study, we asked whether this sampling mechanism is engaged when people generate alternatives for their personal past. To address this question, we asked participants to recall an episodic autobiographical event and then sequentially generate four alternative versions of that event. Employing Bayesian mixed models, our results showed credible evidence that participants sampled more plausible and vivid alternatives first, while showing no credible evidence that simulation difficulty change across sequential generations.

  • Modeling the Link between Confidence and Truth Judgments in the Repetition-Based Truth Effect

    Repeated exposure to information increases its perceived truth, a phenomenon known as the illusory truth effect. Beyond influencing truth judgments, repetition also increases individuals' confidence judgments. Existing theoretical models generally assume that confidence is lowest when statements are maximally ambiguous and highest when their truth status is known. However, it is currently unclear how the relationship between perceived truth and confidence can be formally modeled and whether such models are consistent with empirical data. In this study, we developed a Bayesian model to characterize the link between confidence and truth judgments, with several model variants capturing alternative assumptions about this relationship. The model also accounts for the effect of repetition on both variables. Fitting these models to two datasets provided evidence for an asymmetric, V-shaped relationship between truth and confidence. Specifically, individuals showed higher confidence when judging highly true statements than when judging highly false ones, with confidence reaching a minimum for ambiguous statements.

  • Perspectival Asymmetry in Mandarin Deictic Motion Verbs

    Deictic motion verbs require speakers to anchor motion to a perspective, yet the stability of perspective-taking in language use remains unclear. This study examines how Mandarin speakers produce and interpret lai 'come' and qu 'go' across twelve person-based conditions, focusing on the theoretically challenging qu2 contexts where motion occurs between non-speaker locations. In Experiment 1 (sentence completion), participants sometimes avoided deictic commitment by using the non-deictic verb dao 'arrive'; inappropriate responses were concentrated in qu2 contexts as predicted. In Experiment 2 (forced-choice between lai and qu), overall appropriateness increased, but the same vulnerability of qu2 persisted. These results reveal a robust asymmetry: while lai remains stable across contexts, qu is selectively vulnerable when the speaker is not a spatial anchor. This finding advances our understanding of deixis by highlighting that real-time deictic choice is shaped not only by lexical knowledge, but also by the dynamics of perspective-taking in discourse.

  • Shapes, Colors, and Space: The Role of Concrete Referents in Abstract Concept Processing

    This study investigates whether concrete visual referents modulate the processing of emotionally valenced abstract concepts. Using a picture-word interference paradigm, Spanish-speaking adults evaluated the valence of abstract adjectives while ignoring concurrent visual stimuli varying in shape (Kiki vs. Bouba), color (red vs. green), and spatial position (left vs. right). Accuracy and reaction times were analyzed using mixed-effects models. Results showed that color reliably influenced performance: congruent color-valence pairings (green-positive, red-negative) increased accuracy, and green selectively reduced reaction times for positive adjectives. Shape modulated reaction times only, with faster responses to positive adjectives paired with the Bouba shape. Spatial position did not interact with adjective valence in either measure. These findings indicate that concrete referents do not contribute equally to abstract concept processing and support embodied and multimodal accounts in which perceptual features, particularly color and shape, facilitate the online evaluation of abstract emotional meaning.

  • The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models

    Conceptual analysis---proposing definitions and refining them through counterexamples---is central to philosophical methodology. We study whether language models can perform this task through iterated analysis and repair chains: one model instance generates counterexamples to a proposed definition, another re- pairs the definition, and the process repeats. Across 20 concepts and thousands of counterexample-repair cycles, we find that, although many LM-generated counterexamples are judged invalid by both expert humans and an LM judge, the LM judge accepts roughly twice as many as humans do. Nonetheless, per-item validity judgments are moderately consistent across humans and between humans and the LM. We further find that extended iteration produces increasingly verbose definitions without improving accuracy. We also see that some concepts resist stable def- initions in general. These findings suggest that while LMs can engage in philosophical reasoning, the counterexample-repair loop hits diminishing returns quickly and could be a fruitful test case for evaluating whether LMs can sustain high-level iterated philosophical reasoning.

  • Linguistically attested, nested constituency structures are preferentially learned across learning domains

    A well-known hypothesis for language learning is that learners are predisposed to learn patterns attested in human language, while eschewing alternative hypotheses that are linguistically unattested. Despite the popularity of this hypothesis, supporting experimental evidence has remained scarce. We aim to fill this gap by measuring the effectiveness with which adult learners can infer a linguistically attested pattern (a hierarchically nested constituency structure) in comparison to an unattested one (a constituency structure based on non-consecutive constituents) within a controlled artificial language learning task. We also investigate whether any observed learning biases extend to a non-linguistic, general puzzle-solving task. We find that our study participants are better at learning linguistically attested structures over unattested ones across the two types of learning domains. More broadly, our study serves as a methodological proof-of-concept extensible to testing the strength of other kinds of learning biases in linguistic and non-linguistic learning tasks.

  • The GRASP Model of Belief Updating: The Impact of Source- and Argument Conditional Expectations

    Bayesian models gauge and potentially describe how people integrate source reliability and argument strength when updating beliefs, but no existing account captures their joint dynamics. We present GRASP, a process-level model that combines Bayesian-inspired weighting of source reliability with non-Bayesian extensions (a confidence-based gate and a faint-praise function), organised around five core mechanisms: confidence-based gating, source-conditional expectations, faint praise inference, reliability weighting, and bidirectional reliability updating. Our main experimental findings (N= 591) are: Stronger than expected arguments increased source reliability and moved beliefs toward speaker positions, which contrasted with weaker than expected arguments; argument quality dominated initial credentials in reliability judgments; source-conditional expectations govern reliability judgments more than belief updating. The model demonstrates how multiple phenomena emerge from one principle: evidence is evaluated relative to source- and argument conditional expectations.

  • Causal judgments in complex situations

    The Counterfactual Simulation Model (CSM) proposes that people make causal judgments by mentally simulating what would have happened in relevant counterfactual scenarios. The model assumes that causal judgments incorporate two aspects: whether-causation (would the outcome have occurred without the candidate cause?) and how-causation (did the candidate cause affect how the outcome occurred?). The CSM accurately predicts causal judgments for simple billiard ball scenarios (2 or 3 balls). Here, we explore whether it also captures people's judgments in more complex situations (up to 12 balls). We measured whether-causation (Experiment 1), how-causation (Experiment 2), and causal judgments (Experiment 3) for billiard ball scenarios varying in complexity. A model combining both whether- and how-causation best predicted causal judgments, though a how-causation-only model performed almost as well. The CSM's whether-causation predictions diverged from human judgments in complex scenarios, while how-causation predictions remained robust. This suggests that when participants make causal judgments, they might shift from counterfactual simulation toward simpler force-tracking heuristics as complexity increases.

  • Grammatical Gender Amplifies Social Gender Biases in Lexical Representations: Evidence from Word Embeddings

    This study examined whether grammatical gender, a structural property of language, causally shapes word representations to exaggerate implicit biases about social gender. Because isolating the effect of grammatical gender from semantic, socio-cultural, and (other) syntactic factors is difficult in human studies, we used word embeddings. We created 97 near-English languages from an English corpus, holding content constant but varying gender marking systems (e.g., which genders were marked, which parts of speech were marked). Separate FastText embeddings were trained on each corpus. Introducing grammatical gender increased implicit gender bias for nouns with grammatical gender but no inherent gender (grammatically-feminine nouns became more associated with women, and vice versa for grammatically-masculine nouns). This effect strengthened as agreement extended to more parts of speech and was strongest under asymmetric marking. These results provide causal evidence that grammatical gender alone may amplify gender-related social biases of word representations in artificial and, perhaps, human minds.

  • Referential Production Across the Adult Lifespan: A Virtual Reality Study

    Speakers often use modifiers (e.g., color) to distinguish objects, even when such details are unnecessary. Overspecification was traditionally viewed as inefficient, but contemporary accounts suggest it may be adaptive because modifiers can aid comprehension. Prior work also shows that modifier use changes with age, with older adults overspecifying more than younger adults. However, little is known about modifier use across the lifespan, including middle adulthood, or in realistic settings, as most studies rely on simple 2D displays. To address these gaps, we examined referential production across the adult lifespan in an immersive virtual environment, where participants described everyday objects paired with either a competitor differing in color or state, or an unrelated object. Overall, participants exhibited high overspecification rates, with only a modest increase in middle and older ages and no robust age effects on expression length or modifier use, offering new insights into real-world communication across the adult lifespan.

  • Objects Before Words: Object-First Inductive Biases for Grounding Language in Child-View Video

    Learning grounded word meaning from natural experience requires resolving two ambiguities in infant-view recordings: when the named referent appears and where it is in a cluttered frame. In SAYCam-style data, caregiver speech is sparse and weakly synchronized with egocentric video, so single-frame contrastive pairing yields noisy positives in which the intended object is absent or entangled with distractors. We propose BabyMind, an object-first bias for child-view contrastive learning under sparse, noisy supervision. BabyMind extracts candidate object embeddings using an offline mask-based region interface, links candidates across a short utterance-centered window into lightweight object files via tracking, and aligns utterances to bags of object files with a prototype-space multiple-instance contrastive objective. Track-coherence and global-object agreement regularizers stabilize learning and transfer object-file structure into the global frame embedding used at evaluation. On SAYCam-S, BabyMind improves Labeled-S 15 forced-choice accuracy by +2.6 points over CVCL and yields consistent gains on in-vocabulary out-of-distribution benchmarks.

  • Applying Latent Space Network Modeling to Compare Word Associations of School-Age Children and Adults

    This study uses latent space network modeling to explore how the structure of the mental lexicon might change from middle childhood (ages 7 to 11 years) to young adulthood. English-speaking children and adults (N = 21 per group) completed a repeated word association task where they generated the first word that came to mind in response to a list of 48 cues (24 nouns, 24 verbs) repeated three times. We applied mixed-effects models to examine how response characteristics, derived from robust metrics (e.g., WordNet, BERT), varied by group and list repetition. Semantic and phonological distance between responses and cues increased over list repetitions. Adult responses exhibited stronger semantic (word embedding and taxonomic) similarities, while child responses showed stronger phonological relatedness. To evaluate factors associated with producing specific responses, we modeled the corpus using latent space networks. This confirmed a stronger influence of distributional statistics in adults and phonological relatedness in children.

  • MoEEG: A Sparse Mixture-of-Experts Transformer for Universal EEG Representation Learning

    Electroencephalography (EEG) provides a non-invasive and temporally precise window into neural dynamics, yet learning transferable representations across subjects, paradigms, and electrode layouts remains challenging. Existing EEG foundation models typically rely on patch-wise tokenization and dense Transformer blocks, which may weaken fine-grained temporal continuity and provide limited adaptability to heterogeneous neural patterns. We propose MoEEG, a sparse Mixture-of-Experts (MoE) Transformer for generalizable EEG representation learning. MoEEG combines global-local temporal embedding, electrode-aware channel embedding, a dual-path EEGAttention module for temporal and spatial dependency modeling, and sparse expert routing for adaptive representation learning. Linear-probing evaluations on P300, ERN, and motor imagery benchmarks show that MoEEG consistently outperforms BIOT, LaBraM, and EEGPT, with balanced-accuracy improvements over EEGPT of 2.94, 1.69, 0.72, and 1.43 percentage points on PhysioP300, KaggleERN, BCIC-2A, and BCIC-2B, respectively. These results suggest that sparse expert routing and EEG-specific attention can improve cross-task EEG representation learning. Our open source code is available at https://github.com/YuuSenW/MoEEG.git.

  • EmoVAE: Bridging Affective Neural Dynamics for Cross-Dataset Emotion Recognition

    Emotion recognition from neural signals constitutes a cornerstone of cognitive science, providing objective access to affective dynamics. However, reliable decoding across datasets is hindered by domain shifts in recording protocols and neurophysiological variability. We propose EmoVAE, a domain adaptation framework that bridges emotional neural dynamics via variational latent distribution alignment. EmoVAE uses a multiscale spatiotemporal aggregation encoder to capture spatiotemporal patterns of affect-related activity, and an EEG-Latent module to map these representations into a continuous probabilistic space. Within this latent space, we align class-conditional distributions using conditional maximum mean discrepancy, yielding stable, geometry-aware emotion clusters across domains. Experiments on SEED and SEED-VII show state-of-the-art performance for both cross-subject and cross-dataset protocols, demonstrating robust generalization under domain shift. These findings suggest that continuous latent alignment is an effective strategy for building reliable EEG-based emotion recognition systems for cognitive and clinical applications.

  • Executive Function Variability in Biologically Plausible and Optimized Neural Agents

    Executive functions support flexible, goal-directed behavior in dynamic environments. Existing models typically attribute the substantial variability in human performance on such tasks to abstract, resource-based accounts of attention and working memory, leaving the nature of the underlying resource underspecified. Here, we investigate executive-function variability using a modified, continuous version of the Pong game that imposes sustained demands on multi-object tracking, control, and action selection. To account for the broad distribution of human performance levels, we compared a performance-optimized deep reinforcement learning agent with a biologically plausible spiking neural network that incorporates explicit neural resource constraints. Whereas the unconstrained artificial agent achieved superhuman performance, the spiking model reproduced the full spectrum of human scores. Systematic variation in neural resources, perceptual noise, and processing delays produced performance regimes corresponding to low, median, and peak human behavior, including saturation at object counts consistent with known working-memory limits.

  • Shifted Interpretation of Pointing Gestures in Speech Reports

    We report an experiment on the interpretation of pointing gestures in speech reports. German speakers read a narrative in which a speaker reported another person's utterance using either direct or indirect discourse. The reporting speaker produced a pointing gesture whose interpretation depended on whether it was taken to be part of the reported speech event or a contribution of the current utterance. As only direct discourse is typically assumed to be quotational, we hypothesized that participants would interpret the gesture as part of the reported speech event more frequently in direct discourse. Participants identified the referent of the gesture and rated their certainty. The results confirmed the hypothesis: Participants were more likely to interpret the pointing gesture as part of the reported speech event in direct discourse. At the same time, gesture transformations were rare and associated with lower certainty across conditions, indicating competition from the canonical referential function of pointing.

  • Multi-Granularity EEG Decoding: Bridging Object-Background Semantics and Temporal Motion for Video Reconstruction

    Reconstructing video from brain signals is a pivotal task in neural decoding. However, current frameworks struggle by directly aligning noisy EEG with coarse global captions and VAE latents, inevitably leading to semantic misalignment and temporal flickering. To bridge this gap, we propose M-STAR, a bio-inspired framework that mimics human cognitive decomposition. To our knowledge, M-STAR is the first framework to decompose dense captions into object- and environment-level semantics for robust alignment and employ an "Anchor-plus-Residual" strategy that anchors generation to a stable visual anchor and applies residual dynamics to maintain stability over time. Specifically, this is achieved via the Factorized Semantic Module (FSM) and the Anchor-Guided Dynamics Module (ADM), respectively. Extensive evaluations on the SEED-DV dataset demonstrate that M-STAR significantly outperforms state-of-the-art methods, generating high-fidelity videos with precise semantics and coherent motion. M-STAR establishes a new baseline for fine-grained visual decoding, marking a significant step towards high-fidelity visual brain-computer interfaces.

  • Why do young children fail in the scale-model task?

    Symbol use is considered to be one of the few capacities that are unique to humans. The most influential series of studies on the development of symbolic competence, spearheaded by Judy DeLoache and colleagues, adopted the scale-model task. These studies showed that it is not until three years of age that children display an understanding of symbols, while younger children fail to establish stand-for relations between symbols and their referents. Here we propose that young children's failure in the scale-model task does not reflect a limitation in symbolic competence, but a limitation in mapping depicted discourse episodes onto actual situations. By distinguishing between (reality-directed) representations and (discourse-directed) depictions, we argue that the scale-model task conflates symbolic understanding with an additional operation: linking internally represented episodes to real-world situations.

  • Cross-Subject EEG Emotion Recognition with Periodic Evolution Representation and Prototype Boundary Learning

    EEG emotion recognition is important for affective computing and for understanding the neural encoding of emotion. However, EEG signals are highly non-stationary and susceptible to task-irrelevant interference. In cross-subject settings, such variability causes intra-class distribution shifts and increases inter-class overlap, undermining stable discriminative learning. We propose PEPBNet, a cross-subject EEG emotion recognition model with periodic evolution representation and prototype boundary learning. PEPBNet captures stable emotion-related rhythms by modeling multi-granularity intra-period patterns and evolutionary relations between adjacent periods. We further decompose the learned representations into mutually orthogonal emotion and interference factors to suppress non-emotional components. Prototype learning is integrated with decision boundary learning by jointly optimizing classification and prototype losses, improving intra-class compactness and inter-class separability. Experiments show that PEPBNet achieves accuracies of 96.30% on SEED and 87.04% on SEED-IV, and delivers competitive performance on DEAP for both the valence and arousal dimensions, demonstrating improved robustness and generalization.

  • The Turing Test Is Here to Stay

    The Turing test, first proposed by Alan Turing in 1950, has historically served as a benchmark for evaluating artificial intelligence (AI). However, since the release of ELIZA in 1966, and particularly with recent advancements in large language models (LLMs), AI has been claimed to pass the Turing test. Furthermore, criticism argues that the Turing test primarily assesses deceptive mimicry rather than genuine intelligence, prompting the continuous emergence of alternative benchmarks. This study argues against discarding the Turing test, proposing instead using more refined versions of it, for example, by interacting simultaneously with both an AI and human candidate to determine who is who, allowing a longer interaction duration, motivating participants with a bonus payment, "training'' the participants on past conversations, access to the Internet and other AIs, using experienced people as evaluators, etc. Through systematic experimentation using a web-based platform, we demonstrate that richer, contextually structured testing environments significantly enhance participants' ability to differentiate between AI and human interactions. Namely, we show that, while an off-the-shelf LLM can pass some Turing test environments, it fails to do so when faced with a more robust one. We further demonstrate that GPT-4.5, which previous work has shown to trick 70% of the human participants, fools less than 20% of the human participants in the most advanced Turing test we propose. Our findings highlight that the Turing test remains an important and effective method for evaluating AI, provided it continues to adapt as AI technology advances.

  • Features of Experience-Based Learning

    Risk communication often relies on descriptive probabilities, yet there is a general preference for experience-based impressions that can be biased by small samples. Interactive simulations can combine both modes, but evidence is mixed and key design features remain confounded. In two online experiments (N = 422; N = 360), we tested how features of simulated experience affect expected-value maximization (EV-Max) in monetary loss gambles. Probabilities were communicated using icon arrays only, and each option involved two independently co-occurring losses. In Experiment 1, the presentation format had only marginal effects, with no advantage for active sampling, whereas sampling without replacement robustly increased EV-Max choices. In Experiment 2, sampling control did not affect EV-Max choices; accumulative displays yielded only marginal EV-Max gains but increased EV-Max preference, confidence, perceived informativeness, and search length. Overall, increased autonomy was largely irrelevant, while reducing sampling noise and using accumulative summaries appear beneficial for risk communication.

  • From "A Bridge Too Far" to Infrastructural Integration: Designing Computational Thinking for Early Childhood

    Computational thinking (CT) in early childhood is increasingly understood as a set of foundational problem-solving practices rather than as early programming. Yet, a persistent challenge remains: how cognitively informed designs can be integrated into everyday classroom practice. This challenge echoes the longstanding "bridge" problem in cognitive science and education, concerning the difficulty of aligning cognitive principles with pedagogical routines and institutional conditions. In this paper, we examine this integration problem through a hybrid tangible-digital CT learning design developed through a participatory design process and implemented in public preschool classrooms with children aged 4-6 years. Focusing on sequencing, decomposition, and debugging, we analyze how embodied interaction and peer collaboration support anticipatory reasoning, error-based revision, and shared regulation during problem solving. Preliminary qualitative results indicate that children construct and revise procedural plans through tangible action, engage in collaborative debugging, and sustain attention in rule-based tasks. Crucially, these outcomes depend not only on cognitively plausible tasks, but on their infrastructuring into classroom orchestration and existing pedagogical routines. We argue that participatory design provides a concrete mechanism for addressing the bridge problem by enabling multi-level integration of cognitive theory, interaction design, and educational practice.

  • Anticipating Naming: Multimodal Behavioural Dynamics Leading to Successfully Attended Naming Events

    What leads to successfully attended naming events? To map novel labels onto novel objects, children must attend to the correct objects around during naming events. The current study investigates, through a bottom-up approach, which behavioural dynamics lead to successfully attended naming events. In a semi-naturalistic head-mounted eye-tracking study, we tested 23 caregiver-child dyads (17 signing, 19 non-signing). Children acquiring a sign language are known to distribute their visual attention differently from children acquiring a spoken language. The current study investigates if differences in language modality lead to differences in the behavioural dynamics of attended and unattended naming events. Findings suggest that successfully attended naming events follow different behavioural dynamics than unattended naming events and that effects of language modality are most prominent in the amount of mutual gaze between children and their caregivers.

  • Temperament Predicts Vocabulary Development of Infants from Low-Income Families: A Longitudinal Structural Equation Model Analysis

    This study aimed to elucidate the influence of temperament on vocabulary development of infants from low-income families, using data from the Early Head Start Research Evaluation Study (N = 970). Longitudinal structural equation modeling was used to evaluate contributions of latent temperament, with sustained attention, engagement of parent, and negativity (reverse-scored) as indicators, and emotion regulation as predictors of vocabulary knowledge at ages 14, 24, and 36 months. Latent temperament predicted both vocabulary knowledge and emotion regulation skills at subsequent time points, whereas emotion regulation was non-significant in predicting vocabulary knowledge after controlling for temperament. Although descriptive analyses indicated greater vulnerability of boys than girls across measures (e.g., lower vocabulary scores), structural relations appeared to be invariant across genders. Modeling temperament as a latent variable accounts for the common variance shared across constructs. Future research should aim to uncover how distinct facets of temperament influence vocabulary development trajectories.

  • Now or later: A reinforcement learning model of behavioural delay

    (Other) people notoriously delay the initiation and completion of work, sometimes beyond what is optimal. One prominent, and indeed experimentally validated, explanation is based on the fact that rewards delivered in the future are discounted. However, other factors can interact with discounting and affect policies, such as the amount of effort and the probability of successful completion. These have received less empirical attention. Here, we build and fit a new reinforcement learning model to the working trajectories of students over the course of a semester in a real-world task (P. Y. Zhang & Ma, 2024). We show that discounting, effort, and efficacy are all important in explaining students' delays. In addition, the discount factors inferred from task performance correlate significantly with self-reported measures of impulsivity and procrastination, as well as discount rates estimated from a monetary delay discounting task, highlighting that they robustly capture meaningful individual differences in temporal preferences.

  • Vision-Language Model through the lens of Linguistic Relativity

    Linguistic relativity proposes that the categorical structure of a language systematically biases how continuous perceptual domains are partitioned. While this effect is well studied in human color perception, it remains unclear whether vision–language models (VLMs) exhibit a language-conditioned categorical structure analogous to it. In this work, we treat VLM as an \emph{artificial subject} and adapt psycholinguistic paradigms to test for Whorfian effects in a multimodal neural network. Using a tightly controlled set of color-discrimination triads that vary only in OKLab lightness, we quantify model decisions across English and Russian lexical frames. For a contrastive dual-encoder VLM, we observe a robust shift in the internal decision boundary between light and dark blue when the language of the textual input changes, despite identical visual input. This boundary displacement disappears in vision-only embeddings, demonstrating that the effect is not perceptual in origin but emerges from vision–language alignment. In contrast, when languages share the same color lexicon (e.g., English and French), category boundaries largely overlap, and when generic labels are used, no stable boundary emerges. Together, these findings show that a computational analogue of linguistic relativity can arise in artificial systems, but only when visual representations are explicitly aligned with language. Our results establish a framework for testing cognitive theories in multimodal models.

  • Comparing Inherent Learning Capabilities of Recurrent Neural Networks

    Recurrent neural networks (RNNs) remain central to cognitive modeling, despite the rise of transformers, because their recurrent dynamics are relatively cognitively plausible. But many questions remain about the strengths and weaknesses of RNNs as learning systems and cognitive models. We compare two architectures: Elman networks with hidden-layer recurrence and Jordan networks with output-layer recurrence. Using carefully constructed artificial languages, we show that Elman networks excel at learning long-range dependencies, especially when intervening elements provide no predictive cues. Jordan networks perform equally well when adjacent elements correlate with target outputs or when sequence-level error signals are available. However, Jordan networks falter in particular with item-level error feedback (as opposed to sequence-level error feedback), particularly for nonlinear sequential relations such as XOR. These results clarify how recurrence placement shapes sequence learning, suggest links to cognitive and neural processing, and highlight the representational advantages of hidden-layer recurrence over simple output chaining.

  • MEG Evidence for Shared Neural Representations of Phrase Grammaticality Across Serial and Parallel Presentation

    Language can occur in highly serial (speech) or fully parallel form (text), yet how our brains adapt to seriality differences is not understood. Here, we held modality constant (visual) and examined whether processing mechanisms underlying grammaticality detection are similar under Rapid Serial Visual Presentation (RSVP) and Rapid Parallel Visual Presentation (RPVP). To identify shared neural representations of grammaticality, we trained a classifier on magnetoencephalography (MEG) responses to grammatical and ungrammatical three-word phrases in one presentation mode and tested its generalization to the other. Behaviorally, accuracy was higher and reaction times faster for grammatical vs. ungrammatical phrases across presentation modes. Neurally, cross-presentation decoding revealed shared two-stage activity at word 2 across train-test directions, and an early shared representation at word 3 when trained on RPVP. We take the bidirectional representations at word 2 to capture combinatory processes and the asymmetry at word 3 as reflecting stronger global structure encoding in RPVP.

  • Statistical Properties of the Environment Influence Action Abstraction in Human Planning

    Planning requires significant cognitive effort, as it involves constructing and evaluating numerous potential futures. A key method to simplify planning is the grouping of sequences of actions into cohesive units-often referred to as action chunking. While studies confirm the existence of action chunking in human planning, we explore chunking as a learning process across multiple tasks with different statistical properties. To that end, we ran a large-scale preregistered study using the Lightbot navigation paradigm, in which we manipulated the underlying statistical properties of the planning tasks. We found that the majority of participants adapted their chunking behaviour to match the environment's statistical properties by either learning to use the action chunk that is optimal given these statistical properties or a subsequence thereof. Together, our results suggest that statistical properties of the environment influence action chunking in human planning.

  • Item Distinctiveness Modulates the Relationship Between Hits and False Alarms in Old-New Recognition

    Recognition memory is strongly shaped by similarity structure, yet how item distinctiveness reorganizes the relationship between hits and false alarms is not well understood. Using distances in a multidimensional feature space to quantify item similarity, we examined item-level recognition performance for real-world objects. Low-distinctive items showed positive correlations between hit rates and false-alarm rates, reflecting variation in inter-item similarity as predicted by standard global-familiarity models. In contrast, highly distinctive items exhibited high hits but low false alarms, reducing or reversing the overall hit–false-alarm correlation. Split-half analyses showed that both hits and false alarms were reliable across participants, but their relationship depended on an item's position within the surrounding similarity space. The Hybrid-Similarity Exemplar model combining inter-item similarity with a distinctiveness-dependent self-match signal captured this reorganization. These findings show that memorability is reliable yet relational, emerging from similarity structure rather than intrinsic stimulus properties.

  • When Efficient Communication Explains Convexity

    Much recent work has argued that the variation in the languages of the world can be explained from the perspective of efficient communication; in particular, languages can be seen as optimally balancing competing pressures to be simple and to be informative. Focusing on the expression of meaning---semantic typology---the present paper asks what factors are responsible for successful explanations in terms of efficient communication. Using the Information Bottleneck (IB) approach to formalizing this trade-off, we first demonstrate and analyze a correlation between optimality in the IB sense and a novel generalization of convexity to this setting. In a second experiment, we manipulate various modeling parameters in the IB framework to determine which factors drive the correlation between convexity and optimality. We find that the convexity of the communicative need distribution plays an especially important role. These results move beyond showing that efficient communication can explain aspects of semantic typology into explanations for why that is the case by identifying which underlying factors are responsible.

  • From Private Beliefs to Public Silence: A Multi-Agent LLM Simulation of Psychological Safety

    How do organizational phenomena emerge from individual cognition? We address this question using a multi-agent simulation in which LLM agents possess individual cogni- tive architectures, including private beliefs, pairwise trust, and social-risk calculation. Using a cognitive process-tracing approach, we prompt agents with universal psychological principles rather than prescribed behaviors, thereby avoiding circularity. We observe the emergence of the paradox identified by Amy Edmondson: higher interpersonal trust leads to significantly greater information disclosure, mirroring the find- ings in human organizations. This macro-level pattern arises purely from individual cognitive modeling without explicit team-level programming. Furthermore, we show that team structure shapes these dynamics: inter-faction trust dramatically reduces information suppression, and trust effects amplify with team size, yielding greater returns in larger organizations. These results demonstrate that psychologically meaningful organizational phenomena can emerge from cognitively grounded agent models, opening a path toward a systematic computational science of organizations.

  • MetaCoT-CTRL: Metacognitive Control for Reliable Chain-of-Thought Reasoning in Large Language Models

    Chain-of-thought (CoT) prompting improves the accuracy of large language models (LLMs) by externalizing intermediate reasoning steps, yet these traces are often unreliable, containing local inconsistencies or spurious steps that undermine interpretability and trust. We propose \textbf{MetaCoT-CTRL}, a metacognitive control framework that treats CoT as an explicit cognitive trajectory subject to monitoring and correction. MetaCoT-CTRL first generates a low-cost System-1 draft reasoning chain, then applies a conflict monitor to localize unreliable steps via contradiction, entailment, and task-specific consistency signals. A System-2 targeted repair module selectively revises only high-conflict segments, preserving stable reasoning. To ensure robustness, the repaired chains are further evaluated using causal consistency and counterfactual validation under semantically preserving input perturbations. A cognitive cost regularizer discourages redundant or overly long reasoning. Experiments on ANLI, SVAMP, and CommonQA show that MetaCoT-CTRL improves reasoning reliability and stability over existing CoT-based methods while maintaining efficient inference.

  • A Resource-Rational Analysis of Forgetting in Continual Learning

    In machine learning, forgetting is typically treated as a defect, whereas human forgetting is understood as a rational response to limited resources and environmental change. We ask whether forgetting can likewise be optimal for machines under realistic constraints. We formalize a continual learning setting, parameterized by a hazard rate governing task volatility and a memory cost penalizing stored data. We compare three replay-based agents: naïve (no memory), unlimited replay (infinite memory), and resource-rational (finite memory). The resource-rational agent significantly outperforms the unlimited replay agent in overall utility across volatile or memory-constrained environments. Optimal memory usage decreases predictably as task volatility and memory cost increase. The resulting forgetting dynamics are best described by an exponential model, showing quantitative alignment with human memory laws. Forgetting thus provides a rational strategy for machines to use when engaging in continual learning in volatile environments or when subject to memory constraints.

  • Coordinated Minorities Exploit Social Influence in Multi-Agent Debate Systems

    Multi-agent debate (MAD) systems improve collective decision quality by aggregating diverse perspectives. However, this open interaction introduces security vulnerabilities. While existing research primarily focuses on single-agent attacks, threats from adversarial coalitions remain underexplored. Therefore, we propose the Dynamic Cognitive Manipulation Attack (DCMA), a framework to investigate how adversarial coalitions exploit sociocognitive dynamics to manipulate group consensus. DCMA consists of two components: a coordination mechanism that dynamically synchronizes coalition intentions through a perception-refinement-execution loop, and a strategic engine that selects appropriate persuasion strategies based on debate context. Empirical evaluations across four large language models validate the attack's effectiveness, with Attack Success Rate (ASR) reaching up to 22.9% and convergence delay nearly doubling compared to uncoordinated attacks. Our analysis also reveals three sociocognitive vulnerabilities: domain vulnerability asymmetry, latent process contamination, and scale-amplified vulnerability. Finally, we propose detection schemes and identify future research directions to address the limitations of existing defenses.

  • Uncovering Stable and Domain-Specific Drivers of Cognitive Impairment via Time-Slice Causal Discovery

    Cognitive impairment is a major public health challenge in aging populations. Identifying stable, actionable drivers is crucial for targeted intervention. Causal discovery has been leveraged for its interpretability to explore these factors. However, causal modeling approaches that integrate long-span, multi-wave data to specifically address the heterogeneity of cognitive impairment remain underdeveloped, thus inherently failing to reveal domain-specific causal pathways from long-term, sparse longitudinal data. To address this, we introduce a framework that decomposes cognition into distinct domains and then applies time-slice causal discovery to sparse, multi-wave data. Applied to 8-year CHARLS data, the framework decomposed cognition into Mental State and Contextual Memory, identifying both domain-specific and shared causal drivers. These drivers were validated through robust predictive modeling, external generalization (CFPS), and interpretability analyses, demonstrating their utility. Our findings provide new insights for heterogeneous prevention and targeted early intervention.

  • Conceptual diversity of linguistic structure: a mixed-methods corpus study

    The goal of this paper is to uncover how conceptualizations of linguistic structure differ between the practice of research on language emergence and standard, formalized approaches of philosophy of language and theoretical linguistics. We analyze a corpus of 5000 scientific articles within the field of language emergence. Adopting a mixed-methods approach, we combine computational analysis—semantic search, topic modelling, and collocation analysis—with qualitative close reading. Results indicate that researchers in language emergence view linguistic structure as a complex phenomenon, which cuts across some existing theoretical distinctions, most importantly between "compositional" and "contextual" accounts. These results indicate the need for a new view of the systematicity of linguistic structure.

  • Formal Linguistic Competence is not a Monolith: Consequences for LLMs

    The capacities of large language models (LLMs) remain disputed, but one particular hypothesis has recently become widely endorsed: LLMs have a uniform human-level formal linguistic competence in grammar and semantics, while diverging in functional competence that extends to broader cognitive and social capacities. I raise two challenges with this account. First, the separation between weak and strong generative capacity should not be sidestepped. LLMs can process and produce grammatical text, but this does not reveal how they represent its structure. Second, the dissociability of the two competences does not eliminate influence between them. Distinct functional competences are not expected to coincide with identical formal competences, since the choice between different weakly equivalent grammars is likely regulated by functional competence. I further evaluate this hypothesis by reviewing experimental work on the behavioral and mechanistic analysis of LLMs, which does not point to a uniform formal competence across models.

  • Semantic map learning and externalization in an embodied neural agent: A comparison to human behavioral and neural data

    All mammals can build internal cognitive maps from sensory input, supporting spatial learning and planning. While rodent studies have shown the mammalian navigation system solves the SLAM (Simultaneous Localization and Mapping) problem, its role in human spatial memory and overt recall are less understood. To investigate this, we adapted a spiking semantic SLAM algorithm for a "Treasure Hunt" task, where human participants navigate a 3D beach in virtual reality and later point to remembered object locations. Our agent integrates networks for bipedal locomotion, vision, memory, and arm control to enable first-person learning of place-object associations, and externalizing that knowledge by pointing and expressing confidence. Comparing model observables to human data, we replicate key behavioral and neural effects: monotonic scaling of accuracy with confidence, and recall-dependence on local field potential power observed in the left hippocampus. This work offers a mechanistic framework linking embodied navigation, memory, and communication in human spatial cognition.

  • Effect of Retention Interval on Mnemonic Processing in the Mnemonic Similarity Task

    Pattern separation enables discrimination between similar experiences and relies on hippocampal function. The Mnemonic Similarity Task (MST) indexes this process using the Lure Discrimination (LDI). Recent studies, on the other hand, suggest MST performance is influenced by perceptual processing. We investigated the extent of mnemonic processes in MST performance across increasing delays between encoding and retention phases (0 minutes, 30 minutes, 8 hours, 24 hours). LDI declined in the 8-hour and 24-hour groups, with significantly worse lure discrimination especially for low similarity lures. Response patterns revealed a bias, shorter delays showed an "Old" response bias, whereas longer delays were characterized by a "New" response bias. Additionally, decreasing lure similarity was associated with increased "New" responses. Together, these findings suggest that the MST is sensitive to mnemonic processing over short delays, however performance at longer delays is constrained and may reflect non-mnemonic factors.

  • How Different LLMs Behave in Unverifiable Decision Scenarios?

    The behaviors of Large Language Models (LLMs) as artificial social actors are largely underexplored, particularly in unverifiable scenarios where conventional benchmarking is often not applicable. Thus, examining their behaviors in such scenarios can help understand and improve LLMs' capabilities of simulating real-world social actors in many tasks such as LLM-empowered agents. We draw a typical unverifiable scenario--a simplified pull request scenario on \textsc{GitHub} focusing on decision-making based on Activity Overview signal--to investigate how human and LLMs behave. We introduce a method to collect, compare, and reason about human and LLMs' decisions. We reveal that there are both similarities and differences between human mind and LLMs' decisions, and proprietary LLMs generally behave more like human than open-source LLMs do. We further find that human and LLMs may rely on different information and reasoning mechanisms in decision-making. Our study thus urges more work on human and LLMs decision-making in unverifiable environments.

  • Mapping Music Teachers' Beliefs About Students Gifted for Music: A Latent Profile Analysis

    Stereotypical beliefs about gifted students shape teaching and learning contexts. However, little is known about teachers' beliefs in domain-specific contexts, such as music. The present vignette-based online study examined whether distinct belief profiles could be identified among the 278 participating (prospective) music teachers evaluating a fictitious student gifted for music on achievement-related, socio-emotional, and behavioral characteristics, as well as on personality traits. Latent Profile Analysis revealed two profiles reflecting a generally positive view of musically gifted students, consistent with the so-called harmony hypothesis, differing only in extremity: Moderate Raters and Extreme Raters. A Monte Carlo simulation with 1,000 iterations confirmed that this two-profile solution best fits the present data. The findings suggest domain-specific, positive stereotypical teacher beliefs about students gifted for music and underscore the need to consider their potential effects in teaching contexts, particularly regarding external expectations and the pressures that may be exerted on these students.

  • Prompting Children's Curiosity Through Questions: An Experimental Test of Framing in Science Learning

    Promoting curiosity may support learning, but there is little research assessing effective methods of promoting curiosity in children. Framing a lesson using questions to make uncertainty salient may be one such method. This study examined whether question framing led to higher self-reported and demonstrated curiosity compared to a statement-framing. Children (age 7-8, N = 110 after exclusions) were randomly assigned to hear a science activity introduced using either statements or a question about what they could learn, or they generated their own question about what could be learned. Children then engaged in exploration to learn. Curiosity was assessed using self-report before and after the exploration activity, and time exploring and number of explorations during the task. There were no differences across conditions, but children's self-reported curiosity before exploration predicted their behavioral curiosity, supporting the use of children's self-reported curiosity in future research. Possible explanations for the results are discussed.

  • Measuring Algorithm Understanding in Humans and AI

    Understanding algorithms is a central learning goal for those in computational fields and is frequently used as a cognitive benchmark for contemporary large language models (LLMs). We present a hierarchy that characterizes algorithmic understanding in terms of observable performance on both concrete and abstract tasks to juxtapose human and LLM performance and draw conclusions about the latter's understanding of the subject space. We evaluate the hierarchy in a controlled assessment study comparing 43 university students with eight LLMs across five canonical algorithms. The resulting performance profiles reveal that LLMs often excel at tasks dependent on concept retrieval while failing on tasks requiring computational example construction. These differences mirror distinctions in rote versus meaningful learning and suggest that algorithm understanding in humans and LLMs may rely on different underlying resources.

  • EEG Correlates of Linguistic Error Detection in Turkish Learners and Fluent Bilinguals

    This study used a picture–word verification task to examine EEG correlates of lexical and morphological error detection in adult learners (N=36) and fluent bilingual speakers of Turkish (N=18). Prior to EEG, we trained learners on a miniature version of Turkish featuring 36 nouns inflected for case (dative, ablative) and number (singular, plural) presented in pre-recorded spoken dialogues with corresponding pictures. During EEG, participants judged whether the picture matched the inflected Turkish noun. Both groups showed increased negativity to lexical violations (wrong noun) in the 300–500 ms window, consistent with an N400, with a smaller, more prolonged effect in learners. Fluent bilinguals showed an early negativity (100–300 ms) followed by a P600 (500–700 ms) to case and number violations (wrong inflection). Learners failed to show a P600 and exhibited a weak, sustained negativity. Findings suggest that adult second language learners may initially process morphological errors similarly to lexical errors.

  • GPT surprisal as a limited proxy for human sentence acceptability: Evidence from Korean clausal constructions

    The present study compares human acceptability judgements of Korean clausal constructions—dative alternations, active–passive voice under animacy manipulation, and negative polarity item licensing—with surprisal estimates derived from GPT-style language models. Fifty-six native speakers evaluated 224 sentences, and surprisal was computed using three Korean-capable variants (KoGPT-2, Ko-GPT-Trinity, and mGPT). The models reproduced the strong preference for Dative–Accusative datives, captured voice- and animacy-related effects only partially, and failed to represent clause-bounded NPI licensing. These findings indicate limitations of surprisal as a proxy for acceptability beyond morphosyntax and motivate the development of cross-linguistic benchmarks for explainable AI and diagnostic evaluation.

  • What GPT language models get right—and wrong—about honorific expression: Evidence from Korean subject honorification

    The present study investigates how GPT language models approximate human sentence-processing dynamics observed in Korean subject honorification. Drawing upon two open-access datasets, we compare human responses with surprisal estimates from two Korean-capable GPT models. Across tasks, the models reliably identified overt morphosyntactic violations, assigning high surprisal to clear honorific agreement mismatches associated with low politeness ratings and increased reading times in humans. However, GPT surprisal failed to capture more graded and context-sensitive effects, including the optionality of the subject honorific suffix, animacy-dependent politeness evaluations, and delayed spillover effects in human reading. Correlations between human measures and model surprisal were consistently weak and highly condition-specific. Taken together, while GPT surprisal serves as a coarse indicator of well-formedness, it inadequately represents the socio-pragmatic and discourse-level constraints that shape Korean sentence processing. This calls for the need for more cognitively and culturally grounded metrics in language sciences and explainable AI.

  • Independently learned behavioral strategies support coordination in common marmosets

    Coordination, where two or more individuals align their actions to achieve a common goal, has been documented in several species using the rope-pulling paradigm. Nonetheless, the extent to which subjects understand their partner's role in facilitating successful coordination remains unclear. We tested five pairs of common marmosets (Callithrix jacchus) to address this question using successive phases with increasing demand for coordination (decreasing rope length and delaying one monkey). Two of the tested pairs succeeded on all rope lengths as well as a 5s delay but failed a 10s delay. Behavioral analyses showed that these two pairs achieved success through different mechanisms employed by each subject, suggesting that individually learned behavior may be sufficient to support coordination in the absence of shared intentionality, at least in the context of a physically intuitive task. Our study yields novel insights into the cognitive underpinnings of coordination in a cooperatively breeding species.

  • Possible identities versus possible absence in three-year-olds' modal reasoning

    Modal concepts—representations of possibility, necessity, and impossibility—play a central role in human thought. Some argue that these are uniquely human and dependent on modal language, whereas others propose earlier developmental and evolutionary emergences. Recent success by rhesus macaques on a modified 3-cup task supports the latter view. Here, three-year-old children (N =144), who struggle with modal words, completed equivalent modified 3-cup tasks across three experiments. When preserving most features of the original 3-cup task—contrasting a guaranteed reward with a possible reward or its absence—children mirrored macaque performance and did not exceed chance. However, children successfully preferred a guaranteed reward to a possible reward or non-reward. A follow-up experiment confirmed that our paradigms did not bias participants towards the guaranteed side. Although this pattern broadly paralleled macaques, evidence for a between-experiments difference in performance was weaker, likely reflecting different primary constraints on 3-year-olds' modal reasoning abilities.

  • Exploring the role of music in language development: Caregiver song negatively correlates with infant speech-related vocalization

    Caregiver infant-directed (ID) speech and infant vocal productions are both important for speech-language development. Although music experience is associated with language abilities at later ages, the role of music in infant language development is unclear. We used manual annotation to obtain quantities of different infant vocalization types and quantities of caregiver speech and song during 5-minute excerpts from daylong audio recordings of 3-, 6-, 9- and 18- month-olds. Excerpts in which caregivers produced more infant-directed musical vocalizations (such as singing or humming) also contained fewer speech-related infant vocalizations. This could be driven by musical input causing infants to be quieter as they attend auditorily and/or by adults singing or humming during times when infants are less vocal. We also found a negative association between non-infant- directed adult speech and infant speech-related vocalization, positive associations between quantities of infant-directed speech and infant crying and laughter, and age trends consistent with prior reports.

  • Evidence of context-dependence in people's representations of negative numbers

    Negative integers lack direct perceptual referents for humans and require integrating magnitude with polarity to use effectively. Prior work on integer representation has produced inconsistent findings and multiple hypotheses. We propose that these inconsistencies reflect context-dependent recruitment of multiple coexisting representations. Adults completed a symbolic magnitude comparison task involving positive – positive, negative – negative, and mixed – sign pairs, with presentation format (simultaneous vs. sequential) and numerical context (grouped vs. ungrouped comparison type) manipulated. We evaluated the componential, extension, and reflection hypotheses using individual-level Bayesian models. Group-level analyses revealed no distance effect for mixed-sign pairs. Individual-level modeling unveiled participants flexibly recruiting different representations across contexts, with extension-based representations emerging primarily when task demands were minimized. Reflection-based representations were rare. These findings highlight the flexibility of integer representations and emphasize that group-level analyses can obscure the dynamics of individuals' internal representations.

  • The Persistence of Self-Preference Bias in LLM Evaluations of Creativity

    Large language models (LLMs) are increasingly used to evaluate human performance in high-stakes decisions. Here, we investigate LLMs' ability to evaluate creativity, a hallmark of human cognition. Creativity is a robust predictor of academic and professional success. It is also shown that prioritizing creativity in high-stakes decisions such as college admissions mitigates traditional biases. We examined how GPT and Llama models evaluate creativity in authentic admission essays and AI-generated counterparts. LLM ratings were compared against an independent semantic divergence index that closely approximates human judgments of creativity. Results revealed strong self-preference bias in LLMs, inflating ratings for their own outputs and penalizing human-authored essays. We applied different mitigation strategies including zero- and few-shot prompting, and fine-tuning models on expert creativity ratings. Mitigation strategies reduced bias, with fine-tuning proving most effective, but could not fully eliminate it. These findings raise important concerns about using LLMs as judges of human creativity.

  • Divergence of Metacognition and Performance under Increasing Task Demand in Multiple-Object Tracking

    As task demands increase, fewer cognitive resources remain for executing a task and monitoring one's own performance. This study investigates the point at which performance and metacognition (self-assessment) begin to diverge under increasing mental demand. Participants completed a dynamic visual tracking task in which they monitored target dots moving among distractors on a screen. Task difficulty was manipulated by varying the number of target dots from two to four, among a total of eight dots. Participants were prompted to estimate their performance during the trial, prior to receiving feedback. After each trial, they also completed the Paas scale to assess perceived cognitive workload. Participants accurately assessed the quality of their performance during the task and detected changes in workload afterward. They were less accurate in their judgements at higher difficulty levels. This suggests that as task demands increase, an individual's ability to monitor task performance is reduced.

  • Developmental Trajectories and Individual Profiles in Sentence Comprehension in Spanish

    Sentence comprehension involves integrating multiple cues varying in salience and reliability across structures. A key question is how children weight these cues across sentence types over development. This study examined how Spanish-speaking children aged 3 to 11 years integrate word order, morphological marking, and syntactic embedding. Using a sentence–picture matching task, we assessed comprehension accuracy for five structures differing in canonicity and complexity. Group-level analyses revealed a clear developmental ordering, with early success for structures with canonical word order, more gradual improvement for structures requiring morphosyntactic cues to override linear order, and persistent difficulty when non-canonical order was combined with syntactic embedding. Profile-based analyses showed substantial individual variability, with distinct patterns of strengths and weaknesses across structures among same-age children. These findings indicate that sentence comprehension development reflects both age-related gains in accuracy and differences in cue weighting across structures, highlighting the value of combining developmental and individual-level approaches.

  • Causal Hierarchy-Guided Feature Reconstruction: Enhancing Interpretability and Performance of Postoperative Delirium Prediction in ICU Patients

    Postoperative delirium is a prevalent acute neurocognitive dysfunction in ICU patients, impairing key cognitive functions and deteriorating clinical prognosis, which is closely related to cognitive science research on neurocognitive regulation. Traditional machine learning prediction methods depend on statistical correlations and suffer from poor interpretability, failing to reflect the inherent causal hierarchy of delirium. To solve this issue, this paper integrates causal inference and proposes a Causal Hierarchy-Guided Feature Reconstruction Model (CDM). It adopts Greedy Equivalence Search and graph neural networks for causal feature reconstruction, followed by Bayesian-optimized XGBoost. Evaluated on the MIMIC-IV dataset, our model achieves superior performance with an AUROC of 0.9054, effectively enhancing model interpretability and offering a novel paradigm for cognitive disorder prediction.

  • Semantic Cues in the Learning of Alternating Argument Structures: An Insight from Mandarin Ditransitive Alternation in the Adult Input

    Ditransitive alternation in adult input leads to multiple ways of syntactic-thematic mapping, posing challenges for child language acquisition. Semantic cues associated with arguments may guide children in overcoming this problem, as the semantic bootstrapping hypothesis assumes. This study attempts to explore whether such semantic cues are available and reliable in adult input that might help children to establish thematic-syntactic mappings of alternating ditransitive syntactic frames. We extracted all double-object sentences and post-object GeiP sentences from the adult input of 16 Mandarin-speaking children, which approximate the English double-object sentences and dative sentences, and coded both the indirect and direct objects along the dimensions of animacy, pronominality, givenness, and indefiniteness. The results show that the indirect and direct objects differ systematically across all dimensions. These distinctions are robust across syntactic frames. This suggests that cues for thematic-syntactic mapping are available in the input, which is consistent with the semantic bootstrapping hypothesis.

  • Tiered Social Recursion: A Functional Architecture for the Evolution of Consciousness

    Treating consciousness as an evolutionary adaptation of internal world modeling, we introduce the Tiered Social Recursion framework that classifies biological awareness. Each tier is delineated with behavioral taxonomic benchmarks. The operationalization comprises a hierarchical sensory Parser and an autoregressive loop linking Generative Engine (integrated with Long-term Memory), Short-term Memory, and sensory-grounding Gate Unit. This generative-perception architecture is anchored in neuroanatomy and disorders of self-awareness, accounting for a spectrum of normative and pathological mental states through unit-activation patterns. For example, in the REM state, LTM switches to consolidation mode, so peri-conscious processing of STM content dominates and dreams feel unreal; in Anton syndrome, loop-generated visuals persist without Gate Unit constraints. Motivated by Cogitate's PFC-aligned evidence, we propose that qualia emerges from peak recursive modeling of outer-self in the minds of conspecifics. We posit that lack of persistent identity prevents the modeling of outer-self required for human-like consciousness in AI.

  • Towards Scalable and Ecological Methods for Early Social Development Research: Validating Automated Facial Expression Estimation in Egocentric Video

    Wearable head-mounted cameras offer a window into infant's social world, yet manual coding remains a computational bottleneck. To address this limitation, we evaluate a scalable, open-source pipeline for automated facial expression estimation using the EfficientNet-B0 architecture on naturalistic, egocentric videos. Despite non-canonical viewpoints, poor lighting and motion artifacts, the model achieves performance (F1 = 0.543) comparable to state-of-the-art benchmarks. We demonstrate representational alignment (r = 0.76) between model confusion patterns and human perception, suggesting that the model reflects the adult social-perceptual structure. While categorical accuracy varies, the model's distributional probabilities capture the graded nature of affect. We also identify a critical divergence in neutral expression processing, where model and human ambiguity dissociate (r = -0.26). These findings support automated facial expression estimation as a viable approach for quantifying infants' natural social input, while highlighting category-specific limitations. Together they represent a step toward a "big data" science of early social development.

  • The Information-Computation Gap: Computational Complexity of Belief Updating Predicts Quality of Human Beliefs

    Rational updating of beliefs requires the processing of new information, which, in turn, requires computational resources, which are limited. Limited computational resources may give rise to an information-computation gap: the quality of an agent's beliefs may not reach the information-theoretic limit due to limited computational resources. Here, we quantify the computational resources required for belief updating in an objective, ex-ante, and task-independent way and document, using a laboratory experiment, that indeed the quality of human beliefs declines in proportion to the computational demands of updating. Further, we validated this metric by showing that it makes unique predictions regarding belief quality. These findings advance a prescriptive theory of beliefs updating that accounts for systematic deviations from rational behaviour. Our results highlight the importance of incorporating computational constraints when evaluating decision-making performance.

  • Towards a Hippocampus–Neocortex Inspired Two-Stage Diffusion Framework for Single-Image Dehazing

    Modeling perceptual restoration under severe degradation remains a core challenge in vision science. Inspired by the hierarchical organization of human episodic memory, we propose a cognitively grounded two-stage diffusion framework for single image dehazing. Analogous to neocortical–hippocampal division of labor, the first stage encodes a coarse, stable global structure, while the second stage progressively restores fine-grained details conditioned on this intermediate representation. This staged reconstruction mirrors memory consolidation, where global context is stabilized before perceptual detail is refined, enabling a more efficient and cognitively plausible coarse-to-fine inference process. By decomposing restoration into functionally distinct stages, our framework avoids inefficient multi-scale optimization and promotes stable, coherent convergence through target-conditioned diffusion dynamics. Extensive experiments demonstrate that the proposed model achieves state-of-the-art performance, exhibiting strong robustness, interpretability, and practical applicability in complex real-world dehazing scenarios.

  • When Hallucinations Travel: How Evidence-Like Cues Shift Integration, Verification, and Source Memory in Human–AI Interaction

    AI hallucinations can reshape cognition when false content becomes useful for decisions. We test how AI-generated false claims become usable in reasoning and difficult to separate from their source. Across two preregistered studies (Study 1: N = 295; Study 2: N = 242), participants made a medication choice after reviewing standardized information and receiving an AI reply varying in hallucination severity. Study 1 shows that stronger hallucinations shifted understanding of a risk relation, changed onset-duration weighting, and increased source confusion, especially for a fabricated causal mechanism. Study 2 traced this pathway in speech: stronger hallucinations increased uptake during decision verbalization, reduced critical evaluation and reliance on original evidence, and weakened source reasoning during interaction recall. Later transmission and post hoc reflection were comparatively bounded. These findings locate hallucination risk at the transition from exposure to reason construction, where AI-originated content becomes useful in participants' explanations while provenance becomes difficult to track.

  • Inability to Control Inner Speech Linked to Negative Mental Health

    In this exploratory study, we investigated how task-based inner speech initiation and suppression relate to broad psychological dimensions. A large online sample (N = 453) completed two one-minute inner speech tasks (initiation vs. suppression) and a battery of 11 questionnaires covering inner experience, personality, mental health, and self-regulation. Exploratory factor analysis identified seven latent factors, including Negative Mental Health and Inner Speech. The Negative Mental Health factor predicted higher reported inner speech frequency despite attempted suppression, while showing no association with attempted initiation. The questionnaire-derived Inner Speech factor predicted general inner speech propensity in the tasks. These findings suggest that impaired suppression, rather than elevated initiation, characterizes inner speech dysregulation in mental health risk, highlighting the importance of incorporating regulation and control dimensions into future inner speech measures.

  • Is language needed to construct–ötwo-place predicates? Event imitation in toddlers

    A central question in cognitive science is whether language merely expresses pre-existing concepts or provides new forms of conceptual representation. Some recent studies suggest that representing generic two-participant relations (dogs chase birds) requires linguistic encoding, in contrast to specific events (this dog chased this bird) or one-participant generic relations (dogs jump). We explored whether 17- to 21-month-olds, who do not yet systematically produce transitive sentences, can represent generic two-participant events. Children (n-28) completed an imitation task with either one-participant events or two-participant events. Toddlers reliably reproduced the observed action and generalized role assignments to novel exemplars (new dogs and birds), performing above chance in the two-participant condition. These results suggest that toddlers can represent transitive event relations prior to the productive mastery of transitive constructions, constraining strong claims that linguistic production is necessary for forming such representations.

  • Sensations into Stereotypes: Large-Scale Measurement of Cross-Sensory Bias in Text-to-Image Generation

    Human perception relies on crossmodal correspondences, systematic links between different sensory modalities. As text-to-image (T2I) models increasingly serve as aesthetic infrastructure, they risk codifying cultural biases embedded in sensory language. We analyze the mapping of cross-sensory bias in mainstream T2I models, examining how gustatory, tactile, auditory, and olfactory adjectives map onto visual attributes. Using a modality _ carrier design, we evaluate five bias dimensions in 6,000 images via automated VLM-based assessment: color, demographic, entity, situational, and layout. Our results reveal widespread cross-sensory homogenization, with models projecting abstract sensations onto a narrow set of cultural prototypes. This study presents a framework for quantifying cross-sensory bias in T2I models and offers tools for auditing and mitigating their broader cultural and social impact.

  • Recent linguistic experience influences event role identification

    While events can be identified based on both perceptual and higher-level information, the contribution of language-related factors remains unclear. Here, we examine whether recent linguistic experience modulates event role identification scenes by asking whether verbs with different argument and thematic structures cue abstract event representations that guide subsequent event cognition. Participants briefly viewed agent–patient events (for 300ms) preceded by a transitive verb, an intransitive verb, or a nonsense word control, and then completed a probe-based role identification task. Transitive verb primes facilitated recognition of event roles, yielding faster responses in role-matching trials relative to intransitive and control conditions. These findings show that linguistic cues influence event role identification, especially when their argument and thematic structure aligns with event structures.

  • Diversity and Interaction Structure Shape Performance and Search Dynamics in Joint Cognitive Search: An Agent-Based Simulation

    Cognitive search, conceptualized as information foraging through mental spaces, underpins many daily tasks and is frequently performed jointly. Yet, the mechanisms by which social interaction impacts search dynamics remain understudied. Through three agent-based simulations of a verbal fluency task, we investigated how cognitive diversity and interaction structure modulate collective performance in joint cognitive search. In Experiment 1, under strict turn-taking, moderate diversity benefited performance by inducing increased exploration, but high diversity levels proved detrimental. Experiment 2 revealed that flexible interaction protocols mitigated these costs by enabling distinct, more adaptive search strategies. Experiment 3 demonstrated that enhancing individual cognitive flexibility via working memory mitigated risks related to high diversity. Collectively, these results identify flexibility, whether situated in social interaction protocols or individual cognitive mechanisms, as a critical route for unlocking the benefits of diversity in collective search, and offer a novel computational paradigm to study the emergent mechanics of collective intelligence.

  • Modeling Fairness Judgments of Splits Between Multiple Parties

    Fair divisions are a fundamental problem for moral cognition. Past experimental work has provided evidence in support of three principles of fairness in divisions between two parties: proportional splits, equal splits, and splits that equalize net gains (the Nash bargaining solution). When the number of parties increases, so does the complexity of such decisions, potentially influencing people's cognitive strategies and their reliance on precise explicit heuristics vs intuitive approximations. We design a novel task in which participants can easily and intuitively sample and select among various distributions of resources between multiple people (2 to 18) by adjusting a continuous slider that updates divisions in real time. In two preregistered experiments (n = 378; 2,268 choices), we quantitatively model participants' fairness judgments at the individual level. We find that participants can be categorized into three main groups. Overall, around 50% are best fitted by proportionality, 40% by the Nash bargaining solution, and fewer than 10% by equality. At the aggregate level, proportionality and the Nash bargaining solution perform best. Fairness judgments remain stable as the number of involved parties increases. When they only have access to the slider (Experiment 1), participants best fitted by the Nash bargaining solution seem to intuitively approximate it, but in an imprecise way. By contrast, they tend to precisely select it when three buttons (one for each model, Experiment 2) are available to automatically adjust the slider. These results suggest that a substantial proportion of participants rely on intuitive approximations for the Nash bargaining solution, consistent with bargaining-based (contractualist) theories of fairness.

  • From Thread to Number: Weaving as scaffold for mathematical thinking

    Mathematical cognition is often studied through formal systems or experimental tasks, leaving open how abstract numerical concepts emerge from everyday practices. This paper reviews ethnographic and ethnomathematical studies of weaving crafts, arguing that these practices minimally and systematically recruit a set of cognitive capacities that are central to quantification: grouping, ordering, pairing, memory, exhaustion detection, and cardinality. Across cultures and technologies, weaving stabilizes these capacities through material constraints, procedural regularity, and visual structure, affording the emergence of geometric and numerical knowledge. Therefore, weaving is put forward as a model system for studying the embodied and culturally scaffolded origins of mathematical abstraction, with implications for theories of numerical cognition and material engagement.

  • Causal (In)efficiency: Breaking Markov Violations Through Structural Uncertainty in Causal Chains

    One of the most persistent findings in causal cognition is that individuals frequently deviate from normative Bayesian reasoning, most notably through Markov violations, where they incorporate conditionally independent information into their judgments. Recent research has sought rational explanations for these deviations, ranging from structural uncertainty to memory sampling limitations. In this study, we evaluated four rational models using a novel experimental paradigm that introduced structural uncertainty into a causal chain by manipulating the functional properties of the causal links. We found that participants systematically violated the Markov assumption when causal links shared the same mechanism. Computational modeling identified the Bayesian Uncertainty Model (BUM) as the most plausible explanation for these results, with sampling-based methods like the Mutation Sampler appearing as close contenders. We discuss these findings in light of the need to rethink the definition of rationality in Bayesian causal reasoning.

  • AFSPP: Agent Framework for Shaping Preference and Personality with Large Language Models

    The evolution of Large Language Models (LLMs) introduced a new paradigm for investigating human behavior emulation. Recent research employs LLM Agents in sociological environments, where they exhibit behavior based on unfiltered LLM characteristics. However, these studies overlook iterative development in human-like settings. Human preference and personality are complex, constantly shaped by environmental and subjective influences. We propose the Agent Framework for Shaping Preference and Personality (AFSPP), exploring the impact of social networks and subjective consciousness on Agents' preference and personality formation. AFSPP demonstrates trends consistent with key findings from human personality experiments. AFSPP-based results indicate that planning, perception, and subjective social networks have the most pronounced influence on preference shaping. AFSPP shows potential to enhance the efficiency of psychological experiments and provides insights into preventing undesirable preference and personality development for trustworthy AI.

  • Effects of encouraging vs. reflective feedback on children's explore-exploit decisions

    Explore-exploit decisions are often made without accurate awareness of one's own behavioral tendencies, limiting effective metacognitive regulation. Feedback has the potential to improve metacognitive calibration, but its effectiveness depends on its informative value. This study examined how encouraging and reflective feedback--differing in informative value--influence children's explore-exploit decisions, while also considering the role of feedback frequency. Results showed that high-frequency reflective feedback reduced children's exploration but led to higher reward performance, suggesting the development of more adaptive search strategies. In contrast, children exhibited a more exploratory search pattern under low-frequent reflective feedback. Although encouraging feedback provided relatively low informative value, both low- and high-frequency delivery sustained high exploration, potentially by fostering emotional safety. Overall, these findings offer practical implications for education, highlighting the importance of strategically balancing affective reassurance and metacognitive control to support children's effective learning and exploration.

  • A Study on the Predictors of Problem Difficulty for the Planning Serious Game Tik Tik

    Tik Tik is a serious game designed to study the cognitive mechanisms of planning by combining ecologically valid navigational demands with a highly flexible problem structure. Drawing on data from three studies totaling 122 participants, we show that Tik Tik problem difficulty is primarily driven by a smart, non-exhaustive search of its problem space, in which character trajectories are represented. Furthermore, we find that working memory and visuo-motor skills contribute, highlighting the types of cognitive representations and skills Tik Tik also taps. Our four-parameter model incorporating these factors explains a large proportion of the variance in problem difficulty. Finally, we compare these cognitive demands to those of the most popular planning task, the Tower of London, highlighting commonalities and differences between these two tasks.

  • Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework

    Sarcasm is a pragmatic phenomenon in which speakers convey meanings that diverge from literal content, relying on an interaction between semantics and prosodic expression. However, how these cues jointly contribute to the recognition of sarcasm remains poorly understood. We propose a computational framework that models sarcasm as the integration of semantic interpretation and prosodic realization. Semantic cues are derived from an LLaMA 3 model fine-tuned to capture discourse-level markers of sarcastic intent, while prosodic cues are extracted through semantically aligned utterances drawn from a database of sarcastic speech, providing prosodic exemplars of sarcastic delivery. Using a speech synthesis testbed, perceptual evaluations show that semantic and prosodic cues enhance perceived sarcasm, with the combined system achieving the best downstream F1 while maintaining high subjective sarcasm ratings. These findings highlight the complementary roles of semantics and prosody in pragmatic interpretation and illustrate how modeling can shed light on the mechanisms underlying sarcastic communication.

  • When giving more makes you look worse: paradoxical inferences in a Bayesian model of social evaluation

    People readily infer how much another agent cares about their welfare, for example after observing this agent give them some of what they have. However, these inferences become more difficult when there is uncertainty over the resources someone has to share, a common real-world scenario. We develop a Bayesian com- putational model of how people infer the welfare trade-off ratio (WTR) of another agent under resource uncertainty. This model predicts that under uncertainty people should average over differ- ent possible hypotheses about their partner's resources. In a be- havioral study (N = 129), we found that across donation amounts participants' WTR estimates closely tracked those made by our Bayesian model. Notably, participants exhibited the following paradoxical pattern predicted by our model: they sometimes saw agents as less generous when they gave them more money, in cases where a high donation revealed that the agent is rich and is giving a comparatively small portion of their resources. These findings demonstrate that people are able to make sophisticated inferences consistent with rational Bayesian reasoning, updating their beliefs about a partner's resources and adjusting their WTR estimates accordingly.

  • The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models

    Human intelligence scales through cumulative cultural evolution (CCE), a ratchet process in which innovations are retained against entropic drift. Large language model training, by contrast, still depends primarily on static corpora and parameter growth, leaving little room for endogenous accumulation through interaction. We present POLIS (Population Orchestrated Learning and Inference Society), a framework in which heterogeneous agents generate solutions, verify one another's outputs, retain validated artifacts in shared cultural memory, and internalize them through parameter updates. On mathematical reasoning benchmarks, populations of 1--4B-parameter models achieved average gains of 8.8--18.9 points over base models and narrowed the gap to 70B+ monoliths. Mechanistic ablations identify peer verification as the main ratchet operator and show that internalization sustains accumulation across rounds, providing computational evidence that epistemic vigilance organizes durable knowledge growth. These results position structured social interaction as a scaling lever orthogonal to parameter count.

  • Self-Explaination Improves Multiple Document Integration

    Individuals are frequently presented with information that must be integrated across multiple sources. The current study ex- plores the role of self-explanation strategies in supporting across-text integration during multiple document comprehen- sion. Participants (n=139) read a multiple document text set, while either self-explaining or thinking aloud, then com- pleted verification questions (sentence and inference verifica- tion), a prior knowledge test, and an essay task. Results demonstrated that participants with higher prior knowledge outperformed those with lower knowledge across comprehen- sion assessments. Critically, self-explanation was associated with higher performance on across-text inference verification items, with no corresponding effect for within-text inferences. Computational linguistic analyses revealed that self-explanation was associated with increased cohesion in readers' constructed responses across lexical, semantic, and connective-based in- dices. Together, these findings suggest that self-explanation promotes integration during multiple-document comprehension by strengthening inferential coherence-building mechanisms, and that across-text integration represents a key challenge that benefits from strategic instructional support.

  • What Do You Think I Think? Accounting for Human Beliefs Using Second-Order Theory of Mind

    Discrepancies between an agent's actual knowledge and what a person thinks the agent knows can hinder interactions. If an agent could detect such discrepancies, it could provide feedback to account for them and improve current and future interactions. Using the I-POMDP as a framework for a second-order Theory of Mind (ToM-2), this work endows an agent with the ability to model the evolution of a person's erroneous beliefs about an agent and the cognitive biases and heuristics (CBH) from which they arise. In doing so, the agent can detect when CBH might be at play during an interaction and adaptively generate feedback that accounts for them. An in-person user study shows how a ToM-2 learner can account for the effects of a teacher's CBH to significantly improve the informativeness of teacher actions, and subjective results suggest people find the ToM-2 learner's feedback more useful.

  • Coherence and Bayesian Reasoning in Legal Judgment: How Verdicts Reshape Evidence Evaluation

    Coherence-based theories propose that decision-makers reinterpret evidence to align with emerging conclusions. Prior research on legal decision-making demonstrated coherence shifts in ratings of agreement, but it remains unclear whether similar effects can be found using normative measures. Across three studies, participants evaluated prosecution and defense arguments in a mock burglary case before and after rendering verdicts. In Study 1, participants exhibited clear coherence shifts: guilty verdicts associated with increased agreement with prosecution arguments and decreased agreement with defense arguments, whereas not-guilty verdicts produced the opposite pattern. In Studies 2a–b, the shift of likelihood-based measures of arguments strength across testing periods did not align with verdicts. A constraint satisfaction model predicted mean shifts from pre- to post-verdict evaluations, and a Bayesian network using participants' post-verdict probability estimates closely approximated empirical guilt judgments. It appears that verdicts and post-verdict evaluations of evidence are consistent with Bayesian inference while driven by coherence.

  • First steps in simulating semantic and phonological impairments in aphasia with a computational model of speech processing

    EARSHOT is a model of human speech processing that learns to map real speech to semantic representations (Magnuson et al., 2020). We report first steps towards simulating aspects of aphasia with EARSHOT by damaging increasing proportions of randomly selected weights in different model layers. Although the model is purely receptive, we simulated naming/identification tasks by presenting spoken words to damaged models and measuring the cosine similarity of the output to every word in the lexicon, with the model's naming/identification response operationalized as the word with highest cosine similarity. For errors, we evaluated phonetic and semantic similarity of the response to the target vector. We were interested in how robust EARSHOT would be to damage, and whether we might observe systematic patterns of phonetic vs. semantic deficits following damage to different model components. As expected, lower-level damage impaired phonology most, and higher-level damage most affected semantics. Further development could turn EARSHOT into a valuable tool for enhancing understanding of aphasia.

  • Estimating Task Representations in a Multidimensional Bandit Task: A Method for Revealing Meta-level Learning

    Humans must learn not only which options pay off, but also which features are worth learning about. Yet such meta-level task representations are difficult to observe directly. We propose an inverse inference approach for estimating an individual's hypothesis space over feature dimensions in a multidimensional bandit task. Building on a model previously supported when the hypothesis space is known, we instantiate the model for each candidate hypothesis space and identify the space that best predicts each participant's choice sequence using hierarchical Bayesian estimation and WAIC-based model comparison. In an online experiment, participants' choice accuracy improved both within games and across games, suggesting adaptation at multiple timescales. Moreover, the inferred spaces shifted on average toward the true task structure. These results indicate that changes in task representations are associated with improvements in structured reward learning, and provide a measurement framework for developing process-level models of meta-level learning.

  • Less Talk, More Code: Practice-Based Instruction Improves Programming Skill Acquisition

    Introductory computer science education often emphasizes watching experts write and explain code. We investigate whether practice-based instruction improves programming skills relative to traditional lectures and which features of practice make it effective. In this preregistered experiment (N = 250), we compared three approaches for an introductory programming lesson: watching a video, code tracing visualizations, and writing code with immediate, AI-generated feedback. Participants who engaged in practice-based instruction performed significantly better than those who watched the video on a novel code-generation test, with the highest performance among learners who practiced writing code. Practice-based instruction also resulted in reduced distraction and increased interest. Whereas video and code tracing conditions led to overconfidence about programming abilities, code writing led to underconfident self-assessments. These findings suggest that active practice, particularly generative practice with feedback, supports learning with advantages in cognitive processing, metacognition, and interest in programming.

  • Elements of the Neurobiology of Language: A Neurostimulation Model

    Models of the neurobiology of language often rely on correlative neuroimaging techniques and ambiguous terminology derived from linguistic theory (e.g., 'semantics') rather than regions' functional profiles. The brain's cytoarchitecture and the capacity to functionally affect cognition via neurostimulation imply greater precision may be possible. We systematically reviewed 221 of tES, TMS, and DES papers modulating linguistic processes and qualitatively analysed by-paper outcomes to infer the specific subprocess actually modulated, which we term an element. We uncover 608 distinct elements of the neurobiology of language from 22 different languages and nearly 6,300 participants. These elements apply over a 7-level hierarchy of process specificity, help to redefine the operable terms of language neuroscience, and enable generalised functional characterisation of key nodes of the language network. The model is available for scientific use via an online database (https://www.language-elements.org/) and has clinical applications for planning awake craniotomies with neuro-oncology patients.

  • Testing Temporal Binding and Memory Accounts: A Comparison of Interval and Line Length Judgments

    Temporal binding (TB), the perceived compression in time between an action and its effect versus an observed action and its effect, has been viewed as an implicit measure of sense of agency (SoA). SoA refers to the feeling of control over our actions and their outcomes. Recent work has raised methodological and theoretical concerns about the TB paradigm. Most problematically, the TB effect may be explained by a regression-to-the mean pattern commonly observed as a memory bias, rather than as an effect related to SoA. We tested this possibility using two structurally equivalent tasks. Participants completed a standard TB task and a well-established memory task (line length estimation) that we equated on all aspects to allow for direct comparison. We did not replicate the TB effect; instead, both the time interval and line length conditions showed classic regression-to-the-mean memory effects, consistent with predictions from the Category Adjustment Model (Huttenlocher, 2000).

  • From idiosyncratic narratives to social scripts: Mapping the semantic space of self-descriptions in children and adults

    Self-description is a routine part of daily life, but choosing what to reveal about ourselves is a cognitively sophisticated communicative act. How do children talk about themselves, and how do they differ from adults? To characterize developmental changes in self-description, we asked 3- to 4-year-old children and adults to share 12 things about themselves. We analyzed over 3,000 self-descriptions using qualitative annotations and sentence embeddings to quantify their semantic structure. Children generated small, idiosyncratic "islands" of self-description, sampling narrowly from a specific topic, whereas adults' responses spanned a broader semantic range, reflecting a more "scripted" approach. While preliminary and descriptive in nature, our work demonstrates how advances in natural language processing can be leveraged to quantify and visualize developmental changes in self-description. The results are consistent with the possibility that self-description evolves from idiosyncratic narratives to a normative social act, raising new questions about the cognitive underpinnings of this development.

  • Ability vs Adaptability: What Dynamics Reveal

    Stable trait estimates from sequential task performance fail to separate ability from adaptability: the real-time recalibration of cognitive control parameters that improves execution without acquiring new knowledge. This conflation biases ability estimates. We develop a tractable framework measuring adaptability through trial-by-trial performance dynamics in mathematical cognition tasks. Across large datasets (N>10,000), adaptability strongly predicts math achievement and is uncorrelated with traditional ability measures, establishing it as an independent construct critical for understanding individual differences in cognitive performance.

  • Crosslinguistic patterns of phrasal word order support efficient prediction

    Across languages, word order within phrases shows robust statistical regularities: dependents are typically ordered con- sistently by type, word order tends to be harmonic with the head near an edge, and more strongly associated dependents tend to occur closer to the head. Existing accounts often derive these properties from separate mechanisms or stipulations. I propose a unified efficiency-based explanation grounded in in- formation theory: the idea is that languages are organized so that it is possible to make good predictions using only small amounts of memory. Using simulations of nouns modified by multiple adjectives differing in association strength to the noun, I compare linguistically motivated ordering policies that vary in consistency, harmony, and locality. Across sampled source distributions, I find that memory requirements for prediction are minimal in orders that are simultaneously harmonic, consistent, and local. I analyze why these effects arise: consistency makes upcoming category structure predictable from small amounts of context, and harmony makes phrase boundaries easy to antici- pate. The results suggest that diverse crosslinguistic patterns of phrasal order may reflect a general pressure to make incremental prediction simple.

  • Semantic centrality captures key aspects of knowledge construction: Behavioral and neural evidence of learning from a STEM video lecture

    When learning from educational content, we must simultaneously remember individual pieces of information and integrate them into a coherent conceptual framework. Previous work has shown that semantic network structure predicts what people remember from narratives, but whether these same principles extend to academic learning remains unknown. In this study, participants watched a 15-minute physics lecture and completed free recall during fMRI scanning. Using Sentence-BERT embeddings to quantify semantic centrality, we found that central information was better recalled and, critically, participants who recalled more units of central information demonstrated better conceptual understanding as rated by human raters (r = .73). In a voxel-wise encoding analysis, we modeled brain responses during lecture viewing using the same SBERT features, and this brain-model mapping predicted behavioral outcomes in selective cortical regions. These findings suggest that the embedding space of large language models reflects relevant aspects of how our minds organize newly learned conceptual information.

  • Human-Like Anaphor Resolution in Large Language Models

    Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors.The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.

  • A Computational Model of Action Selection and Action Specification in the Basal Ganglia

    It has been proposed that the basal ganglia is involved in selecting which action is performed, while the motor cortex specifies how the selected action is carried out. However, recent electrophysiological evidence from a reach-to-pull task challenges this view by showing that the motor cortex and the dorsal striatum in the basal ganglia jointly specify continuous action movement parameters, such as reach angle and pull force. This finding implicates the basal ganglia in both action selection and action specification, creating a gap for a mechanistic model that unifies both functions in a single, grounded framework. We address this gap with a biologically plausible model of the basal ganglia that leverages dynamic neural fields applied to high-dimensional representations of action-salience distributions. The dynamic neural field converts high-entropy representations into low-entropy representations without altering the underlying action-salience distribution's peak---a property previously shown to support accurate readout of continuous action parameters (e.g., movement speed or force). Our model is an extension that applies the same dynamic neural field transformation across multiple actions so that we can simultaneously (i) resolve which action is chosen (e.g., lever pull vs. button press) and (ii) specify the continuous action parameter value with which the chosen action is to be executed (e.g., force or speed). In our simulations patterned on the reach-to-pull paradigm, the basal ganglia model performs both discrete action selection and action specification of a continuous action parameter in a single operation, offering a concise mechanism for how basal ganglia-cortex circuits decide what to do and how to do it.

  • Trust Issues: Social Learning Under Misaligned Goals

    Computational models of social learning often assume learners and demonstrators share identical or at least positively correlated goals. Yet this assumption limits applications to real-world scenarios, where preferences may be misaligned or even opposed. We address this gap by extending the socially correlated bandit task to settings where agents need to learn when social information is positively correlated, uncorrelated, or negatively correlated, analogous to learning whom to trust or distrust. We introduce Social Correlation–Adjusted LEarning (SCALE), a multi-output Gaussian Process model that learns the covariance structure between agents' preferences. Using simulations, we characterize the model's performance across social environments and outline a path toward agents that can dynamically infer social correlations from experience. Our model allows us to reframe prior experimental observations, and lays the groundwork for future experimental work on the integration of preferences into individual decision-making.

  • Shallow by strategy: Heuristic parsing in Mandarin-Chinese-speaking learners' comprehension of Korean suffixal passives

    This study examines how L1-Mandarin-Chinese L2-Korean learners comprehend constructions expressing transitivity within a good-enough processing framework. Using picture selection with webcam eye-tracking, we compare learners with native speakers in two studies. Study 1 crosses voice (active vs. passive), word order (canonical vs. scrambled), and verb position (verb-final vs. verb-initial). Learners are accurate for canonical actives, but perform below chance on passives and show reduced target looks in critical regions, consistent with a persistent agent-first bias and limited use of case-marking and passive morphology. Study 2 obscures case markers to isolate verbal morphology; learners favour agent-first interpretations and show only weak sensitivity to passive morphology. Across two studies, proficiency predicts more target-like performance, especially in non-canonical configurations. These findings highlight the role of cue alignment/(re)weighting in L2-acquisition pathways and a proficiency-mediated shift towards more algorithmic parsing, suggesting the dynamic and adaptive nature of L2 good-enough processing.

  • Toward Fair and Diverse Pedagogical Generations: Quantifying Educational Epistemic Bias in Text-to-Image Models

    Text-to-image (T2I) models are increasingly used to generate educational visuals, from textbook illustrations and lecture slides to classroom scenes and institutional marketing materials. However, these models do not neutrally depict education. They often encode stereotyped views of what education should look like, and such patterns can recur across generations and gradually solidify into taken-for-granted visual common sense. Building on cognitive theories, we introduce the notion of Educational Epistemic Bias to describe how T2I models narrow rich educational ideas into a small set of recurring visual patterns. We define this construct along four dimensions and design an education-specific benchmark of 60 prompts that span learning environments, activities, power relations, and roles. We apply this benchmark to nine mainstream T2I models, yielding 2,160 images. To quantify bias, we propose the Educational Epistemic Biases Quantifier (EEBQ), a VLM-based evaluation framework that includes two image-quality metrics and four metrics targeting DEI-related aspects of the images. Our analysis reveals systematic educational epistemic biases across all models. Together, the benchmark and EEBQ offer a concrete way to examine how generative models visually construct education and to inform more careful use of such systems in educational settings.

  • Recovering Event Probabilities from Large Language Model Embeddings via Axiomatic Constraints

    Rational decision-making under uncertainty requires coherent degrees of belief in events. However, event probabilities generated by Large Language Models (LLMs) have been shown to exhibit incoherence, violating the axioms of probability theory. This raises the question of whether coherent event probabilities can be recovered from the embeddings used by the models. If so, those derived probabilities could be used as more accurate estimates in events involving uncertainty, and the embeddings would become more interpretable. To explore this question, we propose enforcing axiomatic constraints, such as the additive rule of probability theory, in the latent space learned by an extended variational autoencoder (VAE) applied to LLM embeddings. This approach enables event probabilities to naturally emerge in the latent space as the VAE learns to both reconstruct the original embeddings and predict the embeddings of semantically related events. We evaluate our method on complementary events (i.e., event A and its complement, event not-A), where the true probabilities of the two events must sum to 1. Experiment results on open-weight language models demonstrate that probabilities recovered from embeddings exhibit greater coherence than those directly generated by the models. Moreover, the recovered probabilities align closely with the true probabilities, while the latent space of the VAE provides interpretable structures.

  • When to Think Deep? Resource-Rational Metacognitive Control for Adaptive Inference in LLM-Based Legal QA

    Legal question-answering systems based on large language models (LLMs) encounter reasoning demands that vary significantly across queries. Consequently, static retrieval-and-reasoning pipelines frequently result in inefficient resource allocation. In this paper, we introduce ALEX, a cost-aware adaptive inference architecture grounded in resource-rational accounts of metacognitive control. The proposed framework implements a metacognitive controller that monitors query complexity via multi-dimensional linguistic and confidence cues, and subsequently routes tasks to hierarchical workflows within an expected utility framework. Experimental results on legal benchmarks demonstrate that ALEX achieves superior accuracy while maintaining an efficient trade-off between accuracy and effort, consistent with the principles of bounded rationality. Furthermore, analyses of routing behavior and error correction driven by verification reveal how metacognitive monitoring can mitigate resource misallocation in LLM-based legal QA.

  • From Neural Control to Cognitive Explanation: Closed-Loop Negative Feedback and Marr's Levels in YIN-CBEHA

    Many cognitive theories explain behavior using linear input–output models grounded in normatively specified tasks. However, these approaches remain weakly grounded in biological organization and real-time organism–environment interaction. This theoretical analysis paper examines Yin's control-theoretic extension of Perceptual Control Theory (YIN-CBEHA) using Marr's levels of analysis as evaluative criteria. In YIN-CBEHA, behavior is not identified with motor output but with the hierarchical control of perceptual variables through closed-loop negative feedback. We show how this framework satisfies Marr's three-level demands within a single control architecture. Converging evidence from neurophysiology and robotic construction supports the functional sufficiency of hierarchical feedback control for posture, movement, and spatial regulation under real-world constraints. By constraining admissible computational problems through continuous-time feedback, YIN-CBEHA reduces explanatory underdetermination and improves biological plausibility. While its strongest empirical support currently lies at intermediate sensorimotor levels, the framework provides a tractable foundation for extending control-theoretic explanation toward higher cognition.

  • Joint Human-AI decision making and sense of agency

    In this study, we address shared Human-AI decision making through the lens of agency: what is the feeling of control experienced by an operator when making a decision that is shared with an AI-based artificial agent? Using an optimization trajectory task, we explored how the level of autonomy and the explainability of the decision support system impact the human operator's feelings of agency, responsibility and confidence. Results indicated that the level of automation modulate the sense of agency, but also the confidence and feeling of responsibility. More interestingly, our results also indicated that explanations could increase the number of mind change, but also the feelings of agency, confidence and responsibility during such mind change. These results are discussed in term of mechanisms underlying the development of a sense of agency and aims to provide tools and models to explore the notion of 'meaningful control' during our interactions with AI-enabled systems.

  • Effects of Egocentric and Exocentric Strategies on Reciprocity: Agent-Based Study Using Ultimatum Game

    This study investigates the formation of reciprocal altruism in human-agent interactions using the ultimatum game, focusing on how people infer others' action-generation principles. We examine the effects of different agent strategies on human rejection behavior and agent impression. An experiment is conducted where two factors are manipulated in the agent's strategy: the presence of a selfish, egocentric strategy based on the agent's internal policy; and an altruistic, exocentric strategy that adjusts to the participant's feedback. The results show that participants rejected proposals less frequently when the agent adopted exocentric/non-egocentric strategies. Furthermore, the exocentric strategy increased the perception of safety toward the agent, whereas the egocentric strategy decreased the perception of intelligence. These findings suggest that cooperative behavior and social impressions are shaped not only by outcome distributions but also by inferences about latent policy parameters, such as responsiveness to social signals and rigidity of internal dynamics.

  • Referring Effectively to a Group of Objects

    In communication, we often need to refer to multiple objects. We have several possibilities: enumerating the objects one by one or referring to a shared feature. This is the first study to systematically compare different ways of referring to groups of objects in terms of response times, accuracy, and memory encoding. In a within-subjects design, participants listened to instructions which either used a shared feature or enumerated coordinates of individual objects. We find that there is no one way of referring that is universally more effective, but the optimal way of referring depends on the type of common feature and the number of objects. When the shared feature is prominent, such as color, and multiple objects need to be identified, referring to it results in shorter identification times and fewer errors. In contrast, when the shared feature is less salient, such as pattern, referring by location tends to be better.

  • How communicatively optimal are exact numeral systems? Once more on lexicon size and morphosyntactic complexity

    Recent research argues that exact recursive numeral systems optimize communicative efficiency by balancing a tradeoff between the size of the numeral lexicon and the average morphosyntactic complexity (roughly length in morphemes) of numeral terms. We argue that previous studies have not characterized the data in a fashion that accounts for the degree of complexity languages display. Using data from 52 genetically diverse languages and an annotation scheme distinguishing between predictable and unpredictable allomorphy (formal variation), we show that many of the world's languages are decisively less efficient than one would expect. We discuss the implications of our findings for the study of numeral systems and linguistic evolution more generally.

  • Do Large Language Models Show Negation Bias? A Replication of Beukeboom et al. (2010)

    LLM bias research has largely focused on content-level associations, overlooking biases arising from linguistic form and pragmatic choice. One such phenomenon is the negation bias, where negated expressions are preferentially used for stereotype-inconsistent behavior and trigger systematic inferences. While LLMs' handling of negation in semantic tasks is well studied, it is unclear whether they reproduce these pragmatic inferences in social contexts. We replicate Beukeboom et al. (2010) with Mistral-7B-Instruct-v0.2, varying the valence and polarity of trait descriptions. The model inferred more negative expectations from negated negative traits (e.g. not stupid) and more positive expectations from negated positive traits (e.g. not smart), and showed stronger situational but weaker dispositional attributions. These results mirror human negation-bias patterns, suggesting that instruction-tuned models can produce pragmatic-like bias in judgment tasks. The findings underscore the need to extend LLM bias evaluation beyond content to linguistic form.

  • How Do People Quantify Naturally: Evidence from Mandarin Picture Description

    Quantification is a fundamental component of everyday language use, yet little is known about how speakers decide whether and how to quantify in naturalistic production. We investigate quantification in Mandarin Chinese using a picture-based elicited description task in which speakers freely described scenes containing multiple objects, without explicit instructions to count or quantify. Across both spoken and written modalities, we examine three aspects of quantification: whether speakers choose to quantify at all, how precise their quantification is, and which quantificational strategies they adopt. Results show that object numerosity, animacy, and production modality systematically shape quantificational behaviour. In particular, increasing numerosity reduces both the likelihood and the precision of quantification, while animate referents and modality selectively modulate strategy choice. This study demonstrates how quantification can be examined under unconstrained production conditions and provides a naturalistic dataset for further analyses of quantity expression in language production.

  • Relating Online Word Predictions and Offline Probability Judgments: Partial Evidence for Shared Probabilistic Processes

    Probabilistic reasoning has been proposed to underlie both word prediction in language comprehension and probability judgment in decision-making, yet the connection between the two cognitive processes remains unclear. This study investigates whether word predictions and probability judgments rely on shared probabilistic processes by examining unpacking effects. An offline probability judgment task and an online self-paced reading task were used to test probability shifts in each task, and further analyses examined the relationship between measures across tasks. Results show that atypical unpacking reliably led to lower probability judgments and slower reading times, whereas typical unpacking didn't consistently lead to higher probability judgments or faster reading times. Interestingly, item-level changes in probability judgments were correlated with changes in reading times only for atypical unpackings. These findings provide partial evidence that word prediction and probability judgment draw on a shared foundation of probabilistic reasoning.

  • Emergent social transmission of model-based representations without inference

    How do people acquire rich, flexible knowledge about their environment from others despite limited cognitive capacity? Humans are often thought to rely on computationally costly mentalizing, such as inferring others' beliefs. In contrast, cultural evolution emphasizes that behavioral transmission can be supported by simple social cues. Using reinforcement learning simulations, we show how minimal social learning can indirectly transmit higher-level representations. We simulate a naive agent searching for rewards in a reconfigurable environment, learning either alone or by observing an expert—crucially, without inferring mental states. Instead, the learner heuristically selects actions or boosts value representations based on observed actions. Our results demonstrate that these cues bias the learner's experience, causing its representation to converge toward the expert's. Model-based learners benefit most from social exposure, showing faster learning and more expert-like representations. These findings show how cultural transmission can arise from simple, non-mentalizing processes exploiting asocial learning mechanisms.

  • Meaning-Making as Symmetry-Breaking: Spectral Thermodynamics of Emergent Communication

    How do interlocutors achieve mutual understanding when communicative mappings are not given in advance? We propose that meaning-making is a symmetry-breaking constructive process: through interaction, dyads compress an open space of communicative possibilities into a low-dimensional set of shared alternatives, namely dyad-specific signal-referent relations that enable coordination. We test this account in 69 dyads using a real-time semiotic game in which participants invent and stabilize dyad-specific communicative conventions, modeling signal–referent relations with spectral graph thermodynamics. Successful dyads exhibited dimensional compression of their signaling space: intrinsic dimensionality was reduced sixfold relative to random networks, reflecting the emergence of stabilized dyad-specific conventions. These dyads also showed lower thermodynamic efficiency, indicating that constructing shared signal–referent structure carries an entropy cost. Thermodynamic and geometric measures converged at a mesoscale diffusion regime, where relational constraints operate. These findings suggest that mutual understanding depends on actively constructing a shared, low-dimensional communicative structure through interaction.

  • Hán Dān Xué Bú (Mimicry) or Qīng Chū Yú Lán (Mastery)? A Cognitive Perspective on Reasoning Distillation in Large Language Models

    Recent Large Reasoning Models trained via reinforcement learning exhibit a "natural" alignment with human cognitive costs. However, we show that reasoning distillation via Supervised Fine-Tuning (SFT) fails to transmit this cognitive logic, leading to a "Cargo Cult" where students only mimic length. Testing the Hán Dān Xué Bú (Superficial Mimicry) hypothesis across 14 models, we identify a "Functional Alignment Collapse": while teacher models mirror human difficulty scaling (r = 0.64), distilled students significantly degrade this alignment (r = 0.34). Crucially, they exhibit "Negative Transfer," dropping below their own pre-distillation baselines. Our analysis reveals a "Linear Inflation Law" where students apply a constant verbosity multiplier (≈ 2.44) regardless of complexity. Consequently, distillation decouples computational cost from cognitive demand, revealing that human-like cognition is an emergent property of active reinforcement, not passive imitation.

  • Cognitive Modeling of Shogi: Effects of Relative Representations and Resource Constraints

    We model shogi move selection in ACT-R to test how expertise depends on representation and cognitive resource constraints. The model integrates visual exploration, a capacity-limited imaginal stack, and declarative memory retrieval. Beyond absolute board encoding, we implement king-centered relative representations (3$\times$3 local patterns) intended for endgame reasoning. Simulations crossed (i) relative representations on/off and (ii) imaginal capacity (9 vs. 18), while expertise was manipulated by storing expert vs. novice game records. Expert advantage was conditional: it was stable when relative representations were available, but collapsed (and sometimes reversed) when relative representations were absent under low capacity, accompanied by a sharp drop in retrieval use and earlier convergence of game states. Larger capacity partially compensated for missing relative representations. These results imply that expertise in this domain emerges from representation $\times$ resource coupling, not experience alone.

  • Acquisition of words and case markers in a novel language: Artificial language learning by Korean and English speakers

    The acquisition of second language morphosyntactic features, such as case systems, is particularly challenging for adult learners, and the extent to which these difficulties are shaped by learners' first language remains a core question in theories of language acquisition. The present study explores the acquisition of lexical semantics and morphosyntax of a case-marking artificial language, Katopu, in an artificial language learning task with Korean-speaking and English-speaking adults. The results showed that both groups were able to acquire Katopu through the exposure to input and feedback, regardless of L1–L2 (dis)similarities and in the absence of explicit instruction, exhibiting comparable learning trajectories in accuracy and response time. Response time data further suggested that the two groups shared a similar processing pattern that semantic and syntactic knowledge were processed using separate mechanisms. The study highlights that sufficient input and feedback can effectively support language acquisition regardless of learners' language background.

  • LLMs Electrified: Early and deep layers differentially correlate with the N400 and P600 in language comprehension

    Recent research has examined the extent to which large language models (LLMs) can model the temporal dynamics of online language comprehension, as indexed by the N400 and P600 components of the event-related potential (ERP) signal. To better understand whether the internal representations of LLMs are consistent with distinct stages of comprehension, we employ representational similarity analysis (RSA) on a German ERP study, that found the N400 to be sensitive to association and expectancy, and the P600 to be sensitive to expectancy alone. We find that earlier layers show a stronger correlation with association, whereas deeper layers are more strongly correlated to expectancy. Similarly, correlations to the N400 are stronger at earlier and intermediate layers, consistent with its sensitivity to association, while correlations with both the N400 and P600 continue to increase in deeper layers, reflecting the influence of expectancy on both components. These results are consistent with independent stages of processing proposed by neurocognitive theories, such as Retrieval-Integration theory, suggesting that LLMs may contribute to our understanding of the temporal dynamics of language comprehension - as indexed by ERPs - at a more mechanistic level.

  • From Insight to Outsight: When Solutions Are Enacted before They Are Recognized

    Creative breakthroughs are typically described as sudden mental events, the "Aha!" moment.. Yet this formulation may fundamentally mischaracterize the process. Rather than solutions being discovered in the mind then implemented in the world, they may emerge through engagement with manipulable objects. We report an experiment comparing interactive and non-interactive conditions across two insight problems (Triangle of Coins, 10 Circles). Survival analyses confirmed that interactivity enhanced solution rates and reduced latencies for both problems. More critically, frame-by-frame multimodal coding of the video data (via ELAN) revealed the temporal coordination between speech, action, and evolving solution prototypes. Detailed analyses of two solution episodes demonstrate how restructuring can occur in the physical configuration of the problem rather than in the mind. These cases illustrate "outsight": solutions recognized after being enacted. By making the solution process visible, multimodal behavioural coding produces qualitative data that challenge cognitivist assumptions about how creative breakthroughs happen.

  • Fact or fiction? How repeated claims influence belief updating

    Individuals often encounter the same claims from multiple sources in complex informational environments. Whether such repetition is treated as independent evidence or as redundant testimony depends on assumptions about source dependence. The present study investigates how prior beliefs about the claims shape perceptions of source dependence and how these perceptions influence belief updating. Using a preregistered experimental design, participants evaluated repeated claims presented in vignettes about an online forum. Participants reported both belief change and perceived dependence among sources. We find that lower prior belief in a claim leads to stronger perceptions of source dependence, indicating that prior beliefs shape not only initial judgments but also inferences about the structure of the informational environment. Participants also tended to treat dependence as a global property applying to several sources in the environment. Together, these findings suggest that belief updating reflects rational simplification of informational structure, with important implications for the persistence of belief polarisation.

  • Domain-Adaptive Transfer Learning with Recurrent Neural Networks for Cross-Subject P300 Speller Classification

    Brain-computer interface (BCI) research has advanced significantly, yet cross-subject P300 decoding remains challenging due to highly variable EEG signals. To address this, we propose a novel domain adaptation framework integrating transfer learning and recurrent neural networks (RNNs). The framework incorporates a gated recurrent unit (GRU) layer to effectively capture temporal dependencies of EEG signals, thereby mitigating temporal misalignment issues. By calculating transfer loss and applying time-step weighting, the framework enhances classification. Furthermore, feature-space adaptive alignment is employed to reduce inter-subject variability, lowering subject dependency and enabling accurate cross-subject character recognition. Experimental results demonstrate that the proposed method achieves an average character recognition accuracy of 92.38%, significantly surpassing traditional P300 recognition algorithms. This indicates that the method effectively mitigates the impact of temporal misalignment in EEG data as well as the high subject-dependence of BCI systems.

  • When is one confession better than two? Rational and observed belief updating under competing explanations

    A central challenge in belief updating is how individuals evaluate evidence when multiple competing hypotheses are plausible. Although prior work shows that belief updating can align with Bayesian principles, this evidence typically involves intuitively directional cases. It remains unclear whether similar rational trajectories are observed in counter-intuitive settings, where accumulating supportive evidence can increase the plausibility of an alternative explanation. In the present study, participants evaluated a scenario in which five sequentially presented confessions were diagnostic either of collective guilt or of coercive interrogation. Early confessions led to increases in belief in both guilt and force, indicating non-zero sum updating. However, contrary to a priori model predictions, participants did not exhibit the non-monotonic pattern whereby belief in force overtakes belief in guilt. Participant-specific Bayesian Network models showed closer alignment with observed trajectories, yet revealed systematic under-weighting of force and over-weighting of guilt. Overall, belief updating reflected sequential evidence integration rather than full Bayesian re-computation.

  • Mathematical Cognition and its Neural Correlates During Development: A Scoping Review with Systematic Search

    Mathematics learning begins in schooling and plays a role in children's cognitive development, predicting academic achievement and professional success. Its development is influenced by linguistic, metacognitive, emotional, and cognitive factors. Cross-sectional neuroimaging studies have highlighted the involvement of the intraparietal sulcus and prefrontal cortex in mathematical tasks. However, longitudinal studies capituring the neural correlates of mathematical development in school-aged children remain scarce and heterogeneous. To provide an overview of current evidence on longitudinal studies of mathematical cognition with neural correlates, we conducted a review with a systematic search strategy. Nineteen studies were identified, revealing mathematics as a dynamic, multi-component system that evolves from a structural foundation into a functionally specialized network. Findings emphasize the central role of parietal and frontal regions, with contributions from temporal, occipital, and subcortical areas, while also underscoring the lack of naturalistic studies and diverse sociocultural samples.

  • Age Cohort Variability in Conceptual Complexity and Concreteness: The Case of Ecological and Technological Concepts

    Concepts are flexible and vary depending on multiple factors. Here, we investigate age cohort conceptual variations using the two domains of ecology (e.g., "deforestation") and technology (e.g., "web") as case examples. 320 Italian older adults (>64 years) evaluated 50 concepts per domain across 39 semantic dimensions. Their ratings were compared with those of younger adults (18–35 years) reported in Falcinelli et al. (2024a). Results revealed clear generational differences: representations were more complex and concrete for ecological concepts in older adults, and for technological concepts in younger adults. Additionally, older adults showed greater personal experience with ecological topics, while younger adults reported higher familiarity with technology. Theoretically, our findings support evidence regarding differences in conceptualizations across age cohorts. Scientifically, the resulting O-TECo database provides the first semantic norms tailored to older adults. Societally, results highlight the importance of age tailored interventions in these two crucial fields.

  • Effects of linguistic context and reading abilities on processing of unknown words

    Encountering unknown words disrupts reading and requires readers to rely on context to infer meaning. We investigate how contextual expectations (science-themed vs. daily-life narratives) and reader literacy modulate processing of unknown words. Using pseudowords as controlled proxies for unfamiliar lexical items, we conducted an eye-tracking experiment (Experiment 1) and a self-paced reading study (Experiment 2). To investigate how participants infer the meaning of pseudowords, each pseudoword was followed by a semantically associated word in the subsequent sentence. Across both methods, pseudowords elicited longer reading times. Crucially, eye-tracking results revealed that higher-literacy readers showed greater sensitivity to the semantically associated word, particularly in science-themed contexts. While both groups of readers showed attempts at reconstruct meaning in an offline task, higher-literacy readers mentioned unknown words themselves more frequently, which can be taken as evidence for incidental vocabulary learning. Together, these findings suggest that literacy supports strategic, context-sensitive processing of unknown words.

  • A Rational Account of Categorization Based on Information Theory

    We present a new theory of categorization based on an information-theoretic rational analysis. To evaluate this theory, we investigate how well it can account for key findings from classic categorization experiments conducted by Hayes-Roth and Hayes-Roth (1977), Medin and Schaffer (1978), and Smith and Minda (1998). We find that it explains the human categorization behavior as well as (or better) than the independent cue and context models (Medin & Schaffer, 1978), the rational model of categorization (Anderson, 1991), and a hierarchical Dirichlet process model (Griffiths et al., 2007).

  • Face Pareidolia and the Implicit Bystander Effect: A Drift-Diffusion Model Analysis with Individual Differences in Autistic Traits

    The bystander effect, reduced helping in the presence of others, has been linked to automatic neural processes that inhibit action preparation. Recent work using the Drift-Diffusion Model (DDM) showed that human faces slow evidence accumulation during emergency classification tasks. We examined whether face pareidolia stimuli (objects in which illusory faces are perceived) produce similar effects, and whether autistic traits predict decision-making parameters. Participants (N = 65) classified scenes as safe or dangerous while viewing human faces, face-like objects, or non-face-like objects. Human faces significantly slowed drift rate, replicating prior findings; face pareidolia showed a similar but non-significant trend. In this non-clinical sample, Autism-Spectrum Quotient (AQ) scores predicted slower drift rates for both scene types, with modest effect sizes (r = .26 to .28), suggesting reduced excitatory efficiency with higher autistic traits. These findings extend the implicit bystander effect to individual differences and link autistic traits to evidence accumulation.

  • Mapping Verb Meaning in Space: Spatial Feature Norms for Rioplatense Spanish Verbs

    Understanding how language recruits spatial cognition to represent events and actions is central to theories of embodied cognition and cognitive linguistics. This study presents a normative dataset of spatial feature ratings for 75 verbs in Rioplatense Spanish, examining the relationship between image-schematic representations and syntactic congruence. A total of 639 native speakers completed an online Spatial Features Questionnaire for Verbal Semantics. Participants judged whether trajectory-based image schemas matched each verb under congruent (subject–verb–object) and incongruent (object–verb–subject) conditions. Chi-square analyses and proportion tests revealed higher acceptance rates for congruent configurations, indicating that syntactic order constrains spatial representations. Item-level analyses showed that a subset of verbs elicited stable spatial trajectory preferences under congruent conditions. These results provide a normative resource that characterizes structured verb–space associations and documents the role of syntactic structure in shaping spatial judgments, offering stimuli for future experimental and cross-linguistic research on embodiment and spatial grounding in Spanish verb processing.

  • When Correlation means Causation: Pragmatic Factors modulate Causal Implicatures in Decision-Making Contexts

    Correlational statements (e.g., "$X$ is associated with $Y$") can be interpreted as conveying a causal relationship. Recent work suggests that such interpretations may be legitimate pragmatic inferences, especially in decision-relevant contexts. Here, we investigate whether causal enrichment of correlational language affects memory recall and message passing, and how pragmatic features of the communicative context modulate it. Our findings support a pragmatic account of causal enrichment of correlational language.

  • Abstraction Modulates Judgments of Intentional Action

    This paper investigates how the different levels of hierarchical plans modulate judgments of intentional action. While previous research has focused on how moral valence influences intentionality judgments of side-effects (outcomes that were not a part of the agent's original plan), we focus on sub-events, which are candidate decompositions of higher-level goals that were a part of the agent's plan. We address two central questions: (1) does the impact of moral valence on intentionality judgments differ between sub-events and side-effects? and (2) does the level of abstraction at which a sub-event is defined affect judgments of its intentionality? Our results show that the effect of moral valence is largely neutralized for sub-events, which instead exhibit a decrease in perceived intentionality as they are decomposed into lower-level actions. This indicates that future models of intentional action should take into account the asymmetry in judgments between higher- and lower-level actions.

  • Syntactic Reanalysis Across the Adult Lifespan: Evidence From Eye Movements

    This study examines the influence of age-related cognitive decline on the time course of syntactic reanalysis in Dutch NP/S-coordination garden-path sentences using eye tracking. 110 native Dutch speakers (30-80 years) read ambiguous sentences and comma-disambiguated controls. Garden-path costs were strongest at the ambiguous noun and smaller at the disambiguating verb, with no spillover effects. To further investigate age-related variability, we related garden-path effects to individual differences in cognitive abilities. Principal Components Analysis identified two components reflecting working memory and cognitive flexibility. Together with reading experience and associative learning, we tested these predictors of garden-path costs. Working memory, reading experience and associative learning moderated early garden-path effects, consistent with efficient encoding and rapid updating upon encountering ambiguity. Cognitive flexibility emerged as the most consistent predictor of late garden-path costs, suggesting a role in flexibly managing competing parses and supporting syntactic reanalysis.

  • Comparing LLM and Human Responses to Human- and AI-labeled Partners During Naturalistic Conversation

    Large language models (LLMs) increasingly interact with humans and other AI systems, raising questions about whether they adjust behavior based on partner identity. In a companion study, humans showed behavioral differences when conversing with partners labeled as human versus AI. Here, we extend this investigation to LLMs. We simulated 2,000 conversations where GPT-3.5-turbo "participants" (N = 50) engaged with partners labeled as either human or AI. All partners were actually identical LLMs, isolating label effects from actual differences. We analyzed transcripts using linguistic measures paralleling human data. LLMs showed robust label effects: more questions, politeness, interpersonal discourse markers, and mental state language with human-labeled partners; more words and positive sentiment with AI-labeled partners. Hedging showed category-specific patterns. Only the discourse marker "like" showed similar patterns across the LLM and human studies. These divergent patterns suggest LLMs have learned partner-type associations, but their social behavior differs fundamentally from humans.

  • Visuospatial Perspective Taking in Multimodal Language Models

    As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks largely rely on text-based vignettes or static scene understanding, leaving visuospatial perspective-taking (VPT) underexplored. We procedurally generate large stimulus batteries for two evaluation tasks adapted from human studies: the Rotating Figure Task, probing perspective-taking across angular disparities, and the Director Task, assessing VPT in a referential communication paradigm. Across both tasks, MLMs show pronounced deficits in Level 2 VPT, with failure patterns indicating reliance on simple mirroring heuristics rather than genuine perspective transformations. These results expose critical limitations in current MLMs' ability to represent and reason about alternative perspectives, with implications for use in collaborative contexts.

  • Quantifier and implicature processing in contexts with full and partial information.

    When evaluating the truth of a sentence, we do not always have full information available. While there exists psycholinguistic research on how an uncertain speaker influences a listener's quantifier processing and interpretation, it remains difficult to disentangle a listener's preferred utterances from ones they merely tolerate as acceptable if uttered by someone else. In the presented research, we investigate how readers process assertable and non-assertable quantified statements in epistemically uncertain situations. We focus on the weak and strong scalar implicatures for the quantifier "some", which arise in such contexts. We show differential ERP effects depending on whether participants adopt the weak or strong pragmatic interpretation. We also show that partial-information contexts are associated with a late and sustained negativity effect, likely signalling epistemic context monitoring.

  • Training Children's Understanding of Numbers Using Compositional Cues

    Infants form exact representations of small quantities (1 to 4) but only approximate representations of larger quantities. How do children acquire exact large-number concepts? We tested whether training showing how sets of 5 or 6 can be decomposed into smaller sets increases children's understanding of five and six. We compared compositional training, using expressions such as "I have five pets: three dogs and two cats" paired with visual displays of the subsets, to training that drew attention to each object in succession, as in counting ("I have five pets, here's one, and another, and another, and another, and another."). Children who did not know the trained number word at the outset performed significantly better at test in the compositional condition. Thus, children can leverage their knowledge of small numbers to build representations of larger numbers.

  • Single-Neuron & LFP Representations of Fear Processing in the Human Brain

    Fear extinction involves forming inhibitory safety memories that compete with the original fear traces. We recorded intracranial microelectrode data from 24 epilepsy patients to investigate how neuronal activity supports these representational shifts. During acquisition, amygdala neurons showed increased firing rates (475–525 ms) for aversive outcomes (n = 19). Oscillatory analyses (n = 24) revealed that the entorhinal cortex exhibited high-gamma enhancement (52–70 Hz and 74–88 Hz; both 0.80–1.10 s) following aversive outcome delivery during acquisition, while extinction was marked by low-frequency suppression (2–10 Hz, 1.15–1.50 s) following reinforcement. When combining acquisition and extinction phases, the hippocampus showed early high-gamma responses (76–98 Hz, 0.25–0.40 s) to aversive outcomes. Critically, hippocampal gamma power (38–62 Hz, 0.30–0.40 s) during extinction distinguished stimuli retaining threat value (CS++) from those undergoing extinction (CS+-), with the latter becoming indistinguishable from always-safe cues (CS--). Our results provide a cellular foundation for how the human mind adjudicates between competing memories and updates mental models through sensory evidence and threat evaluation.

  • Belief Updating: Utilizing a New Paradigm to Separate Evidence Evaluation and Memory Dynamics

    People often judge claims using cues like familiarity and source credibility, but later judgments may also be shaped by memory for what they previously believed. We used a three-phase task to separate these influences. Prolific participants (N = 130) completed 30 pairs of factual statements across five domains. In Section 1, they rated belief in an initial statement (S1) and their familiarity with it. In Section 2, they rated belief in a conceptually related evidence statement (S2) paired with a source and judged source credibility and statement familiarity; S2 either supported or refuted S1. In Section 3, they reported what they thought they answered for S1 earlier (belief-memory) and then re-rated S1. Mixed-effects models showed that credibility and S2 familiarity predicted Section-2 belief, while belief-memory and final belief about S1 were associated with both initial belief and the implications of S2 evaluation.

  • Modeling Context-Sensitive Effects of Inequity on Reinforcement Learning

    People are sensitive to unequal outcomes. Recent studies show that people's ability to learn action-reward mappings depends on the percentage of the reward amount they receive (compared to another person), even when such information does not change which actions are the most rewarding for the learner. However, questions remain about the degree to which this sensitivity to inequity depends on the broader inequity context. We report results from two additional experiments in which participants had the opportunity to learn action-reward mappings, with only part of the reward going to the participant. Crucially, we manipulated whether participants encountered both advantageous and disadvantageous inequity, or only one type of inequity, within the same learning block. We found that the effect of inequity differed between these two designs, suggesting that inequity affects learning in at least partly a context-dependent way. We formalized context-dependency into a new computational model which outperformed its context-independent counterpart.

  • Comparing the Likelihood of Producing Distant Analogs Through Retrieval versus Invention

    Research has long shown that retrieving semantically distant analogs from long-term memory is difficult, whereas retrieving near analogs is relatively easy, a limitation viewed as a major constraint on cross-domain transfer. An alternative route may lie in the invention of analogs. We report two experiments comparing the effects of surface similarity on analog retrieval and analog invention. In Experiment 1, using a production paradigm, participants were presented with simple events drawn from familiar schemas. In the retrieval condition, they recalled analogous events; in the invention condition, they generated them. Results showed that retrieving distant analogs was difficult, whereas inventing them was relatively easy. Experiment 2 replicated this pattern using materials with larger and more complex relational structures and a hybrid paradigm that controlled the availability of intra- and inter-domain analogs in memory, while enabling clearer discrimination between retrieved and invented analogs.

  • Evidence for Representational Shifts in the Learning of Chinese Characters: An Investigation using Machine Vision Techniques

    How does the visual representation of Chinese characters change with expertise? Fluent readers are known to perceive characters in terms of radicals and structural configurations, but the representational basis of perception before such knowledge is acquired remains unclear. We test whether novices' similarity judgments can be explained by generic shape computations from computer vision. Participants with three expertise levels rated pairwise visual similarity of Chinese characters. From pixelated character silhouettes, we computed 109 descriptors spanning multiple families of low-level shape statistics and derived pairwise distances. Using representational similarity analysis and Random Forest prediction, we compared descriptor distances with human judgments. Shape descriptors predicted novices' judgments strongly and learners' moderately but failed to predict fluent readers'. These findings establish a representational bridge between computer vision techniques and human visual cognition, suggesting that the generic shape geometry processing may form the starting point of visual similarity judgments before expertise introduces higher-level structure.

  • LLMs Struggle With Negation, but so do Humans - a Two-step Simulation Approach

    While substantial research has addressed the challenges of negation processing in humans, the parallels between human cognitive models and large language models (LLMs) in this regard remain less explored. This paper investigates how GPT-4o processes negation, drawing comparisons to human negation processing, particularly the two-step model of negation. In two experiments using image-sentence pairs, we assess the model's ability to handle affirmative and negated sentences. Our findings reveal that, like humans, GPT-4o struggles more with negation, exhibiting higher rates of incorrect inferences after negation, and a systematic bias toward completions associated with the negated state of affairs. This study highlights the qualitative similarities between AI and human processing of negation, offering insights into the limitations of current LLMs and suggesting future directions for improving their cognitive alignment with more complex linguistic constructs.

  • Is a Cactus More Animate than a Wildfire? Insights from the Implicit Association Task

    Animacy is widely thought to be conceptualized as a gradient, with humans on one end and artifacts (e.g. towel, fork) on the other. Cognitive scientists lack a clear understanding of the linguistic/conceptual organization of entities in the middle of the continuum (e.g., trees, fire, a river). Progress in this domain is constrained by over-reliance on aliveness judgments (e.g., "is a river alive?") as a way of operationalizing animacy. We present a novel method for studying animacy, an adaptation of the Implicit Association Task (IAT). While plants are judged to be alive more often than natural abiotic entities such as fire, our IAT suggests that U.S. English speakers conceptualize plants and natural abiotic entities as having similar levels of animacy. These findings highlight the need for a more comprehensive theory of how concepts of animacy are reflected in different ways across linguistic contexts.

  • Aging, Misinformation, and Eyewitness Memory: Accuracy and Confidence for Central and Peripheral Items

    Eyewitness memory is susceptible to distortion, especially among aging populations. Although we know that memory deteriorates as an individual becomes older, research on the interaction between aging, misinformation, and eyewitness memory remains limited. In this study, 42 young adults (age = 18-30) and 50 older adults (age = 65-78) first watched a short video clip depicting various central and peripheral items and read a narrative text that described the video but included false information about the items. Participants were then tested for their memory accuracy about the items and rated their confidence in memory. While memory appeared to be less accurate in older adults than in younger adults, misinformation and item location (central vs. peripheral) affected memory accuracy similarly for young and older adults. In contrast, confidence judgments differed by age, whereby older adults reported higher overall confidence, with their confidence was less sensitive to the central vs. peripheral distinction.

  • Recovering Meaning from "Meaningless'' Texts: Semantic Ambiguity Reduction as Recovery of Distributional Structure

    What does "blalfed'' mean in "The fuite twouch blalfed on the splelve snaitch?" Although it may seem impossible to know, people can infer word meanings from such "nonsensical'' contexts. The correct word--"sat''--was guessed correctly by 25% of our participants. We investigate people's ability to infer word meanings in such Jabberwocky texts. Participants guessed the meaning of a target word and selectively unmasked Jabberwocky words, allowing us to study the links between ambiguity, sampling choices, and accuracy of identifying the meaning of a target "nonsense'' word. Ambiguity predicted how participants sampled the text and how successfully they recover meaning, with bidirectional entropy being more predictive than entropy that only considers information preceding the target word. Accuracy increased with early information gain. These results suggest that people combine structural expectations with strategic sampling, and that semantic information is often distributed across a passage rather than localized to a single cue.

  • It doesn't matter how you ask: Many 4-year-olds don't quantify possibilities

    Many 4-year-olds give adultlike responses to questions about what "can" happen and the same responses to questions about what "has to" happen. Why? We argue that children repeatedly apply minimal representations of possibility to answer "can" and "have to" questions. This algorithm (a) only requires evaluating the single possibility each question asks about; (b) generates correct responses to "can" questions and (c) generates the same responses to "have to" questions. We show that half of 4-year-olds answer "have to" questions as adults do, correctly say what all the possibilities are and what the only possibilities are, and show sensitivity to multiple possibilities in nonverbal behavior. The other half make mistakes on "have to" questions, fail to quantify possibilities using "all" and "only", and deploy minimal representations of possibility on the nonverbal task. They make their decisions by evaluating a single possibility.

  • Transformers perform adaptive partial pooling

    Any language model must decide what to say in novel contexts based on information from similar contexts. But what about contexts that are not novel but merely infrequent? In hierarchical regression, the model's predictions for behavior in a context are affected by observations from similar contexts to the extent that 1) the current context is infrequent and 2) different contexts behave similarly. This is called adaptive partial pooling. This paper shows that next-word predictions of a transformer (GPT2) are affected by observations from outside the current context, but this pooling reduces with more training. Pooling is affected by context frequency, context number (type frequency) and context variability in a qualitatively similar way to hierarchical regression. However, there is a "sweet spot" in training at which the transformer best matches the behavior of hierarchical regression. This is the point at which the effect of context frequency on pooling is at maximum.

  • Negative strengthening and speaker adaptation with vague quantifiers

    The interpretation of negated vague predicates, whose meanings depend on underspecified thresholds, is variable. Some positive relative adjectives exhibit asymmetric negative strengthening (not large implies small), whereas negative adjectives generally do not (not small does not imply large). However, the behavior of negated vague quantifiers (e.g., many, few) remains understudied. Previous work also shows that listeners adapt to speakers' thresholds for vague predicates through non-negated assertions, but it is unknown whether such adaptation extends to negated utterances. We investigate whether listeners adapt to speaker thresholds conveyed via negated assertions and whether negated vague quantifiers are interpreted as symmetrically weaker than their antonyms. We find that listeners track speaker thresholds despite negation and interpret negated quantifiers as symmetrically weaker than their antonyms, challenging asymmetric accounts of negative strengthening.

  • Human-AI Synergy Supports Collective Creative Search

    The creation of new ideas and content increasingly relies on collectives: not just of multiple humans, but humans and AI agents together. We study collective generation of new ideas using a controlled word-guessing task that balances open-endedness with an objective measure of task performance. Participants attempt to infer a hidden target word, scored based on the semantic similarity of their guesses to the target, while also observing the best guess from previous players. We found that hybrid human–AI groups showed higher performance than both human-only and AI-only groups. Within hybrid groups, both humans and AI agents systematically adjust their strategies relative to single-agent conditions, suggesting higher-order interaction effects, whereby agents adapt to each other's presence. Although some performance benefits can be reproduced through collaboration between heterogeneous AI systems, human–AI collaboration remains superior, underscoring complementary roles in collective creativity. Together, these findings demonstrate the advantages of human–AI synergy in collective intelligence tasks.

  • Clustering during Scene Description from Memory

    Language production requires speakers to resolve the problem of linearization, where they must choose what to talk about and in what order. Previous work has shown that scene descriptions while viewing the scene show evidence of clustering. Objects which are physically close together are mentioned close together and objects which are semantically similar are mentioned close together. Additionally, the transition time from one object to the next is longer when jumping across pre-defined physical and semantic clusters compared to when staying in the same cluster. We compared previous results with a new experiment where participants produced scene descriptions from memory. We find that descriptions from memory show evidence of clustering using physical and semantic relationships. Results also show that they rely more heavily on semantic relationships than descriptions generated while viewing the image. The work highlights the importance of task demands on linearization strategies in multiutterance production.

  • What do moral rules mean?

    People often communicate social and moral expectations to one another in terms of rules. While these rules often seem simple (e.g. "No walking on the grass'') people understand that much more is being communicated than it at first appears. In this paper we present a novel theory of how people understand moral rules and instantiate that theory in a computational cognitive model. We argue that moral rules are a special kind of speech act, directed at a group and expressing a summary of the agreement that rational actors would arrive at when navigating an interdependent choice problem. Rules, on this conception, are therefore closely tied to their reasons, which reflect the interests of the parties to the agreement. This structure is very powerful: it allows people to interpret and apply rules flexibly in novel situations and edge cases. Specifically, we argue that when judging whether an action violates a rule, people reflect on the reason for the rule and harness mental models of agreement (and their heuristic approximations) to determine whether the action upholds or undermines the reason. Put simply, people ask whether the purpose of the rule would still be upheld if the rule permitted everyone to act in this way.

  • Egocentric Bias in Vision-Language Models

    Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vision-language models. The task requires simulating 180-degree rotations of 2D character strings from another agent's perspective, isolating spatial transformation from 3D scene complexity. Evaluating 103 VLMs reveals systematic egocentric bias: the vast majority perform below chance, with roughly three-quarters of errors reproducing the camera viewpoint. Control experiments expose a compositional deficit--models achieve high theory-of-mind accuracy and above-chance mental rotation in isolation, yet fail catastrophically when integration is required. This dissociation indicates that current VLMs lack the mechanisms needed to bind social awareness to spatial operations, suggesting fundamental limitations in model-based spatial reasoning. FlipSet provides a cognitively grounded testbed for diagnosing perspective-taking capabilities in multimodal systems.

  • Asymmetric Effects of Retrieval Practice on Temporal and Spatial Dimensions in Episodic Memory

    Episodic memories are structured by both temporal and spatial relations, yet how these dimensions shape retrieval remains unclear. We investigated how temporal and spatial dimensions differentially contribute to the benefits of retrieval practice using a behavioral paradigm in which participants encoded items in a structured spatiotemporal environment, followed by selective retrieval practice and a final free recall test (N = 54). Across analyses, we replicated contiguity effects in both domains, confirming that recall is influenced by when and where experiences occurred. Retrieval-induced benefits were selective: non-practiced items studied closer in time to retrieved items showed greater facilitation in final recall, following a graded, distance-dependent function. In contrast, spatial proximity predicted recall overall but did not exhibit comparable retrieval-induced spillover. Using continuous, item-level indices of retrieval influence, we demonstrate that retrieval practice selectively amplifies temporal structure rather than uniformly strengthening all organizational dimensions.

  • DST-GN: Predicting Human Mobility via Disentangled Spatio-Temporal and Category Graph Networks

    Predicting human mobility, specifically next Point-of-Interest (POI) suggestions, requires modeling how individuals integrate spatial, temporal and category cues during decision-making. However, existing computational models often conflate these heterogeneous signals or struggle with data sparsity. We propose the Disentangled Spatio-Temporal and Category Graph Network (DST-GN), a framework that reconstructs cognitive maps through representation learning. DST-GN models spatial, temporal and category contexts as independent graph views, employing Graph Attention Networks (GAT) to learn view-specific embeddings. To ensure coherence across these disentangled pathways, we introduce a cross-view contrastive learning mechanism that aligns signals into a unified semantic space. Furthermore, a global Collective Flow Enhancer is integrated to mitigate data sparsity by leveraging population-level heuristics. Experimental results on three real-world datasets show that our model outperforms state-of-the-art techniques, achieving average improvements of 22.84% in HR@10 and 18.55% in NDCG@10.

  • Preserving Inter-Brain Coupling: A Constraint-Based Preprocessing Method for Collaborative Multi-brain Motor Imagery

    Traditional collaborative Brain Computer Interface (cBCI) pipelines treat participants as independent entities, inadvertently suppressing inter-brain neural dependencies critical for hyperscanning-based tasks. We propose an Inter-brain Coupling-Constrained Independent Component Analysis (ICC-ICA) method that incorporates a coupling gradient into the objective function to preserve shared neural signatures. Validated on a motor imagery (MI) dataset, our approach demonstrates that integrating inter-brain constraints does not compromise signal quality, yielding a stable 1.76% accuracy improvement. Crucially, functional network analysis reveals enhanced recovery of inter-brain links primarily localized in task-relevant motor regions. Furthermore, ICC-ICA shows superior retention of both intra-frequency and cross-frequency coupling across the entire task duration. These findings demonstrate that explicitly regularizing the ICA update rule with inter-participant constraints effectively safeguards social brain markers, providing a robust foundation for group-level neural decoding and more valid collaborative BCI systems.

  • Beyond Compliance: Children's Perceived Agency and Agentic Interpretation of Parental Control in Family Decision-Making

    Abstract Classic socialization frameworks often link parental control with reduced child agency, yet less is known about how children cognitively process external constraints. This study examines the control-agency model by investigating: (1) how maternal beliefs shape children's perceived agency within the Chinese Guan (governance) context, and (2) how children cognitively navigate situations of maximal constraint. Using a novel vignette-based family decision-making task with 82 mother–child dyads (Mage = 6.6) alongside parental belief questionnaires, we assessed maternal beliefs about Guan and child agency, maternal decision dominance, and children's perceived agency across everyday scenarios. Boundary cases of minimal perceived agency were further examined to probe children's interpretations of maternal control and their strategic responses. Results showed maternal beliefs significantly predicted children's perceived agency independently of decision dominance. In the learning scenario, agency beliefs positively predicted perceived agency (_ = .791, p = .043), whereas control-oriented Guan beliefs negatively predicted it (_ = -.510, p = .009; R_ = .151). Additionally, in boundary cases, children's perceived agency significantly informed their interpretations and strategies, yet these two outputs frequently diverged, indicating functional decoupling. These findings indicate that children's agency is not reducible to behavioral autonomy, but also involves internal interpretive processes even under high constraints.

  • Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision

    Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context, but evaluating whether artificial scanpaths are "human-like" is difficult on object-centric datasets with strong center bias. Using Gaze-CIFAR-10, we show that a trivial center-fixation baseline achieves surprisingly strong scores under common scanpath metrics, blurring the distinction between behavioral alignment and central tendency. We introduce GCS (Gaze Consistency Score), a practical center-debiased and movement-aware score that normalizes against human and corner references, subtracts the center baseline, and adds a small movement-similarity term. Applying GCS to a hard-attention classifier under varied fovea-periphery constraints identifies a restricted mid-range regime: a moderate foveal patch with peripheral context yields stronger center-debiased alignment than either narrower or broader alternatives. This regime is not identified by accuracy alone; the highest-accuracy setting differs from the best-GCS setting. These results highlight the need for bias-aware scanpath evaluation and suggest that, on Gaze-CIFAR-10 and under this hard-attention setting, perceptual constraints shape when task-trained policies appear relatively human-like.

  • Eliciting Trustworthiness Priors of Large Language Models via Economic Games

    One critical aspect of building human-centered, trustworthy artificial intelligence (AI) systems is maintaining calibrated trust: appropriate reliance on AI systems outperforms both overtrust (e.g., automation bias) and undertrust (e.g., disuse). A fundamental challenge, however, is how to characterize the level of trust exhibited by an AI system itself. Here, we propose a novel elicitation method based on iterated in-context learning (Zhu & Griffiths, 2024a) and apply it to elicit trustworthiness priors using the Trust Game from behavioral game theory. The Trust Game is particularly well suited for this purpose because it operationalizes trust as voluntary exposure to risk based on beliefs about another agent, rather than self-reported attitudes. Using our method, we elicit trustworthiness priors from several leading large language models (LLMs) and find that GPT-4.1's trustworthiness priors closely track those observed in humans. Building on this result, we further examine how GPT-4.1 responds to different player personas in the Trust Game, providing an initial characterization of how such models differentiate trust across agent characteristics. Finally, we show that variation in elicited trustworthiness can be well predicted by a stereotype-based model grounded in perceived warmth and competence.

  • Decoding Latent Decision Strategies from Think-Aloud Protocols

    Latent decision strategies are central to human behavior and cognition. They may differ across persons and fluctuate over time within a single person. This paper introduces a framework that extracts latent, trial-level decision strategies from think-aloud verbalizations, with the resulting strategy estimates mapped onto formal strategy identification via quantitative model selection. We validated this approach with two intertemporal choice experiments, in which participants verbalized their thoughts while making binary choices between smaller-sooner and larger-later options in free-choice (Experiment 1) or instructed (Experiment 2) conditions. We used three popular LLMs to rate participants' alternative-based versus attribute-based evaluations based on their think-aloud protocols at the trial level. Results suggest that think-aloud protocols effectively capture individual variations in decision strategies, as reflected in comparisons between computational models, and can detect strategy shifts due to instructions. Our framework offers a scalable, powerful tool for understanding latent cognitive processes underlying human choice behavior.

  • Attentional Demand and Subliminal Affect Jointly Regulate Task Performance and Mind Wandering

    Understanding how attention and emotion shape conscious experience remains a central challenge in cognitive science. The present study examined how perceptual load and subliminal emotional information influence task performance and mind wandering. Participants performed a letter search task under low and high perceptual load while being subliminally primed with task-irrelevant happy or angry faces. Thought probes assessed task-unrelated thought, including deliberate and unintentional forms, with block length varied to reduce predictability. High perceptual load slowed responses and reduced accuracy, consistent with load theory. Emotional valence showed limited modulation, with happy faces associated with interference across load conditions and angry faces showing some influence under low load. Mind wandering increased with block length and showed modest modulation by load and emotion, particularly for unintentional task-unrelated thought. Trait mind wandering and ADHD symptoms were associated with higher task-unrelated thought. Overall, the findings suggest that subliminal emotional signals and attentional demands jointly shape performance and conscious thought.

  • Assembly Instructions for the Modular Mind

    How do domain-specific cognitive systems interface with knowledge from beyond their respective domains? In this paper, we draw on the notion of domain-specificity in programming language theory to provide a computational account of how domain-specific systems might interface with domain-general world knowledge, as well as with other domain-specific systems.

  • Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations

    Changing the behavior of large language models (LLMs) can be as straightforward as editing the Transformer's residual streams using appropriately constructed "steering vectors.'' These modifications to internal neural activations, a form of representation engineering, offer an effective and targeted means of influencing model behavior without retraining or fine-tuning the model. But how can such steering vectors be systematically identified? We propose a principled approach, which we call self-alignment, that uncovers steering vectors by aligning latent representations elicited through behavioral methods (specifically, Markov chain Monte Carlo with LLMs) with their neural counterparts. To evaluate this approach, we focus on extracting latent risk preferences from LLMs and steering their risk-related outputs using the aligned representations as steering vectors. We show that the resulting steering vectors successfully and reliably modulate LLM outputs in line with the targeted behavior.

  • The disjunction effect does not violate the Law of Total Probability

    The disjunction effect (DE) refers to an empirical violation of the Sure-Thing Principle, which states that if a person is willing to take an action independently of the outcome of some event, then they must be willing to do so even when the outcome of the event is unknown. In practice, authors report a population-level version of this phenomenon, specifically that fewer people are willing to take the proposed action when the outcome of the event is unknown than for any possible known outcome. This latter condition has received a lot of attention, because it presumably violates the Law of Total Probability. Here we show that this assertion is false and, consequently, this population-level version cannot be interpreted as a DE. This calls for a reevaluation of experimental results that have been interpreted as showing a DE based on the above condition.

  • How the Teaching Style and Interpretation Type of State Interventions Shape Multi-Agent Coordination

    How can an external teacher accelerate coordination among multiple learning agents? We investigate this question in a multi-agent foraging task where a simulated teacher can physically relocate agents—a direct but ambiguous pedagogical signal. We formalize a taxonomy of teaching styles (e.g., Undoing, Correcting) and agent interpretation rules. Our computational experiments reveal that the Undoing style is a uniquely powerful catalyst for spatial coordination, most effectively breaking behavioral symmetry between two agents by imposing spatial constraints that define "home regions." This catalytic effect scales, reinforcing niche specialization in three-agent systems. However, success is governed by a pedagogical matching principle: interventions fail when the agent's interpretation rule mismatches the teacher's intent (e.g., Correcting fails when interventions are interpreted as punitive). Our work provides a formal framework showing how structured guidance and agent interpretation jointly shape the emergence of collective organization.

  • Metacognitive Active Perception with Memory-Guided Hypothesis Verification

    Current multimodal visual models continue to improve in perceptual performance, yet still rely on passive, dense processing, resulting in high computational overhead. From a cognitive perspective, existing methods incorporate limited prior memory and resource allocation during perception, making it difficult to construct perception as goal-directed sequential decision-making. This paper proposes an active metacognitive framework inspired by biological memory priors, organizing visual understanding as a memory-driven hypothesis–verification process. The model forms initial beliefs from low-resolution global information and internal memory, and selects local observations to reduce uncertainty. As evidence accumulates, the system adaptively terminates perception based on confidence, enabling on-demand allocation of computational resources. Experiments across visual complexity settings show that this approach matches or surpasses traditional models while reducing inference time and token usage by approximately 15–30%, and produces structured, interpretable observation sequences.

  • MLDA-STG: A Multi-Level Domain Adaptation Spatio-Temporal Graph Network for Cross-Subject EEG Fatigue Detection

    Reliable decoding of cognitive fatigue from Electroencephalography (EEG) is essential for monitoring sustained attention but is hampered by significant inter-subject variability in neural dynamics. To address the resulting domain shift, we propose the Multi-Level Domain Adaptation Network with an Enhanced Spatio-Temporal Graph (MLDA-STG). Our approach uniquely models the evolving functional connectivity of the brain during cognitive decline through a dynamic graph fusion network, while a collaborative learning mechanism disentangles personalized neural traits from shared fatigue patterns. Crucially, we introduce a Multi-Level Alignment framework that goes beyond global adaptation to align conditional distributions, preserving the semantic structure of cognitive states across individuals. Evaluations on the SEED-VIG and SADT datasets demonstrate that MLDA-STG achieves state-of-the-art performance, offering a robust solution for calibration-free cognitive state assessment. The source code is available at: https://github.com/zjh-sys/MLDA-STG.git.

  • On the Ordinary Concept of Intelligence

    Attributions of intelligence affect political discourse (Hartman et al., 2023; Knöchelmann & Cohrs, 2024; Oktar & Lombrozo, 2026; Stanley et al., 2020), perceptions of animals' moral standing (Piazza et al., 2014; Ruby & Heine, 2012; Sytsma & Machery, 2012), and perceptions of AI (Joo, 2024; Ladak et al., 2024). Thus, it matters how people tend to attribute intelligence. The folk psychology of intelligence, however, does not figure prominently in discussions of folk psychology (Andrews, 2012; Fodor, 1987; Goldman, 2006; Nichols & Stich, 2003; Stich, 1983), and researchers have yet to extensively investigate this topic empirically. Here, we experimentally investigate whether people attribute intelligence differently across different types of activities and their practitioners, depending on whether the activities are intellectual or embodied. In three experiments, we find evidence for the Differential Use Hypothesis.

  • Strategic allocation of memory resources through the lens of dependency locality: Evidence from a controlled reading experiment

    Human working memory resources are strategically allocated in sentence processing. Recent studies have argued that less predictable linguistic units are prioritized for memory resources to be encoded with higher precision in sentence processing, as a principled solution to minimize the overall memory error given noisy encoding and resource limitations. The investigation of this claim has been operationalized from the perspective of memory robustness, predicting that representations of unexpected linguistic units are more robust against memory interference since they are more precisely encoded. In this study, using a controlled reading experiment in the A-Maze paradigm, we examine this prediction through the lens of the dependency locality effect in comprehension, a classic sentence-processing effect arising from memory decay and interference. Focusing on the subject--verb dependency in English, we find that the locality effect only occurs when the retrieval target (i.e., left codependent) is of high predictability, suggesting that high-predictability linguistic units are less prioritized for memory resources, thus encoded with less robust representation against the interference from the intervening material between codependents. Interestingly, we also observe a mild anti-locality pattern for unexpected retrieval target, pointing to a potential competition between memory- and expectation-based sentence processing mechanisms under the impact of strategic allocation of memory resources.

  • Are Recognition Decisions in Visual Working Memory Different from Recognition Decisions in Long-Term Memory?

    What are the evidence distributions underlying recognition decisions in visual working memory, and how do they compare to those in long-term memory? A recent critical test shows that long-term recognition memory is best described by a Gumbel\textsubscript{min} signal-detection model. Our study applies this test to visual working memory. Participants studied shape-colour bindings and completed a recognition task with varying set size but an equal number of studied and non-studied items. Half of the participants had to select one studied item; half had to select one non-studied item. The Gumbel\textsubscript{min} model predicts that accuracy should improve with set size when selecting a studied item but remain unchanged when selecting a non-studied item. Our results are somewhat ambiguous but suggest that accuracy increased with set size in both task variants, which would provide evidence against the Gumbel\textsubscript{min} signal-detection model for visual working memory. Future research needs to replicate our results in a design without potential experimental confounds.

  • Compression-Driven Abstraction Under Limited Inference: A Resource-Rational Account of Rules, Chunks, and Symmetries

    Humans can show abrupt gains in learning and problem solving, e.g., discovering a rule or forming a chunk. We propose a process-level account in which an agent selects among representation families by minimizing a two-part minimum description length (MDL) objective augmented with an explicit inference-cost penalty. Bounded inference is modeled as budgeted search with sparse top-k routing (limited consideration), yielding a retained-mass diagnostic _k that predicts when truncation changes choices. The framework yields simple threshold conditions for when rules, macros, and symmetry codes become worthwhile, predicting step-like changes as experience accumulates under fixed resources. Minimal simulations reproduce key signatures, including compute–coverage tradeoffs and symmetry-threshold scaling with group size. Large language model (LLM) experiments further show a semantic–exact dissociation in symmetry learning and candidate-set bottlenecks consistent with limited consideration.

  • Costly Communication Shapes Networked Social Learning: Accuracy–Utility Tradeoffs in an Agent-Based Model

    Communication in social learning is often limited by time, attention, and explicit penalties, yet many models assume free or fixed exchange. We study how two bundled communication regimes shape networked inference in an agent-based model. Groups of agents on a ring network infer a binary hidden state over 10 rounds from noisy private signals and neighbors' probabilistic messages. The regimes jointly vary per-broadcast penalty, broadcast probability, per-round broadcast cap, and social-update weight. Simulations show that the high-cost, communication-constrained regime produces fewer outgoing broadcasts, slower reduction of logged-belief dispersion, and stronger dependence of final accuracy and score on private-signal reliability. Increasing signal strength therefore yields larger gains in the high-cost regime than in the low-cost regime. These results should be interpreted as regime-level differences rather than as the isolated causal effect of message penalty alone, and they generate testable predictions for human experiments.

  • Haunting Images of Forgotten Films: Reddit Movie Threads Reveal Emotion–Memory Interactions

    Emotional events are subjectively vivid, but whether this vividness reflects superior objective memory remains debated. We distinguish three competing hypotheses: (1) Emotion does not enhance objective recall despite subjective vividness, (2) Emotion selectively enhances central sensory details while impairing peripheral contextual information, or (3) Emotion enhances all memory aspects, with benefits spreading from central to peripheral details. We tested these hypotheses using a naturalistic dataset from three Reddit discussion forums ("subreddits''), where users post remembered details of forgotten movies. Analyzing 1,509 posts from movie identification subreddits, we found that negative emotion strongly predicted which movies were remembered (supporting hypotheses 2 and 3 over hypothesis 1), with horror and thriller movies significantly overrepresented. Scene-level details were recalled more frequently and accurately than movie-level information, suggesting that emotion preferentially affects central, sensory details. However, both scene-level and movie-level details were enhanced by negative emotion without a significant interaction, supporting the third hypothesis that emotion broadly enhances memory while prioritizing central information. These findings suggest that negative emotion enhances both the selectivity and overall strength of long-term episodic memories, with implications for understanding emotional memory in naturalistic contexts.

  • A Heterogeneous Computational Model Reveals Temporal Hierarchy of Brain

    Brain functional activity exhibits a distinct temporal hierarchy, characterized by the increase in intrinsic neural timescale (INT) from unimodal to transmodal cortices. However, existing large-scale brain network models struggle to accurately simulate this timescale hierarchy. This study developed a heterogeneous dynamic mean-field (hDMF) model by quantifying cortical microstructural differences using T1-weighted/T2-weighted (T1w/T2w) mapping. For model training, we proposed a novel multi-objective expectation maximization (MOEM) algorithm guided by bifurcation theory to achieve precise optimization in high-dimensional parameter spaces. Results demonstrated that the hDMF model significantly outperformed traditional models in fitting functional connectivity and metastable states, while successfully reproducing the gradient distribution of INT. Further application to ADNI clinical data revealed its ability to effectively simulate the abnormally elevated INT in Alzheimer's patients. In summary, this study provides a high-fidelity computational model for understanding the spatiotemporal brain dynamics and demonstrates potential clinical application value in biomarker exploration for neurodegenerative diseases.

  • Joint Modeling of Choices and Response Times in Multi-stage Decisions via Likelihood Approximation

    Planning involves a process of considering future states before acting. To understand this process, researchers typically infer planning algorithms by fitting computational models to choices. However, different planning models often predict the same choices, despite relying on different computations. Reaction time can help distinguish among models, since different computations produce different temporal signatures. However, incorporating reaction time into fitting is challenging because analytical likelihoods are typically unavailable. Here we propose a likelihood-free method to estimate the density for choices and reaction times in multi-stage decision making. We validate the method through comparisons with analytical solutions, parameter recovery, and showing robust estimates relative to distribution-free and summary statistic approaches. Through a new human experiment and fitting evidence accumulation models from Solway and Botvinick (2015), we demonstrate that modeling the full distribution is important to explain human behavior. Overall, our method is a valuable tool for modeling reaction times in multi-stage decision-making.

  • Rational communication explains Differential Argument Marking and typological asymmetries between subject and object marking

    Transitive events involve an agent (source of the action) and a patient (recipient of the action), and the successful communication of such events requires conveying these roles. Languages employ various strategies for this, including word order and case marking. Differential Argument Marking (DAM), denoting systems in which case marking only occurs in certain contexts, has garnered considerable attention for its cross-linguistic prevalence. One account posits that DAM arises via efficiency: referents with properties atypical for their role (and therefore less predictable) create pressure for marking, while typical referents do not. This predicts similar numbers of languages with atypical agent-marking and atypical patient-marking. However, patient-marking is not only more common, but agent-marking systems often violate expected typicality patterns. To address this asymmetry, we develop a computational model showing that it emerges naturally from the interaction of typicality biases with production costs.

  • Inferring possible causes from event structures

    Causal reasoning is inherently temporal in nature; humans can infer whether one event is a possible cause of another, even from temporal descriptions. For example, the sentence, "after the book was placed, the shelf collapsed" implies that it's possible (but not necessary) that the book's placement caused the shelf to collapse, and that it's impossible that the shelf's collapse caused the book's placement. These modal inferences challenge existing accounts of causal reasoning. Two experiments reveal how humans infer possible causes from temporal descriptions. Experiment 1 showed that their possibility judgements depend on the event structures described by temporal relations. Experiment 2 varied the consistency of event descriptions and found that people inferred possible causes even from conflicting information. These data challenge extant theories of causal reasoning but corroborate the hypothesis that people comprehend causal relations by building mental simulations of event structures.

  • Mapping Emotion Representations: Evidence for Category–Dimension Relationships and Individual Differences

    A central question in affective computing is how emotions should be represented. Emotion datasets typically adopt either categorical labels (e.g., Ekman categories) or dimensional ratings (e.g., Valence–Arousal–Dominance). Although these frameworks are often treated as comparable descriptions of affect, most work relies on only one representation at a time, limiting cross-dataset comparability and leaving their relationship empirically under-specified. This study explores how VAD ratings map onto categorical emotion judgments, and whether this mapping is consistent across individuals. We fit generalized linear mixed models with participant-level random effects using the MSP-Podcast Corpus to predict categorical labels from dimensional ratings. The results showed a significant effect of Valence, Arousal, and Dominance across several emotion categories (Sadness, Happiness, Anger, Fear, and Disgust) with significant participant-level variability. Overall, our results highlight that categorical and dimensional emotion representations are systematically related but their mapping is shaped by individual differences in affective interpretation. This suggests that affective computing models should explicitly account for rater-dependent variation when using dimensional and categorical emotion representations.

  • Bridging Associative Memory and Logical Reasoning: A Causal-Enhanced Dual-Process Approach with Metacognitive Monitoring

    While Large Language Models (LLMs) exhibit impressive fluency, their reasoning often relies on surface-level statistical correlations rather than robust causal understanding. Current reasoning approaches, such as Retrieval-Augmented Generation (RAG), mitigate knowledge obsolescence but typically depend on vector-based retrieval, which mimics associative memory but fails to support the rigorous logical deduction required for complex queries. To address this, we propose a novel framework titled Bridging Associative Memory and Logical Reasoning, which implements a dual-process cognitive architecture. We construct high-fidelity Causal Knowledge Graphs through a human-in-the-loop method, utilizing expert-supervised LoRA fine-tuning to extract reliable causal dependencies from raw text. During inference, a Query Cognitive Planner orchestrates a Dual-Path Memory Access strategy, synergizing associative vector evidence with causal graph traversal. Crucially, a metacognitive validation acts as a "System 2" critic for iterative self-correction. Experiments demonstrate that this cooperative interaction outperforms associative baselines, confirming that explicit "System 2" monitoring effectively reduces logical hallucinations.

  • The Eye Movement Pattern and Brain Dynamics indicate successful learning during education video viewing

    Online education has become a prominent feature of modern learning. However, what determines learning success in online education remains unclear. To explore this, the present study applies Hidden Markov Model to eye-tracking and EEG data in an open dataset containing educational video-viewing task. The model revealed two eye movement patterns: a global pattern where fixations were distributed across the video, and a local pattern where fixations focused on task-relevant features. The model also revealed 6 distinct brain states involving attention and memory. Experiments 1 and 2 explore the impact of task engagement on eye movement. We found that compared to intentional learning condition, more participants use a global pattern and suffer from worse memory performance in incidental learning condition. Experiment 3 explores the impact of video style. We found that only in the presenter-and-animation style, the local eye movement patterns and learning brain states are associated with learning success. Overall, this study explores factors that indicate learning success and offers insights for improving learning efficiency in online education.

  • Spatial navigation at the extremes: The collective accomplishment of high-speed, high-risk navigation in Rally Racing

    Spatial navigation is a fundamental survival ability, and humans have evolved to navigate at the everyday scales of time and space that are experienced routinely. But some humans seem to shatter these limitations. Teams in the high-risk sport of rally racing must navigate complex courses at incredibly high speeds, where a single mistake can be fatal. How do rally teams accomplish navigation that is seemingly beyond the human scale? Here, we investigate rally racing to understand how distributed cognitive systems can accomplish spatial navigation at the extremes. By combining qualitative analysis, computer vision, and time series methods, we show how high-speed navigation is accomplished by a distributed system in which driver, co-driver, and cultural artifacts come into coordination. We end by discussing how extreme feats of cognition are made possible not by superhuman individuals, but by redistributing the computational burden across time, people, and artifacts.

  • Disfluency in Spontaneous Speech: Social Attribution and Behavioral Consequences

    Speech disfluency such as filled pauses (um and uh) is commonly associated with negative speaker evaluation, yet findings from previous studies relied on scripted, lab-recorded speech. Here, in two experiments, we investigate how filled pauses influence social judgment and decision making using unscripted, multi-sentence discourse extracted from spontaneous speech corpora. We find that filled pauses led to lower perceived readiness and certainty, regardless of speaker expertise (Experiment 1). However, disfluent experts were judged as more careful than disfluent novices, indicating a context-dependent social benefit of disfluency. In a follow-up decision task (Experiment 2), listeners were less likely to choose to ask disfluent speakers again (as opposed to someone else), an effect mediated by perceived readiness but not certainty. By integrating spontaneous speech and speaker profiles, this study demonstrates the importance of socially situated accounts of language processing and the dynamic nature of social evaluation in communication.

  • Are counterfactuals necessary for actual causation judgments?

    Two descriptive accounts of actual causation judgment, the Counterfactual Simulation Model (CSM) and the Counterfactual Effect Size Model (CESM), propose that actual-cause judgments for an observed event are derived from operations over mentally simulated counterfactual alternatives of the event. We argue that counterfactual models face computational and conceptual obstacles and propose the Retrieval of Invariant Causal Knowledge (RICK) model, which utilizes forward mental simulation to map type-level causal knowledge directly to token causal events. We demonstrate two predicted shortcomings of the CSM and compare with our model's performance.

  • Evaluating ___: Global Motion Detection Requires Random Scale Representation and Allows for Threshold Representation

    Global motion detection performance is often described via ___ (i.e., equal-variance Gaussian Signal Detection Theory), de- spite evidence that global motion detection requires a threshold representation which is incompatible with ___. Both ___ and threshold representation share the core assumption of random- scale representation (RSR), which posits a latent evidence scale without committing to specific distributional forms. We apply critical tests to assess whether global motion detection requires a RSR and whether a threshold representation is admissible. Participants completed an __-alternative forced-choice random dot motion task with __ = {2 . . . 8}. Accuracy across __ sat- isfied RSR constraints, providing strong evidence that global motion detection relies on RSR. We derived a hazard func- tion from the accuracy and tested the monotonicity constraint implied by threshold representation. The hazard function satis- fied monotonicity, consistent with a threshold account. Recon- structed ROC curves were piece-wise linear, inconsistent with equal-variance Gaussian SDT, while generally consistent with threshold models.

  • Testing Potential Mechanisms of the Description-Experience Gap in Loss Aversion

    Recent research has found a novel description-experience gap in loss aversion, such that the overweighting of losses relative to gains in risky choice is more pronounced when people make decisions based on experienced rather than on descriptive summary information. Here, we tested two possible mechanisms of this gap: a decision-by-sampling (DbS) account, according to which the relative distributions of gain and loss outcomes that the decision maker encounters differ between description and experience; and an optional stopping account, according to which the role of losses for stopping information search differs between description and experience. Neither of the two mechanisms explained the gap. Whereas in description the relative rank of losses was higher than that of gains, this pattern was not more pronounced in experience (in fact, it was even reversed)—at odds with a DbS account. Further, although individual differences in the tendency of search termination after encountering a loss (vs. a gain) showed some association with individual-level loss aversion, this mechanism did not mediate the description-experience gap in loss aversion. Our results clarify that two prominent mechanisms of preference construction are unlikely to contribute to the description-experience gap in loss aversion, clearing the view on potential contributions of alternative mechanisms (e.g., asymmetric learning of gains and losses).

  • Online library learning in human visual puzzle solving

    When learning a novel complex task, people often form efficient reusable abstractions that simplify future work, despite uncertainty about the future. We study this process in a visual puzzle task where participants define and reuse helpers---intermediate constructions that capture repeating structure. In an online experiment, participants solved puzzles of increasing difficulty. Early on, they created many helpers, favoring completeness over efficiency. With experience, helper use became more selective and efficient, reflecting sensitivity to reuse and cost. Access to helpers enabled participants to solve puzzles that were otherwise difficult or impossible. Computational modeling shows that human decision times and number of operations used to complete a puzzle increase with search space estimated by a program induction model with library learning. In contrast, raw program length predicts failure but not effort. Together, these results point to online library learning as a core mechanism in human problem solving, allowing people to flexibly build, refine, and reuse abstractions as task demands grow.

  • Toward a Mechanistic Account of Cognitive Inhibition: Integrating Computational Modeling and Psychometrics

    Cognitive inhibition is crucial for resolving conflict in decision making, but it is often estimated using reaction time differences between congruent and incongruent trials, which ignores accuracy and conflates inhibition with other processes. To address this, we adapted a Diffusion Model with Inhibition (DMI) to formally quantify cognitive inhibition—the suppression of task-irrelevant information—within a decision-making framework. This approach disentangles inhibition from processing speed and non-decision time. Combining computational modeling with structural equation modeling across multiple conflict tasks, we tested whether inhibition forms a coherent latent factor. Applying the same framework to large-scale data from over 1 million participants in the Implicit Association Task revealed distinct lifespan trajectories of inhibition across adulthood and aging. Overall, our findings show how computational modeling can uncover latent cognitive mechanisms underlying individual differences in inhibition.

  • Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?

    A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the simultaneous presence of multiple causes, while performing better in disjunctive settings. However, most demonstrations of this "conjunctive handicap'' rely on passive observation paradigms with limited evidence, where learners have no control over evidence generation. This paper asks whether this bias persists when adults are granted agency through active exploration. Using a modified "blicket detector'' task, adult participants freely intervened to identify causal objects under conjunctive or disjunctive rule structures. We show that active exploration substantially improves adults' conjunctive causal reasoning, although conjunctive rules still require more tests to infer than disjunctive rules. We further compare human performance to a range of large language models in the same setting. While some state-of-the-art models approach human-level performance on hypothesis inference accuracy, they often exhibit less efficient exploration strategies and similar conjunctive-disjunctive performance gaps.

  • Group Biases and Motivated Reasoning in Media Consumption: Evidence for Intergroup Polarization

    Group-based identification provides motivation for displaying ingroup favoritism across one's cognitive faculties. Social identity theory purports that individuals are driven to construe positive self-esteem from ingroup membership, predisposing one to positively differentiate themselves from other relevant social groups (outgroups). In this context, social narratives help sustain, propagate and rationalize the existing intergroup differences, and polarized news outlets become a common carrier of such biased narratives. In the current study, across three experimental studies using the minimal group paradigm, we investigated how individuals are predisposed to preferentially process, believe and enact on intergroup news headlines. Experiment 1 revealed faster processing of positive ingroup news headlines compared to those of positive outgroup news. Experiment 2 introduced news sources varying in reliability (low and high) and found asymmetric believability for the positive and negative intergroup news headlines, favoring ingroup irrespective of source reliability. Finally, in Experiment 3, simulating a sharing and voting task, participants again showed preferential sharing of positive ingroup headlines while withholding more negative ingroup news, and displayed clear voting preference for their minimal ingroup candidates. The study demonstrates how easily individuals are susceptible to biased intergroup news consumption, highlighting the urgency for media literacy in the increasingly polarizing information age.

  • Modeling How Inhibitory Control Affects False Belief Reasoning Over Time Using Dense Longitudinal Data

    An important puzzle in theory of mind development is the performance shift in false belief understanding between the ages of 3 to 4. Theory of Mind Mechanism (ToMM) theory (Leslie et al., 2005) argues that changes in inhibitory power (IP) account for this shift. Wang and colleagues (2019) developed a computational model to implement this theory, where they quantified the role of IP using group averaged data. Here, we extend the Wang and colleagues' (2019) model to account for changes happening across time at the individual level. We model a novel dataset where 14 children were repeatedly tested during this transitional period on average every 20 days. Their performance in high and low inhibitory demand false belief tasks (FBT), and an independent measure of their inhibitory control, was collected. We find that the model framework is successful in predicting matching inferences on children's IP given their performance in different FBTs.

  • Superstitious Ordering Despite Discouragement: The Role of Visual Features

    Humans readily perceive patterns even in situations with only random noise, but how do they respond when the genuine rule deliberately discourages this tendency? We investigated superstitious ordering under a consistent preference-discouraging (PD) schedule, in which reward probability increases for previously avoided options and decreases for previously chosen options. This contradicts typical reinforcement principles, in which some options are always preferred over others. We also examined the role of visual features, with one group using geometric shape combinations as stimuli and another group using natural images. We found that participants in both groups developed significant subjective linear orderings of stimuli, with a marginally stronger effect in the Geometrical Image condition. Comparisons with the benchmark agent, optimized Q-learning agent, and random agent showed that none of the models could explain participants' response patterns, indicating a unique and stable superstitious ordering effect in human learning.

  • Behavioral and Functional Near-Infrared Spectroscopy (fNIRS) Study Testing for Shared Magnitude Processing for Numerosity and Proportion

    A large body of research has investigated humans' representation of quantity through the Approximate Number System (ANS). In contrast, it is less clear how proportions are represented, and whether proportion processing is fully analogous to numerosity processing. In the current study, we compared numerosity and proportion discrimination within the same participants, while measuring cerebral hemodynamic responses using functional near-infrared spectroscopy (fNIRS). Behavioral results showed substantial commonalities between the two tasks, including highly comparable response accuracy, and shared ratio-dependent psychophysics. At the same time, proportion judgments were slower than numerosity judgments, suggesting additional processing demands beyond shared magnitude sensitivity. Neural evidence was less diagnostic, perhaps due to insufficient power, although a group-level conjunction analysis revealed overlapping parietal activation centered on the intraparietal sulcus for a general task contrast. Together, these findings speak to the possibility of shared magnitude processing substrate spanning number and proportion, while also cautioning the equivalence of the two forms of quantitative processing.

  • A Deep Entanglement Model of Attentional Bottleneck: Brain Plays Psi-Phi Plus Dice

    Human attention exhibits an attentional bottleneck during which conscious perception is impaired, a phenomenon that has not yet been explored directly with quantum models. This emerges when competing stimuli are presented serially and briefly. The representation of the stimuli was encoded as two-qubit entangled states, either correlated or anticorrelated. The mask intensity was encoded through random angle rotations to a qubit. The two qubits formed Bell states (Psi and Phi), undergoing deep entanglement through repeated layers. Measurement in terms of probability at each depth level simulated the attentional refraction and oscillation consistent with patterns observed in human studies, supported by extra measurements altogether, such as separate and combined probabilities, expectation values, and measures including entropy, fidelity, and concurrence. The model reproduced several effects found in attentional blink, specifically the well-known lag-1 sparing and masking effects, as well as two additional effects newly termed as partial lag-2 sparing and lag-7 divergence.

  • Attentional Inertia and Residue in Dual-Tasking Performance: Examining Within- and Between-Task Variability Using Multilevel Vector Autoregression Models

    When individuals attempt to complete multiple tasks simultaneously, constraints on how attention is deployed influence performance. However, experimental findings from dual-tasking paradigms often rely on analysis of aggregated data that obscure study of how the carryover or spillover of attention over time – attentional inertia and residue – changes within person and over time. This study introduces an analytical framework for examining if and how attentional control settings from an initial task persist both within the same task (attentional inertia) and between two different tasks (attentional residue) across trial, session, and person levels. Leveraging reaction time data from 58 participants who completed 20 online sessions of dual-task performance blocks, we employed multilevel vector autoregression models to simultaneously model the lagged and cross-lagged effects. Our results demonstrate that both positive inertia and positive residue effects manifest and that they are positively correlated, especially for those originating from the same initial task.

  • Production Efficiency in Written Hindi Falsehoods

    This study examines linguistic differences between truthful and false written narratives in Hindi using psycholinguistically motivated measures of language processing complexity. Adopting a corpus-based, processing-oriented approach, we investigate how the cognitive demands of producing falsehoods shape lexical and syntactic choices beyond surface-level cues. Participants produced written opinions that were either truthful or intentionally false, which were analyzed using metrics including surprisal, dependency distance, part-of-speech distributions, and type–token ratio. Our results show that lower mean surprisal, along with higher verb counts and reduced use of nouns and adverbs, significantly predicts false narratives. These patterns suggest that falsehood production in Hindi relies on linguistically efficient strategies characterized by predictable constructions and reduced informational specificity to ease planning and monitoring under cognitive load. By presenting evidence from an underrepresented South Asian language, this study advances cross-linguistic research on deception and highlights the value of processing-based metrics for understanding deceptive language production.

  • Cognitive Efficiency and Perceived Text Quality in Children: how Concreteness and Specificity shape Clarity and Informativeness

    Readability is often conceived in terms of processing efficiency, assuming that texts that are easier to process are also more effective communicatively. However, lexical choices may support different aspects of textual quality. This study investigates how concreteness and specificity shape children's subjective evaluations of textual quality, focusing on perceived clarity and informativeness. Children aged 9–12 read short texts manipulated for concreteness and specificity and rated them on four dimensions: ease of understanding, confusion, quantity of information, and relevance of information. Mixed-effects analyses reveal that concreteness increases perceived clarity across age groups, whereas specificity primarily tends to enhance perceived informativeness, with effects varying across development. These findings suggest that different lexical properties contribute in distinct ways to clarity and informativeness, and that their relationship is not fully captured by processing ease alone. Overall, the results support a multidimensional view of communicative efficiency and highlight developmental changes in how children evaluate informational content.

  • When Tutor Errors Match Learners' Misconceptions: A Congruency Effect in LLM Tutoring and a Scaffolding Intervention

    Errors in LLM-style tutoring are not uniformly harmful. We test a misconception-congruency effect: when an explanation's key error aligns with a learner's pre-existing misconception, the error is harder to detect and more likely to be adopted. We introduce epistemic scaffolding—claim decomposition with evidence/counterexample selection—to shift evaluation toward evidential cues. In a controlled experiment with engineering students (N = 124) using researcher-written feedback that simulates LLM explanations, congruent errors were detected less often than incongruent errors (36% vs. 58%). Scaffolding improved congruent-error detection to 62% and reduced adoption, while signal detection analyses showed higher sensitivity (d_) without increased response bias. These findings suggest learner-aware safeguards should target the interaction between tutor errors and learner priors.

  • Progressive Coherence Building in Relational Understanding: A Hierarchical Cognitive Model

    Human readers derive coherent relational understanding from information distributed across a document through a progressive, hierarchical process: local semantic associations are established first and then integrated into global coherence. How the mind achieves such robust integration while limiting combinatorial explosion and error propagation remains an open question. We propose progressive coherence building as a key computational principle, in which complex inference is constrained through sequentially layered, verifiable representations. Guided by this hypothesis, we introduce a Hierarchical Cognitive Model (HCM) for cross-context relational learning. HCM instantiates three cognitively motivated stages: (1) lexical-semantic grounding via prompt-based semantic priming; (2) local proposition formation through intra-triple attention that enforces local semantic consistency; and (3) global coherence optimization using prior-guided axial attention to integrate evidence across the narrative. This staged design stabilizes intermediate representations and mitigates error propagation. Experiments on CDR, GDA, and DocRED demonstrate state-of-the-art performance. Ablation results reveal cumulative degradation when higher stages are removed, empirically supporting the proposed hierarchical processing account.

  • Animal but not Dog: Children's Computation of Implicatures for Hierarchically Organized Categories

    Children succeed at context-specific ad hoc implicatures early, but struggle with generalized scalar implicatures until much later. This developmental trajectory is often explained by children's failure to access relevant alternatives for scalar implicatures. We tested this hypothesis by investigating scalar implicatures involving concrete nouns (superordinate – basic-level – subordinate nouns), termed hierarchical implicatures (HI). To create scenarios where a speaker can felicitously forgo a more informative term, we also manipulated speaker knowledge of different nouns. We found that 4- to 5-year-old children computed HI with superordinates (contrasted with basic-level nouns) only in certain contexts, and did not compute HI with basic-level nouns (contrasted with subordinates). Meanwhile, adults computed HI in both cases. Our findings suggested that difficulties accessing relevant alternatives might not be the main reason for children's late success in computing generalized scalar implicatures. Other factors, such as the ability to represent asymmetric relations between lexical terms, might play a bigger role.

  • Metaphors We Campaign by: Political Metaphor Analysis in U.S. Presidential Campaigns (1960--2024)

    We present an exploratory computational analysis of cognitive patterns in U.S. presidential campaign manifestos through metaphorical concept mappings. We examine documents from 1960 to 2024, using MetaPro to extract metaphors from both the Republican and Democratic parties' rhetoric. Our goal is to propose a novel way of political analysis by generating insights into the cognitive foundations of election rhetoric and revealing distinct patterns. Our research focused on three key ideas: (1) the evolution of the conceptual landscape of campaign rhetoric over six decades, (2) the different metaphorical strategies of incumbent and challenger campaigns, and (3) the divergence of Democratic and Republican parties in their metaphorical repertoires. The analysis reveals a moderate decline in cross-cycle mapping turnover over the observed timeline, as well as structural differences between incumbents and challengers, and the parties themselves, suggesting that political polarization is observable even at the level of metaphors.

  • Emotional State or Deviation? Emotion Deviations Consistently Predict Working Memory Performance in Daily Life

    Previous studies have reported contradictory findings, showing that emotion can both enhance and interfere with working memory (WM). We propose that these inconsistencies partly arise from differences in how emotional effects are conceptually defined, rather than from emotional properties alone. Using a smartphone-based experience sampling method (ESM), we repeatedly assessed multidimensional emotional states and performance on two WM tasks over approximately three months in daily life. We conducted two analyses on the same dataset: one defining emotional effects based on emotional states at the measurement occasion, and another defining emotional effects as within-person deviations from each individual's typical emotional level using linear mixed-effects models. Analyses based on emotional states yielded mixed enhancement and interference effects depending on tasks and emotional dimensions. In contrast, emotional deviations from typical level were more consistently associated with interference effects on WM performance. Together, these findings demonstrate that how emotional effects are defined critically shapes their interpretation and consistency in emotion–WM research.

  • Tri-Domain Cross-Attention Transformers for Cross-Subject EEG Decoding of Cognitive States

    Mental fatigue and high cognitive loads reduce cognitive efficiency and task performance. Electroencephalogram (EEG) can track these states with high temporal resolution, but informative patterns span time, frequency, and distributed scalp topology. Most decoders handle these domains separately or fuse them only at the output, which can miss cross-domain dependencies. We propose TD-CAT, a tri-domain cross-attention Transformer that models interactions among temporal, spectral, and spatial representations. TD-CAT builds domain-specific features via temporal and spectral convolutions and a topology-aware graph module, then integrates the three streams with bidirectional cross-attention and low-rank fusion. Under leave-one-subject-out evaluation on SADT and SEED-VIG (driver fatigue) and MAT (cognitive load), TD-CAT reaches 80.7% and 92.6% accuracy on SADT and SEED-VIG, and 82.4% on MAT, demonstrating strong cross-subject performance across multiple datasets and task settings.

  • Decoding the Neural Representation underlying Underwater Acoustic Signal Feature Perceptions in Musicians

    Long-term musical training enhances cognitive efficiency in processing music- and speech-related sounds, particularly in timbre, frequency, and temporal structure. However, the perception of underwater acoustic signals characterized by complex spectra and strong noise in musicians remains largely unexplored. We compared electroencephalography (EEG) responses of musicians and non-musicians during underwater acoustic detection tasks. Although no significant group differences were observed in event-related potentials, musicians exhibited stronger theta–alpha synchronization and more focal information flow under high task demands. Network analyses further revealed that musicians' timbre perception was associated with strengthened connectivity between the precuneus and visual cortex, whereas non-musicians relied on more distributed default-mode network pathways. Moreover, musical expertise could be decoded from combined EEG spatiotemporal and spatiospectral features with an accuracy of 82.77%. These findings suggest that musical training shapes network-level information routing during underwater acoustic perception, rather than altering early sensory encoding. Keywords:underwater acoustic target signal; auditory specificity; EEG; musical training; timbre; frequency

  • Argument Evaluation Strategies in Human and Machine Reasoning: LLMs Resemble Sound Deliberate (But Not Intuitive) Thinkers

    Humans routinely evaluate arguments to decide which reasons warrant belief or action. Large language models (LLMs) are increasingly used as evaluators, yet little is known about how their argument evaluation compares to human evaluation. In this study we examined how humans and LLMs evaluate arguments supporting answers to classic reasoning problems. Human participants rated arguments either under time pressure and cognitive load or without constraints, allowing us to situate LLM evaluations relative to intuitive and deliberate human judgments. Arguments varied in response correctness and justification type. Modeling argument ratings as a function of argument validity, surface explicitness, and alignment with the evaluator's own response revealed that intuitive human evaluations were dominated by belief-consistency bias, whereas deliberate evaluations –especially among correct responders– prioritized justification validity. In contrast, LLMs generally showed stable, validity-centered evaluative strategies, closely resembling deliberate human evaluations from correct responders and diverging from intuitive human judgments.

  • The Effects of Hallucination Warnings and Source Credibility on Content Learning and Source Memory when Learning with Large Language Models

    More and more people use large language models (LLMs) as sources of information. However, LLMs are prone to hallucinations, meaning they can produce plausible but false information. In a pre-registered laboratory experiment, we analyzed the effects of hallucination warnings (between-subjects: with vs. without) and source credibility (within-subjects: high-credible vs. unknown vs. low-credible) on content learning and source memory. N = 97 learners first received pieces of information accompanied by a source label from a simulated LLM-based chatbot, before completing learning and source memory tests. Source memory was analyzed with multinomial processing tree models. Content learning was not affected by warnings and source labels. With hallucination warnings, high- and low-credible sources were remembered better than unknown sources. However, without a warning, only low-credible sources were remembered better than unknown sources. Disclosing the source (credibility) seems promising for source memory, but especially high-credible sources may require additional highlighting in contexts without warnings.

  • When agreement looks like copying: Testing a Bayesian model of source dependency inference

    When multiple sources agree, their testimony provides stronger evidence if they are independent. But how do people infer whether sources are coordinated? We developed a Bayesian model predicting that dependency inferences should depend on three factors: verbal similarity between reports, the knowability of the topic, and the diversity of possible expressions. In a pre-registered experiment (N = 156), participants viewed social media posts varying in these dimensions and judged the likelihood of coordination. Results revealed a striking dissociation: participants were highly sensitive to verbal similarity (__ = .283), inferring substantially more dependency when posts were near-identical versus substantively similar but differently worded. However, participants were insensitive to knowability and expressive similarity, contrary to normative predictions. These findings suggest people rely on a simple "copy detection" heuristic based on surface similarity rather than engaging in full Bayesian integration of domain-relevant probabilistic information.

  • Education Shapes Executive Control: More Evidence from Turkish Literate and Illiterate Adults

    Executive functions are known to be shaped by formal education, yet direct evidence from adult populations with minimal schooling remains scarce. The present study examines the effects of education and reading experience on inhibitory control by comparing literate and illiterate adults using a simple inhibition task. Participants also completed measures of reading fluency with real and nonce words. Results revealed significantly higher inhibitory accuracy in literate participants compared to illiterate participants. Importantly, inhibitory performance showed limited variability among literate adults, suggesting a ceiling effect, whereas illiterate adults displayed substantially large individual differences. These findings support the view that formal education and sustained engagement with written language contribute to the development of inhibitory control mechanisms. The study provides novel behavioral evidence that literacy-related experiences continue to shape executive functions in adulthood from a Turkish population, highlighting the cognitive consequences of unequal access to formal education.

  • Evidence for Joint & Complementary Hemispheric Lateralization, Within Subjects, Across Tasks & Dependent Measures: Evidence from Visual Half-Field Paradigm

    Previous studies link the right hemisphere (RH) to visuo-spatial/configural processing and the left hemisphere (LH) to language, although evidence shows variability across tasks, measures, and individuals, with RH also contributing to verbal processing (Whitehouse & Bishop, 2012). The current study tested for hemispheric specialization using a visual half-field paradigm (face-matching, symmetry-judgment, and word-categorization tasks), analyzing accuracy, sensitivity (d_), RT, and decision bias (c). Mixed-effects modeling revealed a clear LVF/RH advantage for face-matching and symmetry-judgment, with higher sensitivity, faster RTs, greater accuracy, and a liberal bias for faces but a slightly conservative bias for symmetry. Word-categorization showed a subtle RVF/LH advantage in RT, both at task level and for abstract words, but no reliable visual field differences in accuracy. These findings demonstrate joint and complementary hemispheric asymmetries across tasks and measures for RH and LH functions, respectively, challenging the statistical independence view (Bryden, 1982; Harms & Elias, 2014).

  • A Rational Model of Growth Mindset Theory

    Mindset theory proposes that believing intelligence is improvable fosters academic achievement. Despite its influence, the theory lacks a mechanistic foundation. Here, we examine mindset-driven behaviors from the perspective of the computational problem the mind is solving—optimizing cumulative reward under uncertainty about skill malleability. We formalize this problem as a Markov decision process, where agents balance cultivating for future gains against harvesting immediate rewards. Growth- and fixed-mindsets are represented as the agent's priors over skill malleability. Through simulation, we demonstrate that mindset-driven behaviors, such as persistence and challenge seeking, arise as rational decisions under different prior beliefs, with optimistic priors promoting sustained engagement, belief updating, and higher rewards in favorable environments. Crucially, mindset effects vanish when environmental structures disincentivize long-term investment, leading agents to converge on a harvest-only strategy. Our model offers a unified account of the heterogeneous findings of mindset interventions, highlighting the importance of supportive environments for mindset effects to manifest.

  • An evolutionary model of recombination in social learning

    Social learning is often valued for reducing the costs of individual exploration and avoiding mistakes. Its benefits, however, extend beyond simple error prevention in domains where knowledge is compositional: when ideas are generated by combining existing elements, meetings of minds are a fertile ground for novel innovations. We introduce an evolutionary agent-based model in which agents pursue diverse learning strategies. Some agents create knowledge independently, others exchange and recombine ideas. Agents interact repeatedly in a compositional task environment, building, sharing, and combining knowledge components. We find that social interaction produces a rich pool of partial ideas, and recombination among these ideas can connect disconnected knowledge strings, though only alongside accurate discoveries made by individual learners. Social recombination allows populations to explore combinatorial solution spaces more effectively than individual learning alone, or success-biased social learning. By highlighting the interplay between idea exchange, recombination, and selective copying, our results reveal a novel pathway through which social learning enhances adaptive knowledge accumulation.

  • Underspecification and communicative efficiency of visual affixes of motion in comics

    In a corpus of 331 comics, we analyzed the size and frequency of two visual affixes in comics: motion lines (trailing behind movers) with the meaning of directed motion and circumfixing lines (imitating the contours of a mover) with an underspecified meaning of motion. Given that the meaning of circumfixing lines is underspecified compared to motion lines, we asked whether their sizes and frequencies behave in communicatively efficient ways consistent with ambiguous words in language. We hypothesized that circumfixing lines should be smaller and more frequent than motion lines, in line with the ambiguous words being shorter and more frequent than less ambiguous words in spoken language. We found that circumfixing lines are smaller but not more frequent than motion lines. We argue that circumfixing lines are smaller than motion lines because of their underspecification and lower information content and we discuss possible reasons for their lower frequency.

  • Probability Weighting from Intrapersonal Aggregation: Internal Sampling, Metacognition, and Divergence Minimization

    Probability weighting is a central component of cumulative prospect theory, but its cognitive basis remains contested. We propose a process-level account in which weighting emerges from \emph{intrapersonal aggregation}: decision makers integrate an externally provided probability with internally generated estimates from memory or simulation. We formalize this as minimization of a weighted Kullback--Leibler divergence and show that the solution is logarithmic pooling in log-odds space, yielding the standard two-parameter weighting function. Curvature reflects the relative influence of external information, while elevation captures directional bias in internal sampling. The model predicts that sampling depth and external reliability primarily affect curvature, whereas framing or affective biases primarily shift elevation. An analytic Beta-sampling benchmark makes these dependencies explicit.

  • Why do wolf howls sound like "uw"? Behavioral observations and exploratory modeling

    We report observations of Hudson Bay wolves (Canis lupus hudsonicus) producing howl-like calls accompanied by visible dynamic tongue movements. We show that the wolves actively changed the shape of their vocal tract through apical tongue gestures (i.e., the tongue tip arches anteriorly toward the palate), a phenomenon that changes vowel quality in human speech. A combination of video content analysis and articulatory modeling indicates that the apical lingual gesture changes vocal-tract resonances along the close--open dimension, producing a consistent perceptual shift from a schwa-like quality toward a close central rounded quality. These findings provide evidence that a non-primate animal can perform dynamic lingual filtering during vocal production.

  • The Supply Side: Are Biased Information Sources Necessary for Belief Polarisation?

    Do biased minds or sources cause issue polarisation? Demand-side accounts emphasise motivated reasoning where people selectively attend to, and favourably evaluate, attitude-consistent information. Supply-side accounts emphasise the information environment, such as ideologically committed sources. We present an agent-based model to test if motivated reasoning mechanisms can create polarisation without biased information sources. Our simulations suggest that populations converge to consensus, regardless of motivated reasoning strength. Polarisation only emerges given ideological broadcasters. Critically, pre-polarised populations converge without persistent biased sources, indicating that motivated reasoning does not maintain existing polarisation. We also find that selective exposure counterintuitively reduces polarisation from consensus by symmetrically filtering all extreme sources. These results support a supply-side thesis: within the logic of the model, motivated reasoning mechanisms were not sufficient for belief polarisation without biased information sources. Here, motivated reasoning and echo chambers are amplifiers, not generators, of issue polarisation.

  • An attention-to-thoughts model of stream of thought: Affective dynamics depends on attentional scope

    The stream of thought has received less empirical attention than external cognition, with existing approaches relying on thought probes or continuous verbal reports rather than modeling internal representational dynamics. The Attention-to-Thoughts (Amir & Bernstein, 2022) model offers a dynamical-systems framework in which thought trajectories emerge from moment-to-moment interactions among lower-level components, capturing several thought patterns, including those consistent with repetitive negative thinking (RNT). However, the model uses simplified thought representations and does not capture the breadth of internal attentional selection: a dimension extensively studied in external attention and implicated in internal attention, too, in studies of creativity (broad) and rumination (narrow), for instance. We extend A2T by grounding thoughts in semantic embeddings and affective ratings, and introducing an attentional scope parameter modulating internal selection breadth. Simulations reveal systematic affective and dynamic differences between thought trajectories in broad and narrow attention runs. Control analyses confirm these effects arise from model dynamics rather than embedding structure alone. This gives us an empirically testable computational basis for understanding how the scope of internal attention shapes affective experience during spontaneous thought streams.

  • Linking Neural Dynamics of Prediction to Decision: Evidence for N400-Drift-Diffusion Parameters Relationship

    Prediction is central to language comprehension, yet how neural markers of prediction translate into decision behavior remains unclear. We combined EEG with drift-diffusion modeling (DDM) to dissociate prediction-specific mechanisms from semantic priming facilitation effects. Participants (N=30) completed a lexical decision task following either one or three semantically related primes, manipulating prediction strength while holding semantic relatedness constant. The three-prime context elicited early prediction-sensitive activity (eN400, 200–300 ms) alongside classic N400 effects. Critically, DDM revealed a double dissociation: classic N400 predicted drift rate, while eN400 predicted boundary separation, reflecting different computational pathways. These findings demonstrate that even minimal predictive context (three primes) engages early neural processing. Moreover, the observed neurocomputational dissociation between neural indices and decision parameters supports hierarchical predictive accounts of language comprehension by specifying not only whether predictive context matters, but when it influences semantic processing and which task-relevant decision mechanisms it modulates.

  • The Contagious Sense of Understanding Effect: Really a Contagious Sense of Incomprehension

    The community of knowledge hypothesis holds that people estimate their level of understanding by relying on the perceived knowledge of others. Prior work describes a "contagious sense of understanding," in which self-assessed understanding increases when experts are believed to understand an issue. However, because prior studies examined only two conditions, experts understanding or not understanding, the effect's direction remains unclear as it lacks a baseline without information about others' knowledge levels. Across three experiments (N = 1080), a neutral condition was added while replicating the paradigm. Across all experiments, self-rated understanding was lower in the expert-not-understanding condition compared to both the neutral and expert-understanding conditions, with no difference between the latter two, for both public policy proposals and novel scientific phenomena. Thus, the effect reflects decreases in perceived understanding when experts lack understanding, not increases when expert understanding is explicit, suggesting a contagious sense of incomprehension rather than a contagious sense of understanding.

  • Is Ignorance About "Being Ignorant"? A Study Across Varieties of Ignorance

    In the epistemology of ignorance, it is often assumed that ignorance is best investigated by examining how the expression "to be ignorant" is used. This contrasts with many non-English languages, which lack a direct equivalent and instead express ignorance as the absence of knowledge. But even in English, is "to be ignorant" a privileged way of expressing ignorance? In two preregistered experiments, we asked English speakers to assess the naturalness of different ignorance-related expressions in brief vignettes sampling four varieties of ignorance introduced by Peels (2023). In Experiment 1, expressions involving simple negation of knowledge (i.e., "not to know") were judged more natural than "to be ignorant", both overall and across all varieties. Experiment 2 tested whether "to be ignorant" is especially appropriate when the agent is expected to know the relevant fact. This proposal was not supported. These findings suggest that, even in English, ignorance is primarily represented as the absence of epistemic good (e.g., knowledge) rather than as a distinct epistemic predicate.

  • Concepts in terms of other Concepts: a method for the analysis of collective word use

    Sometimes how concepts are used in day-to-day talk differs from the expectations set in explicit, socially accepted, definitions. For example, many occupations which are inherently gender-neutral exhibit a gender bias in their usage across different corpora. We propose a method to quantify interactions between concepts in collections of data using autoregressive language models. Instead of prompting or extracting internal representations (i.e. embeddings), we measure differences in token probability distributions on minimal pairs of subject-verb-complements sentences, where the main noun in the subject varies across different 'concept-words'. We demonstrate the expressivity of this approach with a case study where we recover well-known social biases in language from contemporary internet data: ranking concepts in relation to men-women reveals gender biases in what would be expected to be gender neutral vocabulary. We also show how our method can be used to examine concepts across different domains –time periods, authors or ideologies– by selecting sentences from different corpora (e.g. news articles from the 19th century). We extend this to arbitrary conceptual 'dimensions', enabling the study of concepts in terms of other concepts with state-of-the-art language models.

  • Musical Keys Diverge in Perception but Not in Performance

    Musical key is associated with distinctive sensory qualities, yet it remains unclear whether such impressions reflect pitch height alone or notation-related key representations beyond the major--minor distinction. This study tested whether sharp and flat keys evoke different sensory impressions and whether these differences are reflected in performance. In Experiment~1, musicians rated brightness, warmth, and shape while listening to Mozart excerpts transposed across all keys and into enharmonic key pairs. Sharp-labeled keys were rated as brighter, colder, and more pointed than flat-labeled keys, even when pitch content was identical. In Experiment~2, violinists performed excerpts in the same key conditions, and acoustic features were analyzed. No reliable sharp--flat differences were observed in the acoustic measures examined here. These findings suggest that sharp--flat distinctions can shape explicit perceptual judgments beyond pitch height without necessarily producing corresponding differences in measured performance acoustics.

  • Set Size and Repetition Effects in Cued-Recall from Long-Term Memory

    The set size effect refers to the decline in memory performance as the number of to-be-remembered items increases. However, memory retrieval in everyday life is often fast and accurate, suggesting the existence of factors that support efficient retrieval. One such factor may be repetition of items in memory. Theoretical views differ in how repetitions affect memory representations. Single-stronger image views suggest that each repetition strengthens an existing memory trace without increasing effective set size. Multiple-image views suggest that each repetition creates separate memory traces, inflating effective set size. In this study, we used a similarity-based sequential sampling model to generate hypotheses distinguishing these two views and tested them in a cued-recall paradigm. Results showed a robust set size effect on recall accuracy and a main effect of repetition on response times and accuracy, but no interaction. These results suggest that repetitions strengthen existing memory traces, without alleviating set size effects.

  • Complement Size and Epistemic Effects in Perception Reports

    Interpretation of perception reports with verbs like see varies depending on what the verb combines with. When see combines with a full tensed clause, the perception report is thought to be epistemically non-neutral: the perceiver must form an event-consistent belief. But it is `perceptually' non-committal --- the perceiver need not have directly witnessed the event. When the verb combines with something smaller than a full tensed clause, the pattern is claimed to flip. The perceiver must have directly witnessed the event described by the complement. But, it does not commit them to a belief about what is perceived. We experimentally test speakers' interpretations of these two types of perception reports by holding constant direct perception and manipulating epistemic neutrality. We found that irrespective of complement size, participants did not treat perception reports as epistemically-neutral. We discuss two possible avenues to account for these unexpectedly strong readings, one semantic and one pragmatic.

  • Action understanding with presentational goals

    Humans balance many goals while navigating social interactions. We are motivated not only by what we want for ourselves and for others, but also by what we want others to think about us. Inferences about agents' personal and social goals have been formally captured using Bayesian inverse planning models. However, less is known about the mechanisms underlying our ability to infer presentational goals: an agent's desires over how another agent sees them. We introduce a novel paradigm that allows us to test joint inferences about social and presentational goals. We propose an extension to inverse planning where agents derive additional utility from shifting others' beliefs about their social goals. Because the model places no constraint on the valence of presentational targets, it naturally captures cases where agents wish to appear prosocial or adversarial. Across a variety of scenarios, participants make systematic joint inferences about social goals, presentational goals, and presentational targets. Our computational model captures participants' inferences better than feature-based alternatives.

  • Let's be friends! People work together even when there is no incentive to do so

    One distinguishing feature of human social intelligence is our capacity to flexibly coordinate complex behaviors, such as when working together to prepare a meal. What cognitive or motivational processes make such joint action possible? Here, we investigate the role that a \emph{preference to enmesh behaviors} plays in human coordination using a novel two-player virtual baking paradigm. Participants could either independently bake bread by bringing their own water and flour, or jointly bake bread by bringing only water or only flour to share. Depending on the ingredient options available, trials in our experiment could correspond to either the classic Prisoner's Dilemma, in which acting jointly is disincentivized by the payoffs, or a "Friend's Dilemma", in which acting jointly is neither incentivized nor disincentivized (though it remains risky since one's partner might not share). Formal analysis of equilibria and dyadic learning dynamics indicate that people work together even when there is no incentive to do so. These results provide preliminary evidence that a motivation to enmesh behaviors with social partners supports successful coordination.

  • Bongards at the Boundary of Perception and Reasoning: Programs or Language?

    Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans possess the puzzling ability to deploy their visual rea- soning abilities in radically new situations – a skill rigorously tested by the classic set of visual reasoning challenges known as the Bongard problems. We present a neurosymbolic approach to solving these problems: given a hypothesized solution rule for a Bongard problem, we leverage LLMs to generate parameterized programmatic representations for the rule and perform parame- ter fitting using Bayesian optimization. We evaluate our method on classifying Bongard problem images given the ground truth rule, as well as on solving the problems from scratch.

  • Belief as Self-Endorsement: A Bayesian Model of Commitment Under Uncertainty

    Humans routinely treat propositions as settled for the purposes of reasoning and action despite remaining uncertainty. Standard probabilistic models capture graded belief, but they leave the formation and epistemic role of such commitments underspecified. We propose a minimal Bayesian network model in which commitment is represented as an endogenous endorsement event regulated by metacognitive confidence. Within this framework, endorsement functions as a confidence-gated internal signal: conditioning on endorsement can increase subjective probability once, without new external evidence, while its dependence structure prevents epistemic bootstrapping. Extending the model to include explicit evidence yields a distinctive empirical prediction: the effect of evidence on commitment is systematically moderated by confidence. This prediction supports experimental designs in which evidence strength and confidence are independently manipulated and both commitment and probability judgments are measured. The model thus provides a tractable and empirically testable account of commitment under uncertainty.

  • Explaining "Why" Matters: Mechanistic Explanations of Misconceptions Promote Learning of Counterintuitive Conceptual Content

    What type of information best supports learning? In educational settings, instructional information may describe patterns of error and provide corrections, or explain the mechanisms that generate those errors. This distinction may be especially consequential in scientific domains where misconceptions are common. We tested whether mechanistic explanations of why misconceptions arise support adults' learning of a counterintuitive scientific idea beyond descriptive error correction. U.S. elementary school teachers (N = 127) were assigned to a Control, Description, or Mechanistic explanation condition in a professional development intervention on natural selection, where they heard different accounts of common student misconceptions. Teachers' generalized learning of natural selection was assessed using near and far transfer measures. Near generalization improved across all conditions, but far generalization improved only in the Mechanistic condition, suggesting that fostering a deeper conceptual framework for understanding misconceptions supports broader, more flexible learning of counterintuitive conceptual content than descriptive correction.

  • Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses

    Images vary in how memorable they are to humans. Inspired by findings from cognitive science and computer vision, we explore correlates of image memorability in pretrained transformer-based vision encoders for the first time. Focusing initially on activations, attention distributions, and the uniformity of image patches, we find that these features correlate with memorability to some extent. Additionally, we explore sparse autoencoder loss over the representations of vision encoders as a proxy for memorability, which yields results outperforming past methods using convolutional neural network representations. Our results shed light on the relationship between model-internal features and memorability. They show that some features are informative predictors of what makes images memorable to humans; revealing that, in particular, the reconstruction loss from our autoencoders is a strong correlate of image memorability.

  • Memory Engrams and Vehicle Indeterminacy

    Memory engrams are hypothetical neural representations that explain the persistence of memory. Finding memory engrams requires the identification and individuation of engram vehicles, which are determinate neural entities or processes bearing memory content. However, not only are there many candidate vehicles for memory engrams, but each candidate is ill-suited for explaining memory persistence due to instability via degradation, turnover, and dynamics. I argue that memory engram models presume the truth of a mere heuristic: that changes in neural activities correspond to changes in cognitive content. As such, explanations of persisting memories are thought to require persisting vehicles. This presumption in turn drives issues with determining exactly what remains stable in the brain to explain how memory persists. But to explain the persistence of memory, all that needs to persist is the capacity of the system to recall a memory. I conceive of engrams according to their functional role as causal motifs for memory recall.

  • Collapsing Waste Fractions Without Reducing Sorting Accuracy

    Waste sorting signs present an inviting avenue for comparison of categorisation theories as they often contain both a rule and either an abstract prototype or visual exemplars. Focus in waste sorting studies is often on increasing the number of waste fractions, but some public settings might require fewer-than-ordinary fractions. This paper investigates two approaches to reducing a waste sorting system's granularity at a music festival: explicit combining waste fractions by merging waste signs and implicitly combining fractions by removing alternatives to 'residual waste'. The digital sorting results are modelled as a probit regression allowing direct interpretation in terms of signal detection theory's senstitivy and criterion parameters. The paper further compares four alternative waste sorting signs including visual exemplars to improve waste sorting accuracy for the implicitly combined waste fraction. The results suggest that explicitly combining waste fractions outperform implicit combinations – even with visual exemplars added, despite resulting in more false positives.

  • Folk Theories of Well-Being: Variation Across Contexts?

    The nature of well-being has long been a central question in philosophical ethics. In a series of studies we tested lay support for three theories of well-being: hedonism, desire theories, and objective list theories. Participants were presented with vignettes where they were asked to judge which of three options would be best for a given individual, each corresponding to one of the three theories. Our findings suggest considerable variation in the criteria for well-being people employ in different contexts, especially for people of different ages. Hedonism and objective list theories garnered the most support, for older and younger individuals respectively. Desire theories received very little support. These results cast doubt on the acceptability of desire theories of well-being to much of the public and offer suggestive evidence for what might be called contextualist conceptions of well-being among the folk, perhaps lending some support to philosophical views of this sort.

  • AI doesn't produce speech like caregivers do

    AI-generated child-directed speech is increasingly positioned as a supplement or replacement for caregiver speech, either for use in early-learning technologies (AI-powered toys) or as a substitute to human data for research. We compare child-directed speech from AI and human caregivers along dimensions linked to child language outcomes: utterance length, lexical diversity/complexity, pronoun use, wh-words, and verb types. While the AI-generated speech is longer than caregiver input, it is less variable and exposes children to fewer word types. AI and caregiver speech also show different distributions of grammatical elements: AI favors first- and second-person pronouns, while caregivers use a broader set of referential forms; AI produces primarily argument wh-words, whereas caregivers use wh-words that appear in more complex syntactic constructions; and AI tends to use different kinds of verbs. These findings suggest that AI may not sufficiently imitate child-directed speech with the properties required for early language learning.

  • Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues

    Humans typically use natural language to update teammates on task states. Since not all updates are communicated, discrepancies arise between the team members' mental models that negatively affect overall team performance. How can we categorize such discrepancies? Do misalignments detected in team dialogue predict future mental model misalignments? Traditional shared mental model (SMM) assessment methods rely on retrospective expert coding that cannot capture real-time coordination dynamics. We propose a framework to identify and categorize four types of mental model discrepancies: {\em unsupported beliefs}, {\em false beliefs}, {\em belief contradictions}, and {\em omissions}, all of which can naturally emerge in team dialogues. Using dialogues from twenty dyad teams performing collaborative object identification tasks across four sequential levels, we demonstrate that these discrepancy patterns contain predictive signals. Averaging historical discrepancy counts achieves meaningful prediction accuracy using uniform weighting as an exploratory baseline, with differential predictability across discrepancy types.

  • Overconfidence without Understanding? Evidence for an Illusion of Understanding in AI-Assisted Explanations

    The Illusion of Explanatory Depth (IOED), i.e. the tendency to overestimate one's understanding, has been shown to increase when people search for explanations online. However, it remains unclear whether this effect extends to explanations obtained from AI chatbots. We examined whether chatbot use increases IOED and how it affects explanatory quality. In a between-subjects experiment, one group consulted ChatGPT on four topics, a second group read texts identical to the chatbot's responses without knowing their source, and a control group received no materials. Participants then rated their ability to explain each topic, produced written explanations, and evaluated their explanations. Results showed that the AI group exhibited the greatest overestimation of explanatory ability. In addition, both coder ratings and automated text analyses indicated lower explanation quality in the AI group compared to the non-AI group.

  • Neural Fields as World Models

    Humans rehearse possible futures offline, as in mental practice and perhaps dreaming, suggesting that world models may support task learning away from the environment. Standard machine learning world models compress visual input into latent vectors, discarding the spatial structure that characterizes sensory cortex. We propose isomorphic world models: architectures that preserve sensory topology, so physics prediction becomes geometric propagation rather than abstract state transition. We implement this idea with motor-gated neural fields, where activity evolves through local lateral connectivity and motor commands multiplicatively modulate specific channels. Across three experiments, the same architecture learns ballistic prediction without "teleporting," improves a catching policy offline by propagating task error through a frozen learned world model, and develops body-selective motor channels without body labels. These results provide preliminary evidence that physical prediction, offline task learning, and body-linked representation share a common computational substrate: action-conditional prediction within a spatial map.

  • Gaze Cuing Effect in Intentional Action: Yes, Action Matters!

    One of the core components of non-verbal communication and joint attention is gaze cuing – our tendency to shift attention in the direction of another's gaze. While extensively studied, little is known about whether and how internal states like the sense of agency influence sensitivity to gaze cues. In this study, we investigate whether control over action (action-outcome consistency) modulates gaze cuing effect in a color discrimination task. Participants viewed a synthetic face with lateral (left/right) pupil motion, followed by the appearance of a colored target to the left/right of the face. During each trial, participants either passively observed the gaze movement (control condition), or actively initiated it by selecting a gaze direction and pressing a corresponding button. Crucially, in active trials, the gaze direction could be consistent or inconsistent with the chosen direction. The results showed a larger in magnitude gaze cuing effect in action-outcome consistent condition; moreover, time distribution analyses suggested that the effect appeared early on, in contrast to control and action-outcome inconsistent conditions. Together, these data point towards a possible link between action control and social attentional orienting.

  • Longitudinal Associations Between Early Language Input and Caregiver–Infant Neural Synchrony

    Language acquisition is shaped by infants' engagement in socially embedded, synchronized communicative interactions. At a neural level, this is instantiated as neural synchrony. Extant literature has focused on in-the-moment predictors of neural synchrony, yet little is known about how early language environments shape neural coordination across development. The present study investigated the joint and independent effects of language quantity (i.e., number of adult words) and interaction quality (i.e., caregiver-infant reciprocal exchanges) on neural synchrony. We used day-long audio recordings collected between 12-24 months to characterize speech in infants' everyday environment and examined effects on caregiver-infant neural synchrony during play at 24 months. We found that greater conversational turn-taking at 12 months significantly predicted neural synchrony at 24 months, suggesting that early reciprocal interaction may be central in shaping neural attunement over time. These findings suggest that early reciprocal language experiences may scaffold the development of caregiver–infant neural attunement.

  • Photovoice as a Window into Social Cognition: Investigating Holistic Expressions of Minoritized Identity

    The present study used a community-oriented photovoice approach to explore the lived experiences of n = 20 individuals with minoritized identities through participant-taken photographs, investigating identity in everyday contexts. The study aimed to understand how identity can influence daily life and social interaction, as well as validate this photovoice methodology for future work on identity expression, minoritized experience, and its potential applications for studying social cognition. Capturing local Hawai'i individuals' minoritized identity expression further allows for a cultural exploration of identity in this multicultural environment. Utilizing a qualitative photovoice methodology provided a nuanced first-person understanding of minoritized identity outside of an academically imposed view in participants' everyday lives to observe themes of community, nature, and culture. Using a photovoice method to capture visual evidence of identity expression at participants' convenience enhances ecological richness, which is crucial for innovating future methodological designs studying minoritized identity expression for social and cognitive implications.

  • Slow But Motivated: Social Learning in Open-ended Intellectual Innovation

    Social learning is known for preserving existing cultural designs, but is it also important for generating novel, effective solutions? Building on theoretical and empirical work, we identify intellectual innovation as a knowledge-based form of innovation that has been largely overlooked and examine how social learning shapes it. We introduce a riddle-crafting paradigm combined with a closed diffusion-chain design, enabling open-ended creation, multidimensional measurement, and quantifiable performance. In Experiment 1, participants in social and asocial chains composed and refined riddles distinguishing poisonous from edible mushrooms; in Experiment 2, independent participants attempted to solve them. Although the two modes showed comparable objective solvability, social riddles were more efficient and diverse in form and content. Social participants also showed greater motivation to improve, with hints of incremental gains that may emerge more fully over extended generations. Therefore, socially transmitted intellectual innovations become better tailored for communication, open to mutation, and poised for cumulative refinement.

  • Navigating Binary Identity Systems: Intersectionality and Safety Cue Monitoring by Minoritized Individuals

    Identity labels are embedded in social hierarchies that constrain how intersectional identities can be expressed, often producing dissatisfaction among minoritized individuals. Drawing on the ADDRESSING model framework, this qualitative study (n = 12) explored how individuals with multiple minoritized identities express and conceal aspects of themselves through nonverbal and interpersonal cues. Semi_structured interviews revealed that participants continuously navigate competing identity expectations and engage in vigilant safety_cue monitoring during social interactions, which we interpret as involving adaptive use of social_cognitive processes such as Theory of Mind (ToM) to anticipate prejudice. Thematic results indicate that intersectionality is associated with unique psychological burdens, particularly inner dissatisfaction when lived experience conflicts with rigid categorical labels. Within this sample, the polycultural context of Hawai_i appeared to offer potential protective factors, with participants describing collectivistic orientations as linked to greater awareness of possible prejudicial mindsets but lower reported expectations and experiences of discrimination.

  • A social inverse-planning model captures how relationship intimacy constrains action planning and interpretation

    Cognitive models of inverse planning typically treat a planner's utilities as functions of physical properties. However, many everyday actions are also shaped by utilities that are based on sociological properties, like social relationships. In this research, we extend inverse planning models to include social relationships as entities that change the utilities of actions. Using food sharing as a case study, we show that models require this sociological structure to capture human observers' judgments about actions, desires, and relationships. This work situates inverse planning models within everyday social interactions, where ongoing relationships constrain how people plan and interpret actions.

  • What Makes a Good Example? Modeling Exemplar Selection with Neural Network Representations

    Teaching requires distilling a rich category distribution into a small set of informative exemplars. Although prior work shows that humans consider both representativeness and diversity when teaching, the computational principles underlying these tradeoffs remain unclear. We address this gap by modeling human exemplar selection using neural network feature representations and principled subset selection criteria. Novel visual categories were embedded along a one-dimensional morph continuum using pretrained vision models, and selection strategies varied in their emphasis on prototypicality, joint representativeness, and diversity. Adult participants selected one to three exemplars to teach a learner. Model–human comparisons revealed that strategies based on joint representativeness, or its combination with diversity, best captured human judgments, whereas purely prototypical or diversity-based strategies performed worse. These results highlight the potential utility of dataset distillation methods in machine learning as computational models for teaching.

  • When Personal Experience and Evidence Conflict: Using Refutation to Correct the Learning Styles Misconception

    Refutation texts are effective for revising beliefs in neutral, science misconceptions. We investigated if refutation could be similarly effective for revising identity-based misconceptions (i.e., learning styles). In two studies, participants took a pretest assessing misconception endorsement, read a refutation text, and completed a posttest. In Study 1, refutation reduced belief in the misconception, but did not impact identification with a learning style. However, participants who revised their misconception demonstrated better memory for the text. Study 2 added a feedback manipulation to improve memory. Surprisingly, feedback condition did not predict memory for the text or belief change. However, Study 2 replicated key findings from Study 1 about belief, identification, and memory for the text. Open-response data from both studies suggest personal experience plays a role in the persistence of the misconception while learning explicit evidence facilitates belief revision. We discuss these findings and their implications for belief change in identity-based misconceptions.

  • Impulsivity Shapes Metacognitive Accuracy in Domain-General Ways

    Although inaccurate metacognition impairs decision-making, its sources are still debated. Early theories suggest a hierar- chical metacognitive structure where a domain-general system shapes domain-specific metacognitive processes. Recent the- ory attributes metacognitive inaccuracy to systematic errors in the computation of decision-making confidence. The present study combines these two theories to examine the extent to which general and specific systems affect decision-making con- fidence, and therefore metacognitive accuracy. The results show evidence for domain-general and domain-specific sources of metacognitive accuracy. Impulsivity has domain-general ef- fects on decision-making confidence and metacognitive accu- racy, while the tendency to consider suboptimal choices has domain-specific effects

  • Homogenized Memory: Impacts of Digital Systems on Word Recollection

    The tools we use change the way we think. Thus, the proliferation of generative AI underscores the importance of research on the cognitive implications of increasingly ubiquitous AI use. Collaboration with generative AI and other digital systems increases productivity. However, previous research has shown that AI exerts a homogenizing influence on users. This study extends the investigation of homogenization effects of digital collaboration to memory. We administer a memory task and ask participants to recall words while collaborating with generative AI, internet searches, or alone. We find that collaborating with digital systems homogenizes memory content towards words with greater lexical and category-dependent frequencies. Although the effect size is small, the repeated usage of digital systems compounds homogenization. Unchecked memory homogenization risks collective vulnerability, and has implications for recollection of past, present, and future experiences.

  • Inferring Cognitive Strategy Transitions in Multiplication Fact Learning with Graph-Based Skill Dependency Modeling

    Cognitive theories propose that multiplication learning progresses from slow calculation-like responding toward faster responding often attributed to retrieval and automaticity. We test whether a data-driven graph-based skill dependency model, GraafTel, fit with k latent components (k=3 to 6), recovers an item representation that includes a robust RT-linked "slow" component across learning stages. Using 315,690 practice trials from 540 children ages 6 to 10 in an adaptive fact-learning system, we relate item-level component requirements to response time and to a system-estimated forgetting parameter _ that is partly RT-mediated in Level 3. Across k, one component shows a strong positive association with RT across levels and encounter positions, whereas the remaining components show weaker or more stage-dependent time associations. These results suggest that graph-based latent structure can serve as a computational marker of strategy-sensitive dissociations in large-scale learning traces.

  • Enhanced Brain Activation with Reduced Functional Connectivity in Stroke Patients: An EEG Study on Short-Term Working Memory

    While stroke rehabilitation emphasizes motor recovery, the neural mechanisms underlying short-term memory (STM) impairments—and corresponding compensatory processes—remain underexplored. This study investigates the neurodynamics of STM deficits by analyzing stage-specific brain activation and functional connectivity during perception and decision-making. Utilizing frequency-resolved electrophysiological measures, we identify a distinct dissociation in neural plasticity: increased theta-band power near lesion sites suggests adaptive compensation restoring global network efficiency, yet connectivity within core STM regions remains localized and weak, particularly during memory output. In contrast, higher-order cognitive deficits are marked by significant reductions in gamma-band power and network efficiency. By providing a stage-specific analysis of functional communication, this research pinpoints where compensatory efforts succeed or fail. Our findings offer a novel framework for understanding post-stroke plasticity limits and a targeted foundation for neuromodulatory interventions for cognitive recovery.

  • We Definitely Know What Attention is

    There is a tendency to criticize traditional attention research and argue for abandoning the term attention. This paper identifies the main critique as reification, circularity, false dichotomy, and one–many problem, and offers a systematic response. Circularity refers to treating attention as both explanans and explanandum. As argued, the underlying assumption of the reification critique concerns the specification of neural processes of attention. Drawing on Marr's distinction between levels of analysis, I argue that attention is a phenomenon at the computational level, generated by neural processes, which helps address these two critiques. To address false dichotomy and one–many problem, I adopt the active inference approach and offer a definition of attention as the rule-governed filtering of inputs. This definition justifies the conceptual distinction between intention and attention, attention and preattention, and reshapes the endogenous–exogenous distinction. It also explains how various types of attention fall under a single definition.

  • The Cognitive Mechanisms Underlying Testimonial Injustice in Collaborative Decision-Making

    At its best, collaboration allows for humans to overcome individual limitations by pooling the knowledge of many minds. However, realizing the full potential of collaboration is difficult because some team members' perspectives may be systematically excluded based on their social identities. Here, we use a classic marble-urn paradigm to empirically test the underlying causes of and specific conditions that give rise to testimonial injustice. Across two experiments (one preregistered, one pilot; N = 147), we find that men are consistently (a) more confident than women in expressing their testimony, and (b) less likely than women to revise their beliefs after incorporating others' testimony, specifically when the evidence participants have access to is moderately strong (Experiment 1) and highly diagnostic (Experiment 2) but not definitive. Our work provides important insights into when and how collaborations fail, as well as the psychological root causes that underlie persistent gender discrepancies in group decision-making contexts.

  • The Temporal Discounting of Moral Judgments

    When in time do people stop judging past actions by today's moral standards? Across two experiments, we tested whether and how time—ranging from today to medieval times—affects moral condemnation. In Experiment 1, we found that as temporal distance increased, blameworthiness decreased, and participants were less likely to say their moral standards applied. However, actions high in harmfulness (e.g., sexual assault) showed attenuated discounting rates. Experiment 2 replicated and extended these findings to investigate the role of harmfulness and knowledge in shaping these judgments. We found that beliefs about what someone from the past could have known help to explain blameworthiness and whether people decide to apply their current moral standards. We also found that harmfulness moderates these effects. Together, the results suggest both harmfulness and beliefs about knowledge shape when in time, and whether, people feel their current moral standards no longer apply.

  • Is It Better Not to Know?: Adults' and Children's Understanding of How Uncertainty Affects Emotion Over Time

    How would you predict someone's feelings over time while awaiting an important decision? Would they feel better not knowing compared to receiving unfavourable news? Here we investigated adults' and 5–8-year-olds' (n = 35 / age group) ability to infer how others' emotions change over time in certain and uncertain contexts. Participants saw stories about protagonists receiving positive, negative, or uncertain news and rated their feelings immediately, and after one, two, and four weeks. Following certain outcomes, adults judged strong initial reactions that attenuated over time, whereas following uncertain outcomes, they judged reactions as initially mild but increasingly negative over time. Children showed a developmental progression toward this pattern, with 5.5-year-olds beginning to exhibit adult-like responses for certain outcomes and 6-year-olds for uncertain outcomes. These findings suggest that adults and children construe emotions as dynamic states shaped by what is known and unknown, with this capacity emerging around ages 5–6.

  • The Benefits of Staying Local: Bounded Adaptive Control by Modeling Linearly and Acting Locally

    We present a novel experiment and perspective on human adap- tive control. In our experiment, participants repeatedly adjust, zero, one or two variables with the goal of controlling a third variable, targeting a moving reward region. Across tasks, we vary the function that maps the control variables to the tar- get variable, and use computational modeling to examine how participants represent and solve the tasks. While broadly suc- cessful, we find evidence suggesting that participants fall back on projecting a locally linear monotonic relationships, while also taking control actions that are conservative, preferring to adjust one variable rather than both relative to their previous action. We suggest that this allows for robust performance even when interacting with nonlinear non-monotonic functions.

  • A Quantitative Analysis of Cultural Universals and Variation in Moral Association

    People's moral views may vary across cultures. Previous work has explored certain aspects of this cultural variation, but the extent to which moral variation and universals are reflected in the mental lexicon is underexplored. We present an analysis that quantifies cultural moral universals and variation based on large-scale word association constructed from English and Rioplatense Spanish speakers. We find that concepts related to `wrongdoers' and `medicine' are moralized more prominently among Spanish speakers, while `conflict' and `military action' are more moralized among English speakers. Furthermore, we show that the mental lexicon viewed through semantic categorization and structure explains 13% and 7% of the variances in the English and Spanish speakers in terms of moral association. Our study demonstrates that the mental lexicon, particularly the semantic network, offers a rich resource for characterizing the relations of morality and culture through the lens of their cognitive underpinnings.

  • Narrative Communication as a Learning Tool in Exploration-Exploitation Dilemmas

    Storytelling is an important avenue for social learning, proposed to have enabled humans to acquire, preserve, and transmit rich environmental knowledge over generations. Through stories, learners can be taught the causal structure of their environment and guided in decision-making by the knowledge or behavior of others in similar situations. We explored this process of learning and transmission in the context of an exploration-exploitation dilemma. Using a multi-armed bandit and narrative transmission task (N=262), we investigated if adaptive vs. non-adaptive narratives influence decision-making behavior and performance, and whether participants could identify and transmit adaptive narratives to help future participants. We found partial evidence that adaptive narratives facilitate navigating the explore-exploit trade-off. Surprisingly, we found that participants did not identify and transmit adaptive information; instead, they transmitted the narratives they received. We interpret these findings in terms of experimental limitations and cultural evolutionary processes.

  • Context-Aware Automatic Coding of Category Fluency Using LLMs

    Category fluency tasks, which involve retrieval from semantic memory, help reveal the clustered structure of human semantic memory. However, traditional coding schemes for this task rely on fixed groupings that do not capture humans' varied semantic clusters and retrieval strategies. We evaluate whether the attention patterns of Large Language Models (LLMs) can help generate tailored coding schemes that better reveal the structure of semantic memory within and across people. Using the Hills et al. (2012) data, we prefill each human-generated sequence into an LLM to derive a per-sequence coding scheme. These attention-derived groups (from middle model layers) show inter-item response time increases at subcategory boundaries -- a key prediction of optimal foraging theory -- that are up to 30% greater than when using the standard, fixed Hills et al. (2012) scheme. We also evaluate the cognitive plausibility of LLM-generated sequences, finding that they are broadly consistent with the predictions of optimal foraging theory.

  • Causal inference and learned helplessness

    Prolonged exposure to uncontrollable situations can cause individuals to become and remain dysfunctionally passive. This pattern, known as learned helplessness, is typically induced in lab settings using simple tasks, but real-world control involves complex, non-linear causal systems. In these environments, the ability to influence an outcome often diverges from the ease of achieving the specific result one wants. Moreover, ascribing agency to oneself is a non-trivial process that depends on prior mechanistic beliefs and counterfactual inference. To investigate these dynamics, we systematically manipulated structure, controllability, and reward prevalence while participants interacted with dynamic causal variables in real time. Whilst low levels of practical control reliably induced helpless behaviour, we found that this did not depend on reward prevalence or the accuracy of learners' causal beliefs.

  • Hempel's Paradox in Psychological Experiments: Empirical Tests and Large Language Model Simulations

    This paper investigates whether typical experimental practices in cognitive psychology instantiate Hempel's paradox. Using a hypothetical priming experiment, it argues that a failure to observe both concept activation and judgment bias—usually taken as disconfirming—can logically count as confirmation via the contrapositive of the hypothesis. Two studies asked participants to judge whether four possible experimental outcomes confirmed, disconfirmed, or were irrelevant to a hypothesis linking concept activation to biased judgment. Participants consistently failed to treat the contrapositive outcome as confirmatory, both in original and replication contexts. Latent class analyses suggested reliance on satisficing rather than formal logical reasoning. Simulations using a large language model showed partial alignment with human judgments, especially for disconfirmatory outcomes, but clear divergences for confirmatory cases. Overall, the results suggest that standard psychological experiments can give rise to Hempel's paradox and that both humans and LLMs systematically deviate from formal confirmation logic.

  • Interpretive constraints shape cue learning in early L2 sentence processing

    This study examines how cue accessibility shapes both sentence comprehension and grammatical learning in children acquiring Japanese as a first (L1) or second language (L2). Adopting a good-enough processing framework, we propose that learners rely on the most accessible cue to construct a plausible interpretation, rather than integrating all available cues. We conducted two experiments on the interpretation of transitive sentences and the acquisition of case markers. Experiment 1 showed that Chinese-native children learning Japanese as an L2 relied on word order prior to acquiring case-marker knowledge, whereas Japanese monolingual children showed no reliable word-order bias. Experiment 2 showed that argument-omitted input facilitated case-marker learning, with a stronger effect for bilingual children. We argue that when word order yields a plausible interpretation, attention to case markers is reduced, whereas its failure promotes reliance on case markers and facilitates learning.

  • Domain Distance as a Criterion for Selecting Instructional Analogies in Psychology

    The present research examined whether domain distance constitutes a useful criterion for selecting instructional analogies for psychological concepts. In Experiment 1, Psychology students and instructors rated interdomain analogies as more instructionally useful than intradomain analogies. Instructors' justifications suggested that this advantage stems from a higher perceived clarity and validity rather than from a higher familiarity. To explore the extent to which these ratings and justifications translate into learning outcomes, Experiment 2 assessed whether the comprehension elicited by both kinds of analogies surpasses that of a control condition receiving the literal, non-analogical explanations. While interdomain analogies boosted objective comprehension compared to the non-analogical control, intradomain analogies did not. These findings suggest that domain distance constitutes a more broadly applicable criterion than recipient familiarity for selecting instructional analogies, particularly in abstract domains such as psychology, where intradomain base analogs tend to be as abstract as their targets.

  • Informativeness Guides the Selection of Basic-Level and Superordinate Verbs

    Classic work on categorization has documented a robust preference for basic-level over superordinate labels, shaped by multiple factors including conceptual organization and pragmatic expectations of informativeness. However, it remains unclear whether similar expectations guide the evaluation of verb descriptions, whose semantic organization is less taxonomically stable and more context-dependent. In two experiments, we investigated adults' pragmatic evaluation of basic-level versus superordinate verbs. Experiment 1 showed that superordinate verbs (e.g., in "a woman is moving") were systematically rated lower relative to basic-level alternatives (running) when describing actions, despite being literally true. Experiment 2 demonstrated that this penalty was modulated by contextual relevance: superordinate verbs were judged more acceptable when visual contrast rendered them informative for identifying the target action, but remained dispreferred overall. These findings indicate that pragmatic expectations of informativeness extend to the verbal domain, while highlighting constraints imposed by conventional event construal and the flexible nature of verb abstraction.

  • In-The-Moment Detrimental Effects of Grounding in a Statistics Simulation

    We developed and experimentally tested three versions of an interactive statistics simulation designed to foster a qualitative understanding of how sample size influences statistical estimation and inference. Guided by a grounded cognition framework, two versions used perceptually rich, manipulable tokens (Token and Token-to-Rectangle) and were compared to a conventional Histogram version. In a laboratory study with undergraduates (n = 252), all groups improved similarly from pretest to posttest. Contrary to our hypotheses, however, the token-based (grounded) conditions showed less productive in-the-moment engagement during simulation use; that is, interactions were more often directed toward perceptually salient token operations rather than sampling-relevant actions. Grounded conditions also showed weaker performance on within-activity conceptual items. These findings suggest that grounding can be seductive rather than supportive in the moment, even when dynamic and perceptually rich representations are aligned with key concepts and avoid noninformative decoration.

  • Visual preferences and understanding of goal-directed actions across cultures: Evidence from Tsimane' infants and children

    Previous research with infants from mostly western, developed countries has suggested that infants from a very young age have a sophisticated understanding of the world around them. However, it remains unclear whether the methodological approaches used to explore these topics, and their corresponding findings, extend across cultures or whether they vary with the specific cultural context. In the current study, across three experiments, we explore whether Tsimane' infants and children, a farmer-foraging group living in the Bolivian Amazon, have similar visual preferences as well as expectations regarding an agent's goal-directed actions compared to infants from the United States. Our results are the first to suggest that Tsimane' infants and children performed similarly to infants from the United States, holding similar expectations for agents to engage in goal-directed actions as well as similar preferences for faces and complex patterns.

  • Why Inefficiency Survives: The Off-Book Matrix of Rationality

    Why does "inefficiency" survive? We propose that apparent inefficiency is not a cognitive defect but an artifact of evaluative mismatch. The Off-Book framework identifies two survival requirements unmeasured in laboratory ledgers: biological survival (avoiding ruin) and social survival (avoiding isolation). Our core insight: the subject's condition determines T (trial capacity), and T determines whether expected value becomes a valid reference. Empirical results (N = 100; loss aversion d = 0.99, sunk cost d = 0.47, interaction p = .004) confirm that rationality judgments shift when survival risks arise. We argue for a transition from Universal to Relative Rationality—where the scientist's N = 100 ensemble and the subject's N = 1 intuition both derive from the same principle. Under the same ecology, subjects may hold contradictory rationality due to the attributes they carry in their position. The bias is not in the subject—it is in the ruler.

  • Effects of Spatial Attention in Visual Word Recognition: Evidence from Hindi

    Early literature on spatial attention and visual word recognition has often treated the two as independent processes (Stolz & McCann, 2000) and offered mixed opinions regarding their interaction. Existing theories argued for early or late interaction, or proposed moderation by familiarity of word targets (McCann et al., 1992). However, these results are mainly derived from investigations in English and their generalizability to languages written in different orthographies is not clear. The current study investigated the interaction between spatial attention and word recognition in Hindi language which is written in Devanagari orthography—a spatially dense abugida. Across two experiments, using a spatial-cueing paradigm, we found that cueing effect was robust for words but entirely absent for non-words. Within the word category, this effect remained stable regardless of word frequency, suggesting attention interacts only with lexical-status. Our results align with a late selection view for spatial attention and visual word recognition.

  • Modeling Generalizable Physical Reasoning as Language-Guided Synthesis of Probabilistic Simulation Programs

    People use physical knowledge in remarkably flexible ways, yet most prior work models one aspect of physical reasoning at a time. Here we study the flexibility of people's physical reasoning, building on theories that language guides the construction of ad-hoc mental models grounded in capabilities for forming and manipulating representations of the world. We instantiate this theory in the Physical Reasoning via Interpretable Synthesis of Models (PRISM) framework, which uses structured reasoning with language models to generate task-specific programs in response to a question, relying on primitive functions for perceiving, editing, and simulating physical scenes. We compare PRISM against people's judgments on four distinct physics scenarios inspired by prior research, finding that it explains people's behavior as well as custom-written models from prior research while outperforming vision-language model baselines. PRISM thus provides a cognitively plausible framework for understanding how we assemble physical concepts to flexibly reason about the world.

  • Early Concepts of Caregiving: Do Children Recognize Caregivers in Third-Party Interactions?

    Caregiving relationships are central to survival, yet little is known about how humans represent them. We investigate whether 3- to 7-year-old children (N=426) and adults (N=134) attend to intimacy (who is family) and asymmetry (who is an adult) to predict who provides care. We predicted that when the potential caregivers were both adults (mom, teacher) or both children (sister, friend), participants would expect the family member to respond, and when they were both family (mom, sister), participants would expect the adult to respond. Our hypothesis was partially supported. Adults and children expected a mom to comfort and provide food when the alternative was a teacher. However, unlike adults, children did not have strong expectations that a mom would provide care when the alternative was a child's same-aged sister. Ultimately, understanding caregiving relationships as both intimate and asymmetric may develop relatively slowly.

  • Gathering Clues to Meaning: Children Recruit Past Linguistic Context to Infer Word Meanings Across Exposures in a Discourse

    Children's word learning often occurs under referential ambiguity, but linguistic context can drastically reduce this ambiguity and facilitate accurate word learning. Here, we examined whether linguistic context also facilitates word learning across exposures, testing 3-year-olds' ability to use a word's prior linguistic context to constrain subsequent mappings. To test this, we conducted three studies using semantically informative verbs, comparing children's use of linguistic context to infer meanings within a single exposure (Study 1 & 3) or across exposures (Study 2 & 3). We also varied the noun's salience in the discourse context between exposures. Results indicated that children successfully used familiar verbs to infer novel noun meanings both when the verb and referent co-occurred within an exposure and when they were separated across exposures. However, children were less successful when linguistic and referential contexts were separated. Thus, 3-year-olds use prior linguistic contexts to infer word meanings as a conversation progresses.

  • Emotional Sharing and Heart Rate Synchronization in Joint Art Appreciation: An Exploratory Field Study

    Art appreciation is often a collective, embodied process, yet research remains predominantly individual-focused. This study investigates emotional sharing and physiological coordination during joint art exploration in ecologically valid environments. We conducted two art tour studies where small groups explored artworks in culturally meaningful spaces while their emotional states, behaviors, and heart rates were monitored. Results indicate that joint appreciation is a dynamically evolving social process rather than a set of independent reactions. Crucially, dynamics varied by exploration scale: appreciating single, constrained artworks increased interest and stable physiological synchronization, whereas exploring multiple artworks across broader spaces elicited relational affects, such as harmony and gratitude, with fluctuating physiological synchronization patterns. These findings suggest that shared aesthetic experience is an emergent cognitive process shaped by embodied interaction and spatial scale. This research underscores the necessity of field-based, multi-participant approaches to capture the complexity of real-world art appreciation.

  • Why We Aesthetically Value What We Do Not Like: The Effects of Simultaneous Repetition on Judgments of Liking, Aesthetic Value, and Awe

    Repetition is ubiquitous in art and nature, yet its effects on aesthetic experience remain unclear. In a preregistered experiment, we demonstrate that increasing group size (1, 9, or 25 repeated faces) increases judgments of aesthetic value but decreases judgments of liking. We also identified the possible mechanisms underlying their dissociation. Mediation analyses revealed that both judgments (liking and value) were comparably influenced by fluency-based hedonic processing. Awe-related affective states contributed differently: threat-based awe was more closely associated with reduced liking, whereas positive awe more strongly enhanced aesthetic value. These findings highlight that liking and aesthetic value are not interchangeable indicators of a single hedonic response: liking is more tightly linked to survival-related processes, whereas aesthetic value is more strongly shaped by elaborated meanings and higher-order cognitive processing.

  • Normative Model and Connectome-Driven Target Prioritization Decision Model for Personalized Precision Modulation of Alzheimer's Disease

    Transcranial direct current stimulation (tDCS) is a promising non-invasive intervention for Alzheimer's disease (AD). However, current clinical practices rely on generic stimulation targets, which fail to account for individual differences among AD patients, leading to heterogeneous outcomes. To enable personalized neuromodulation, we propose a normative model and connectome-driven target prioritization decision (NCTPD) model. Using a functional connectivity (FC) normative model, we identify each patient's abnormal FC patterns and abnormal regions. We further quantify the coupling strength between candidate targets and abnormal regions via individual structural connectivity (SC), simulating the propagation mechanism of SC-guided current effects to multiple abnormal brain regions, thereby establishing target prioritization. Finally, we perform virtual stimulation on the prioritized targets in the digital twin brain models of 21 AD patients. Higher-ranked targets show more significant regression of abnormal FC toward the normative reference, indicating superior modulation effects. These results validate the effectiveness of the NCTPD model.

  • From Neural Signal to Embodied Action: A Cognitive Framework for BCI-Mediated Intelligence

    Brain–computer interfaces (BCIs) are typically framed as channels for neural signal transmission, yet their implications for embodiment and agency remain insufficiently theorized. We argue that BCI-mediated systems provide a productive testbed for embodied and extended cognition. Drawing on enactivist and 4E theories, we propose five interaction paradigms: (1)behavior-based embodiment through minimal sensorimotor coupling; (2)shared agency via human–machine co-negotiation; (3)structural coupling through closed-loop co-adaptive plasticity; (4)perceptual extension via integration of non-biological sensors; (5)intentional binding through temporally aligned multimodal feedback. We contend that BCIs function not as passive conduits but as active scaffolds that reconfigure embodiment, transforming tools into quasi-bodily extensions[2] while preserving cognitive sovereignty. This framework positions BCI-mediated intelligence as a dynamic interplay between neural intention and embodied action, offering a principled approach to studying agency, selfhood, and human–machine integration within 4E cognition.[23]

  • Spontaneous mental restructuring in the Connections game

    Insight is often theorized to arise from mental restructuring—abandoning an initial conceptual framework to adopt a new one. However, because of its covert nature, directly observing this process has proven difficult. We developed a modified version of the NYT Connections game, where participants organized sixteen sequentially-presented words into four categories. By manipulating word order, we induced "semantic lures''—compelling but incorrect groupings—and varied whether participants received feedback after incorrect submissions. The order manipulation reliably modulated lure formation, but did not modulate insight reports on its own. A dissociation emerged at the level of individual trajectories: among participants who formed and then abandoned the lure, those who did so without any feedback reported insight credibly more often than those who restructured following an external error signal. Our proposed paradigm provides directly observable evidence that the "aha'' experience tracks the trajectory of restructuring rather than simply reaching a solution.

  • When the Subtitle Doesn't Match: Semantic Tracking of Naturalistic Speech Predicts Cross-Modal N400 Priming for Visual Word Probes

    Researchers studying semantic processing during continuous speech have proposed temporal response function (TRF)–based measures of semantic tracking as a complement to the N400 event-related potential (ERP). However, the interpretation of these measures and their relation to the well-established N400 remain unclear. Here, we asked whether semantic tracking during naturalistic podcast listening relates to cross-modal semantic priming, as reflected in the N400 elicited by intermittently presented visual word probes. EEG was recorded while participants listened to podcast excerpts containing time-locked visual probes that varied in probe–context relatedness. In parallel, TRF decoding models were used to estimate individual semantic tracking performance from the speech stream. We found that less related probes elicited larger N400 amplitudes, and N400 sensitivity to probe–context semantic distance scaled with semantic tracking performance. These results suggest that TRF-based semantic tracking and probe-evoked N400 priming reflect a shared discourse-level semantic representation during naturalistic listening.

  • Cognitive offloading and the speedup illusion in human-AI interaction

    Large language models (LLMs) have the potential to boost human productivity by speeding up task completion---provided users know when to offload cognitive work to them. But we do not know if users are well-calibrated in estimating these potential time savings. We conducted a preregistered large-scale behavioral study (N = 1237) to characterize mismatches between expectations and reality, with a focus on simple cognitive tasks. While actual completion times between independent completion and AI-assisted completion did not differ, participants predicted AI to be significantly faster. The same bias was not observed when imagining help from another human participant. We identify a speedup illusion where people have accurate forecasts of independent completion times but significantly underestimate AI-assisted times. Additionally, time and effort dissociate: participants reported lower subjective effort with AI despite equivalent completion times. This suggests that completion time itself is not sufficient to characterize efficiency gains.

  • Two Distinct Processes in Understanding Number Word Structure

    Hurford (1975) proposed that number words across languages follow a set of phrase-structure rules (syntax), together with a constraint known as the packing strategy. The present study taught participants a novel numeral system to test whether people's mental representations of number words align with this framework. We tested speakers of English and Turkish, which contain irregularities relative to the syntax, and Chinese, which aligns closely with it. Participants' task was to generate large numbers (e.g., 27) in a base-3 artificial number system. In the training conditions, participants learned expressions consistent with both syntax and the packing strategy at different magnitudes (larger vs. smaller numbers); the baseline condition received no training. Across conditions, Chinese speakers were more likely than English or Turkish speakers to produce number words consistent with Hurford's proposal. This advantage was due to greater sensitivity to syntax, not packing. Across languages, training increased sensitivity to syntax relative to baseline, whereas sensitivity to packing did not differ between conditions. Together, these findings suggest that people represent the structure of number words through two distinct cognitive processes rather than through a single unified system.

  • Reflecting the Readers of Today: An Update of the American English Author Recognition Task

    The Author Recognition Test (ART) is a test that has been widely used for decades to estimate print exposure, due to its reliability and objectivity. In this contribution, we present an updated, balanced, and inclusive version of the English ART. We created this by constructing a contemporary item pool based on recent bestsellers and award-winning authors which balances author genre and literary level. We validate the updated English ART using item-response theory analysis and compare it with a commonly used version for American English, the Acheson et al. (2008) ART. Additionally, we examine how participant-level variables such as age and educational level interact with item-level properties such as genre and literary level to modulate recognition probability. Our results indicate that the updated ART is reliable and the test is comparable with the Acheson version in terms of difficulty and discriminative power. Additionally, reader characteristics and author characteristics systematically modulate print exposure. We discuss the implications of our findings and the suitability of the new version of the ART.

  • Fine-tuned Large Language Models predict human word associations and improve performance across lexical and semantic processing tasks.

    Recent studies have shown that LLMs can predict various aspects of semantic cognition when prompted with instructions that closely mimic those of humans. This study examines to what degree LLMs can predict human word associations, asking two previously underexplored questions: (1)~Does fine-tuning LLMs on a limited amount of human data increase prediction quality? And (2)~Does a more human-like association prediction also facilitate predicting other behaviors such as lexical processing and similarity judgments? Our findings suggest that the answer to both questions is yes. We first show that word associations can be predicted more accurately after fine-tuning. This was especially evident for the strongest associates, as the median rank of the predicted first associate was 1, a substantial improvement over previous approaches. Fine-tuning also enhances reflection of human response biases (e.g., favoring short, frequent responses). This close alignment with human associations further benefited the prediction of lexical processing and similarity judgments. Our results have implications on using LLMs fine-tuned on word associations in studies of a wide range of cognitive behavior, and on capturing factors that drive diversity in human associations.

  • Differential Associations between Autistic Traits and Eye Movement Consistency in Face and Object Recognition in Neurotypical Adults

    Recent research has found an association between reduced eye movement consistency in face recognition and autistic traits in social skills in autistic individuals, which may be related to their less well-developed visual routines for social stimuli resulting from lack of social interest. Here we showed that this association was not observed in neurotypical adults, which may be related to their well-learned visual routines that have converged to a similar consistency level, obscuring the potential association. In contrast, neurotypicals' eye movement consistency in subordinate-level object recognition was associated with autistic traits in Social Skills, and their subordinate-level object recognition performance was associated with autistic traits in Patterns/Details. These results suggested a link between autistic traits and subordinate-level object processing, consistent with a domain-general view of autism where atypical visual processing may underlie the social cognitive differences in autism. Our findings thus have important implications for early screening and intervention for autism.

  • Emotion Language Density and the Development of Internal and External Emotion Categorization

    In this study, we examine how the structure of emotion language, both in caregiver emotion language input and children's own vocabularies, relates to children's organization of emotions along an internal-external dimension. English-speaking children aged 4-6 years completed an internal-external emotion categorization task and a vignette-based emotion inference task. Caregiver emotion language was sampled during a semi-naturalistic book-reading interaction. Results indicated that repeated exposure to a small set of emotion words in caregiver input was associated with more adult-like internal–external categorization, whereas greater lexical diversity in children's own emotion vocabularies was associated with less consistent categorization during this period. Together, these findings indicate that early emotion concept development is shaped not simply by the amount of emotion language children hear or use, but also by the timing and structure of that input over development.

  • A Generative Model of Conspicuous Consumption and Status Signaling

    Status signaling drives human behavior and the allocation of scarce resources such as mating opportunities, yet the generative mechanisms governing how specific goods, signals, or behaviors acquire prestige remain a puzzle. Classical frameworks, such as Costly Signaling Theory, explain stable equilibria but struggle to explain how semiotic meaning changes based on context or drifts dynamically over time. In this work, we propose a computational theory of status grounded in the theory of appropriateness, positing that status symbols emerge endogenously through a feedback loop of social observation and predictive pattern completion. We validate this theory using multi-agent simulations of Large Language Models (LLMs) with the Concordia framework. By experimentally manipulating social visibility within naturalistic agent daily routines, our model suggests that social interactions transform functional demand into status-seeking behavior. We observe the emergence of price run-ups and positive price elasticity (Veblen effects) for both real-world luxury items and procedurally generated synthetic goods, ruling out pretraining bias as the sole driver. This work provides a generative bridge between micro-level cognition and macro-level economic and sociological phenomena, offering a new methodology for forecasting how cultural conventions emerge from interaction.

  • Manipulating Causal Priors Affects the Outcome-Density Bias

    Non-contingent learning data can lead to the (seemingly) illusory perception of causality if the potential cause and effect often co-occur -- an effect known as "outcome-density bias." Bayesian models of causal induction explain this effect as the result of a rational learning process in which the data do not fully override non-zero causal priors. Convincing evidence for this rational explanation requires the demonstration of an experimentally manipulated effect of causal priors, which has been lacking. We successfully manipulated participants' N = 300 causal priors through visually conveyed causal mechanism information. This manipulation influenced the outcome-density effect in the predicted way: in a condition inducing low causal priors, the effect disappeared. The results demonstrate the effectiveness of visual mechanism information in manipulating priors and strengthen the Bayesian view of causal induction.

  • Perceptually Grounding and Determining Truth Values of Propositions Joined by Logical Connectives: A Neural Process Account

    When descriptive sentences are linked by disjunctive logical connectives (or, xor), they may be true for multiple possible world configurations, unlike sentences linked by the conjunctive which reference a single possibility. Moreover, some of the component propositions may be false in some of the possible world configurations, while the compound sentence is still true. Determining the truth values of such sentences requires, therefore, dealing with propositions that cannot be grounded because they are false and evaluating sentences relative to multiple possible outcomes of the grounding process. Building on Dynamic Field Theory (DFT), we propose a neural process account that (1) represents affirmative and negative relational propositions linked by logical connectives, (2) autonomously drives perceptual grounding via sequential attention shifts and evaluations, and (3) determines the truth values of each component proposition and of the compound sentence by matching grounding outcomes to the possibility representations implied by each connective. The model reproduces the negation-by-truth-value interaction effect at the proposition level and extends it to the compound level.

  • SVD-Based Broad Learning System with Optimized Network Structure for EEG Emotion Recognition

    In recent years, EEG-based emotion recognition has emerged as a research hotspot in affective computing, yet mainstream deep learning approaches demand considerable time and substantial computational resources. To address these challenges, this paper proposes the HTK-BLS model, which accelerates both training and testing while maintaining competitive accuracy and reducing resource consumption. First, high-order singular value decomposition (SVD) is employed to enhance the network's capacity to capture tensor-level information from EEG frequency domain features. To further improve BLS's feature extraction capabilities, a novel stacked architecture employing approximate kernel functions is introduced with an optimized objective function to eliminate redundant nodes. In addition, weight-matrix computations are refined through truncated SVD, improving the robustness of the network. Experimental results on the DEAP dataset demonstrate that HTK-BLS achieves competitive recognition accuracy while enabling efficient training without backpropagation.

  • Self-adaptors while listening about novel and repeated topics in explanations

    In this study, we investigated the production of self-adaptors by explainees (N = 53) while listening about novel and repeated topics conveyed by explainers (N = 20) during board game explanations. Based on previous research linking self-adaptors to increased cognitive load and non-understanding in novel stimuli, we hypothesized that self-adaptors would occur more frequently during novel than repeated topics. To add granularity to topic repetition frequency, we considered topics that repeated once, twice, or three or more times. Contrary to our prediction, self-adaptors occurred significantly more often during topics repeated two or more times than during novel topics. This suggested that frequent topic repetitions could be associated with increased cognitive load, consistent with prior research linking self-adaptors to non-understanding. Further, we tested whether explainees' self-adaptors differed by topic initiator. Similar distributions of explainees' self-adaptors across topics initiated by themselves or the explainers indicated no effect compared to topic repetitions.

  • How Prior Experience Shapes Subsequent Decisions Under Uncertainty: Insights from the Iowa Gambling Task

    Experience from repeated analogous situations can bias subsequent decision-making, yet it remains unclear whether such effects arise from prior experience itself or the valence of that experience. We investigated whether prior gain or loss modulates subsequent uncertain decisions compared to a no-experience control. Participants completed a manipulated Iowa Gambling Task (M-IGT), inducing predetermined gain or loss, followed by the standard IGT. Prior experience group showed altered exploratory behavior in the early phase, with the loss group exhibiting early risk aversion. Behavioral differences across control and prior experience groups converged later in the task. The PVL-_ model revealed that prior experience initially reduced outcome sensitivity. In the later blocks, the gain group showed reduced learning rate and higher outcome sensitivity, indicating strategy stabilization. Prior loss was associated with higher choice consistency across all blocks. These findings suggest that prior experience transiently biases early exploration, while valence shapes decision consistency and learning flexibility.

  • Does Contextual Informativeness Predict Preschoolers' Word Learning from Stories?

    There is a strong relationship between book-rich environments and vocabulary size in early childhood, and for preschoolers, storybook sharing remains an important source of new vocabulary. In this work we ask, what makes a story an effective tool for word learning and vocabulary enrichment? We use data from 49 preschoolers sharing AI-generated picture books with a caregiver, with each set of stories containing 20 target words (nouns, verbs, and adjectives), to assess the feasibility of using both linguistic and visual contextual informativeness metrics to evaluate the quality of stories as support for learning new vocabulary. Results show that contextual informativeness impacts learning differently across word types: visual metrics and linguistic ground truth measures both correlate with learning for nouns and verbs (but not adjectives), and our automated metric for approximating linguistic informativeness shows significant predictive performance specifically for verbs. These results speak to the importance of considering different sources of information—including linguistic and visual information—when designing materials that support learning. We discuss these findings in the context of improving automatically generated child-directed stories to support the learning of a variety of different types of words.

  • Expectation-based Linearization: Evidence from German Word-Order and Information Structure

    Languages with flexible word order, such as German, typically privilege canonical linearizations (subject-object), with non-canonical orders (object-subject) increasing processing costs. Discourse givenness, however, has been shown to mitigate this cost, in line with expectations that given information should precede new information. The present self-paced reading study investigates how expectation based on linearization preferences influence language comprehension, examining the interaction of word order with three levels of information status (given, implied, new). Results show robust word order effects and graded information-status effects on NP1: given referents were processed fastest, new referents slowest, and implied referents exhibited intermediate processing costs. Crucially, object-first structures were disproportionally facilitated when the object was given, as compared to the other two conditions. Effects on NP2 were largely shaped by expectations established at NP1. These findings indicate that comprehenders form early expectations jointly determined by word order and discourse accessibility, supporting expectation-based accounts of linearization preferences.

  • Orthographic Competition in Word Recognition During Development: An Eye-Tracking Visual World Experiment

    How orthographic representations encode letter identity and letter position, and how these mechanisms develop with reading experience, remains a central issue for cognitive models of reading. The present study examines developmental changes in orthographic processing by assessing real-time lexical competition elicited by transposed-letter (TL) and substituted-letter (SL) competitors during visual spoken word recognition. Using the Visual World Paradigm, we tracked eye movements in third-grade children, fifth-grade children, and adult, all native speakers of Spanish. Orthographic processing was examined as a function of competitor type and word frequency. We adopted a confidence-interval estimation approach to characterize the magnitude and time course of competition effects, prioritizing effect estimation over null-hypothesis testing. Results showed increasing reading accuracy across development and a developmental increase in TL competition, while SL competition remained consistently weaker across groups. Word frequency modulated the magnitude and time course of competition effects, particularly in more skilled readers. This pattern supports predictions from the Multiple-Route Model, while frequency-sensitive effects are consistent with aspects of the Lexical Tuning Hypothesis. Together, these findings provide developmental constraints on cognitive theories of orthographic processing and highlight the value of eye-tracking methods for model evaluation.

  • Quantifying and extending the coverage of spatial categorization data sets

    Variation in spatial categorization across languages is often studied by eliciting human labels for the relations depicted in a set of scenes known as the Topological Relations Picture Series (TRPS). We demonstrate that labels generated by multimodal large language models (MLLMs) align relatively well with human labels, and show how MLLM-generated labels can help to decide which scenes and languages to add to existing spatial data sets. To illustrate our approach we extend the TRPS by adding 42 new scenes, and show that this extension achieves better coverage of the space of possible scenes than two previous extensions of the TRPS. Our results provide a foundation for scaling towards spatial data sets with dozens of languages and hundreds of scenes.

  • When Executive Function Meets High Cognitive Demand: Psychophysiological Signatures Modulated by Anxiety

    Executive functioning relies on the coordinated interaction of attention, working memory (WM) and cognitive control and is particularly vulnerable to anxiety-related disruption. Despite extensive research, few studies have simultaneously examined behavioral performance, visual attention, and physiological arousal across systematically varying cognitive demands and anxiety severity. Using a multimodal approach, we assessed WM performance, eye-movement indices, and electrodermal activity (EDA) in 83 participants, including healthy individuals with low or high trait anxiety and individuals with clinically diagnosed generalized anxiety disorder (GAD). Participants completed a dual n-back task, a baseline visual task, and executive tasks with graded complexity. The results indicated that the GAD participants showed reduced WM performance, delayed attentional engagement, altered oculomotor dynamics, and elevated autonomic arousal during continuous WM update. Anxiety severity interacted with task complexity for performance time and EDA amplitude, whereas eye-movement measures were less sensitive to increasing executive load. Together, these findings support Attentional Control Theory and Processing Efficiency Theory, highlighting anxiety-related inefficiencies under executive demand and identifying performance timing and autonomic arousal as sensitive markers of executive dysfunction in anxiety, with implications for cognitive assessment and intervention design.

  • Cross-Linguistic Variability in the McGurk Effect: A Computational Comparison of Polish and Turkish Native Speakers

    This study explores how native language influences audiovisual speech perception (the McGurk effect) by comparing Polish and Turkish speakers. Using a novel computational approach to quantify articulatory feature distances across diverse syllable pairs, the research reveals that Turkish speakers exhibit higher susceptibility to the effect, supporting the auditory intelligibility approach. However, the underlying audiovisual integration mechanism remains consistent across both languages. These findings suggest that while native linguistic structures modulate overall susceptibility, the articulatory drivers of speech perception are universal.

  • From chopped to cooked: Design and inference in physical environments

    Physical spaces are often designed to support specific uses. But how do people create such environments, and how do users infer their intended function? We propose that design and inference about design are complementary processes, grounded in a capacity to mentally simulate goal-directed actions. We tested this using "Overcooked"-style kitchens where participants either judged what a kitchen was designed for (Study 1) or designed kitchens for cooks with varying goals and beliefs (Study 2). In Study 1, participants inferred that kitchens were designed for tasks the layout made easier to complete, consistent with the prediction of a simulation-based computational model. In Study 2, participants made designs that helped cooks efficiently complete their task, adjusting their choices when cooks faced uncertainty about which task to perform. Together, these findings point towards a study of design as a cognitive activity grounded in the same mechanisms that support planning and social reasoning.

  • Perceptually Grounding and Determining Truth Values of Propositions Joined by Logical Connectives: A Neural Process Account

    When descriptive sentences are linked by disjunctive logical connectives (or, xor), they may be true for multiple possible world configurations, unlike sentences linked by the conjunctive which reference a single possibility. Moreover, some of the component propositions may be false in some of the possible world configurations, while the compound sentence is still true. Determining the truth values of such sentences requires, therefore, dealing with propositions that cannot be grounded because they are false and evaluating sentences relative to multiple possible outcomes of the grounding process. Building on Dynamic Field Theory (DFT), we propose a neural process account that (1) represents affirmative and negative relational propositions linked by logical connectives, (2) autonomously drives perceptual grounding via sequential attention shifts and evaluations, and (3) determines the truth values of each component proposition and of the compound sentence by matching grounding outcomes to the possibility representations implied by each connective. The model reproduces the negation-by-truth-value interaction effect at the proposition level and extends it to the compound level.

  • CMDA-GCN: Adaptive Graph Convolution with Confidence-Margin Semi-Supervised for Cross-Subject Working-Memory Load Decoding

    Accurate assessment of working memory load (WML) is important for cognitive neuroscience and brain-computer interfaces, as WML reflects the recruitment of attentional control and memory resources during encoding, maintenance, and updating. However, WML-related EEG is highly nonstationary: ERP signatures vary in latency and amplitude, and these variations are amplified across subjects, leading to distribution shifts. We propose CMDA-GCN, an adaptive spatiotemporal graph convolutional network based on Confidence-Margin Semi-Supervised Domain Adaptation (CMDA). The model combines multi-scale temporal convolutions with time attention to capture transient ERP dynamics and accommodate latency variability, and uses adaptive graph convolutions to learn task-dependent inter-electrode interactions. To mitigate domain drift, CMDA-GCN integrates multi-source adversarial alignment with confidence-based pseudo-label learning and a hard-sample margin constraint, jointly suppressing pseudo-label noise and sharpening decision boundaries. Experiments on two datasets show consistent performance gains, and visualizations highlight workload-sensitive time windows and distributed interactions, supporting robust and interpretable WML estimation.

  • Representational Similarity Reveals Sleep-Dependent Memory Consolidation Impaired by Simulated OSA

    Obstructive sleep apnea (OSA) is strongly linked to memory consolidation deficits, yet the neural mechanisms by which OSA disrupts memory consolidation remain poorly understood. Clarifying these mechanisms is crucial for identifying early biomarkers and developing targeted interventions to mitigate OSA-related cognitive decline. This study used a within-subject pre-sleep versus post-sleep design in healthy volunteers, comparing normal sleep with simulated OSA (via overnight airflow restriction) while measuring behavioral performance and EEG during visuospatial and verbal memory consolidation tasks. Results revealed that normal sleep produced significant post-sleep behavioral gains, enhanced representational similarity (RSA), strengthened long-range functional connectivity, and increased theta-gamma phase-amplitude coupling (PAC) with optimized phase preference. Simulated OSA abolished these improvements or even caused deterioration, accompanied by reduced RSA, weakened connectivity, and diffuse PAC patterns. These multi-level findings provide mechanistic insights for future analyses of OSA-related cognitive impairments.

  • Measuring IA Cognitive Competences: A Gap in Current Benchmarks

    To what extent do current LLM benchmarks pose the same reasoning challenges that arise in real-world settings? Instead of relying on anthropocentric notions, here we analyze the cognitive competences demanded by a benchmark in terms of (a) formal complexity (e.g. number of inference steps required to reach the desired answer) and (b) amount of contextual interpretation needed to reduce language ambiguity to a level where the required inference can be performed. Our analysis of a representative sample of benchmarks reveals that there are almost no benchmarks with both, high formal complexity and high contextual interpretation requirements. This limitation may help explain the discrepancy between the strong benchmark performance of language models and the shortcomings observed when these models are deployed in real world applications.

  • Interpretational alignment: How agents learn from physical guidance depends on how they interpret it

    In many pedagogical contexts, teachers guide learners through direct physical intervention—a modality we term state intervention. While common in human interaction, the computational mechanisms that allow agents to learn from being physically moved remain underexplored. We propose that the effectiveness of state intervention hinges on the interpretational alignment between teacher and learner: the learner must infer whether an intervention functions as a suggestion, a correction, or a non-pedagogical event. We investigated this problem using a 1D navigation task where human teachers (N=64) trained Q-learning agents under four different interpretation types. Our results show that agents learn most effectively when they interpret interventions as pedagogical signals—either as recommended actions (suggestion) or as discouraging actions (impede). In contrast, interpreting interventions as environment resets (reset) or mere interruptions (interrupt) led to significantly poorer performance and learning plateaus. Crucially, providing teachers with a visual representation of the agent's internal beliefs (Q-table) did not improve teaching effectiveness, suggesting that aligning the agent's learning rules with human intuition is more critical than information transparency. These findings highlight the importance of designing "socially-aware" learners that treat physical interaction as intentional communication.

  • Evidential Dependence and Polarization in Social Networks

    It is an inherent part of communication across social networks that the same evidence can reach us by different paths. This creates the possibility of inadvertent double counting of evidence. Recent work has shown that some degree of double counting seems to be an inevitable part of group deliberation. This makes it imperative to understand what impact evidence-reduplication has on beliefs or opinions across social networks. In this paper, we use agent-based modelling to show how and why double counting gives rise to polarization.

  • Recognizing Estimation-Relevant Structural Differences in Predictive Causal Reasoning

    Predictive causal reasoning requires estimating the probability of an effect given a cause, P(e|c). This estimate is context-dependent: a change in the causal structure may render a previous estimate biased. We investigate whether reasoners can discern estimation-relevant from estimation-irrelevant structural changes of causal models. Using the Causal Bayes Net framework, we compare situations where adding a causal link is either estimation-relevant (e.g., introducing a confounder) or irrelevant. In an experiment ( N=204), participants successfully discerned these cases, adjusting estimates in the normative direction only when structural shifts warranted it. However, we also found that few participants provided normative point-estimates. We identified two dominant heuristic strategies: a "normative change direction" strategy, where reasoners adjusted estimates but overemphasized alternative causes, and a persistent "alternative cause neglect" strategy, where participants ignored alternative causes entirely. Our findings suggest that while people recognize whether structural shifts require re-estimation, they struggle with the numeric integration of multiple causal pathways.

  • Adolescents' Cognitive and Verbal Demands During Online Irony Comprehension: An Eye-tracking Reading Study

    This study investigates how adolescents process written irony and how individual differences in vocabulary and mentalizing abilities relate to comprehension. While prior research has focused on adults, less is known about how developing readers integrate linguistic and contextual information to derive non-literal meaning. One hundred eighteen Spanish-speaking adolescents (16–18 years) read short texts containing ironic or literal utterances while their eye movements were recorded. Comprehension accuracy and reading-time measures were analyzed using mixed-effects models. Ironic utterances elicited longer regression-path durations and total reading times than literal statements, indicating increased integration and reanalysis demands. Accuracy was lower for ironic texts and positively associated with vocabulary knowledge, although this effect did not differ by condition. Mentalizing abilities did not significantly predict performance. These findings suggest that written irony comprehension in adolescence involves increased processing demands and is associated with linguistic proficiency, while evidence for mentalizing abilities in this task remains limited.

  • Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

    Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans and three large language models (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal fluency data. By applying trajectory-based NLP metrics to the items generated by 82 human participants and LLM output across eight temperature settings, we quantified three complementary dimensions: entropy (step size predictability), distance to next (successive semantic steps), and distance to centroid (global dispersion). Humans exhibited higher entropy, larger semantic steps and broader dispersion than all LLMs, indicating more variable and exploratory search. Temperature tuning produced only partial alignments, as individual metrics matched between humans and LLMs at specific settings, but no configuration reproduced the complete human profile (in all dimensions). These findings suggest that human semantic search implements a distinctive balance between local exploitation and global exploration that current model architectures fail to reproduce.

  • DBNet: A Dual-Branch Network for Single-Trial Feedback EEG Decoding and Supporting Avoidance Coupling in MDD

    Major depressive disorder (MDD) involves heightened sensitivity to negative outcomes and altered avoidance learning. While deep learning has advanced EEG-based depression detection, most studies focus on resting-state signals rather than single-trial task feedback and its behavioral relevance. Using a public feedback-locked EEG dataset from a probabilistic learning paradigm, we decode correct versus incorrect feedback trials separately in MDD and healthy controls (HC) and propose DBNet, a dual-branch network combining time-domain modeling with learnable wavelet time-frequency representations. The time branch captures feedback-related ERP dynamics and trial-wise temporal variability, whereas the time-frequency branch learns task-relevant spectral components via adaptive wavelets. DBNet outperforms baselines, with higher decoding performance in MDD. Visualizations indicate that the model emphasizes FRN/P3 time windows and fronto-midline activity, and its learned wavelet spectra show a stronger theta-band emphasis in MDD, consistent with error-monitoring processes. Subject-level representations from DBNet further exhibit stronger coupling with test-phase NoGo accuracy in MDD.

  • Drawing boundary on the mitigating effect of unintentional ignorance

    While violating prescriptive norms draws unfavorable judgments from observers, unknowingly doing so attenuates them, especially if ignorance is not willful. We test boundary conditions on this effect for legal violations. Using the Anglo-American reasonable person standard, Experiment 1 shows that ignorance of law substantially reduces blame for reasonable actions but provides little excuse for unreasonable ones. Experiment 2 examined whether regulatory expectations explain this asymmetry. Indeed, unreasonable actions generated stronger expectations that an average individual should know such action is prohibited, diminishing the exculpation conferred by being ignorant of the law. Reasonable actions, by contrast, were associated with lower regulatory expectations, making ignorance a more acceptable excuse. Counterfactual reasoning revealed parallel asymmetry: for unreasonable action, participants emphasized changes in the agent's behavior for preventing the violation; for reasonable actions, primary locus of change was external factors.

  • Modeling the fan effect during learning with a log-normal race model

    The fan effect is a well-known phenomenon in the study of associative memory. It refers to the finding that as the number of associations linked to a concept increases, retrieval becomes slower and less accurate. Modeling accounts have largely focused on reaction times or accuracy in isolation, and have primarily modeled testing-phase performance. In the present study, we examine whether fan effects already arise during learning and propose a unified account of reaction times and accuracy at the trial level using a Bayesian hierarchical log-normal race model. We report data from a Dutch fan experiment showing that fan effects are already observable during the learning phase and increase across repetitions. The hierarchical Bayesian framework allows us to capture both reaction times and accuracy simultaneously while accounting for individual differences. Our model successfully captures slower incorrect responses, consistent with predictions from activation-based theories such as ACT-R. Our results demonstrate that race models can reproduce key empirical patterns associated with fan effects during memory acquisition.

  • The influence of dependency length on expectancy: Evidence from reading times and ERPs

    Expectation-based models of sentence processing predict lower processing costs for highly expected words, whereas re- search on long-distance dependencies consistently shows that greater memory demands increase processing difficulty. The Lossy-Context Surprisal (LCS) framework integrates these ap- proaches, postulating that surprisal is determined on the basis of imperfect memory representations, predicting reduced ex- pectancy effects as dependency length increases. In both a self- paced reading and ERP experiment, we investigated whether the distance between predictive elements of the context and a tar- get word modulates expectancy as predicted by LCS. While the SPR findings revealed an interaction of these factors, providing preliminary support for unified accounts like LCS, neurophys- iological responses revealed additive effects of expectancy and memory demands, but not the predicted interaction.

  • An efficiency-based effect of frequency on lexicalization: a dyadic experiment

    Zipf's Law of Abbreviation states that the more frequent a word is, the shorter it tends to be. Zipf's own explanation of the law in terms of the trade-off between cognitive effort and communicative accuracy, i.e., Principle of Least Effort, predicts that a similar pattern holds not only for word length but also for the lexical/compositional distinction. That is, lexical forms tend to express more frequent meanings compared to compositional forms. In this paper, we report on a dyadic communication experiment that supports this prediction.

  • How do preverbal infants represent geometric shapes?

    Humans' sensitivity to geometric structure begins to emerge in infancy. However, representational strategies underlying infant shape processing remain unknown. Do infants encode shapes through abstract combinations of geometric primitives, or do they rely primarily on low-level perceptual features? To address this question, we examined shape processing in 10–12-month-olds using a novel eye-tracking intruder task. Infants viewed arrays of 4 geometric shapes consisting of 3 standard and 1 deviant shape. Shapes varied in regularity and number of sides/angles (equilateral triangle, isosceles triangle, square, rhombus, regular pentagon). We assessed whether shape complexity influenced deviant detection, defined as the proportion of looking time to the deviant shape. Results revealed a significant effect of shape, with above-chance deviant detection only for squares and real-world objects. This selective success suggests that infants' geometric representations may be sensitive to particular structural properties (e.g. right angles, symmetry, parallelism), rather than reflecting a fully-fledged combinatorial system of geometric primitives.

  • Voluntary actions influence learning of foreperiod distributions in temporal preparation

    Temporal preparation is influenced by the distributions of foreperiod durations, as reflected in differences in foreperiod _ reaction time (RT) curves. Recently, it has been shown that initiating foreperiods with voluntary actions influences temporal preparation. Here, we investigated whether this effect of actions on preparation is related to differences in learning of foreperiod distributions when intervals are self-initiated. Participants indicated the orientations of Gabors preceded by foreperiods with variable durations. Foreperiods were initiated with a voluntary keypress (action condition) or automatically, after a random interval (external condition). Across eight blocks, we manipulated the distributions of foreperiods. RT curves showed that participants learned the distributions in both conditions. However, in the external condition, learning was restricted to the first blocks, whereas in the action condition participants adapted to the changing distributions across the whole experiment. We discuss the implication of these results for the relationship between actions and temporal cognition.

  • The Sound of Suggestion: Priming Emotional Interpretation with Non-Linguistic Audio Cues

    Recent research in affective computing has shown growing interest in how multisensory integration and metaphorical expression shape emotional experience. However, the specific role of non-linguistic sounds in shaping affective information processing remains underexamined. Here we engaged thirty participants in an adapted Thematic Apperception Test featuring affectively ambiguous black-and-white line drawings. We exposed them to 300 Hz or 800 Hz quasi-white noise while selecting a color for a central figure and an emotional interpretation for the scene. The 800 Hz sound elicited significantly more positive interpretations and warm color choices, whereas the 300 Hz sound elicited negative interpretations and cool color choices. These findings show that frequency-specific auditory cues act as affective signals biasing cross-modal meaning construction. Consistent with evolutionary accounts of frequency-based emotional coding, they support using non-linguistic sound as an implicit emotional prime, offering insights for art therapy and multisensory design.

  • Knowledge-Informed Dynamic Strategy Adaptation Multi-Agent Framework for Psychotherapy

    Psychotherapy is an inherently interactive and adaptive process, in which therapists continuously adjust intervention strategies according to clients' evolving emotional states. However, most existing LLM-based psychotherapy dialogue generation methods rely on single or fixed therapeutic techniques, failing to explicitly model the turn-level intervention decision-making. To overcome this drawback, we propose a knowledge-informed dynamic strategy adaptation multi-agent framework (K-DAF), which formulates psychotherapy as a turn-level emotion perception and adaptive strategy intervention process explicitly grounded in a psychotherapy knowledge base. K-DAF mainly contains the modules of client state tracking, knowledge-grounded strategy retrieval and recommendation, and strategy-conditioned response generation. Based on K-DAF, we construct a multi-turn psychotherapy dialogue dataset, PsyDAF, and fine-tune a psychotherapy-oriented model, DAF-Chat. Comparison experiments demonstrate the superiority of PsyDAF and DAF-Chat in terms of therapeutic alliance, empathic understanding, and counseling skills. The ablation study indicates the significance of the knowledge base and client state tracking.

  • A cognitive account of how institutional pressures influence rule enforcement decisions

    Violating regulations often prompts a punitive reaction from enforcement authorities, but not always. We examine three potential contributors to how enforcement officers might make such choices: epistemic expectations (whether the violator should have known about the regulation), mentalistic inference (whether the violator acted in bad faith), and institutional pressures (top-down directives to enforce the regulation). The first two of these factors have attracted considerable attention through research on norms and causation; however, how role-based demands on enforcers shape their enforcement decisions remains under-explored. To bridge this gap, we conducted two experiments examining how institutional pressures interact with epistemic expectations and mentalistic inference. Manipulating perceived legality, prevalence, and enforcement demands, we find that institutional pressures increase enforcement but do not distort judgments about violators' knowledge or intentions. We thus identify institutional pressure as a separable source of enforcement motivation: it shapes regulatory decisions without contaminating the underlying moral evaluations of non-compliance.

  • From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents

    Wide applications of LLM-based agents require strong alignment with human social values. However, current works still exhibit deficiencies in self-cognition, emotion determination, and proactivity in social behaviors. To remedy this, we propose an evaluation scenario in which the agent needs to align with three famous psychological theories: Maslow's Hierarchy of Needs, Plutchik's Wheel of Emotion, and Moral Foundation Theory. We then design a novel value-based framework that employs GraphRAG to extract and index the prescriptive, human-written seed principles, forming a knowledge graph of emotions, needs, and moralities. The framework further conducts online query-based summarization based on a semantic retriever, with a top-k ranking mechanism. This dynamic instruction finally steers the agent such that it behaves as expected by the descriptive theories. We define the alignment metric as the ratio of expected behaviors, as well as the similarity-based metrics, when the golden responses are available. By experimenting with our method on the DAILYDILEMMAS benchmark, we observe significant performance gains over both prompting, finetuning, and retrieval-based baselines. Our method provides a basis for the social value alignment of LLM-based agents.

  • People Intuitively Schedule Tasks to Improve Collective Efficiency

    People routinely collaborate to accomplish goals more quickly, but deciding who should do which tasks, and when, poses a formidable coordination problem. How do people allocate work across collaborators? Related problems have been studied in distributed computer systems, where load balancing algorithms divide heterogeneous tasks among processors with heterogeneous capabilities. Here, we investigate whether human scheduling reflects sophisticated principles from load balancing or simpler heuristics such as taking turns. In a restaurant management task, participants (N=99) assigned customer orders to chefs with heterogeneous processing speeds, either planning schedules in advance or allocating tasks dynamically. People substantially outperformed simple heuristics, discovering efficient assignments that jointly accounted for collaborator and task heterogeneity. We also uncover systematic biases, including differences between upfront and sequential planning and a preference for completing shorter jobs early. More broadly, linking computational load-balancing theories with cognitive theories of coordination offers a valuable framework for understanding human collaboration.

  • Observing Pitch Gestures Modulates Sleep-Dependent Consolidation in L2 Lexical Tone Learning

    This study examined the effect of sleep on L2 Mandarin lexical tone acquisition by L1 English speakers and how observing pitch gestures modulates this effect. Sixty-four L1 American English speakers without prior tonal language experience learned monosyllabic L2 Mandarin words with or without observing pitch gestures and were tested at pre-test, immediate post-test, and a 24-hour delayed post-test. Higher sleep quality predicted improved accuracy and decreased response latency in the delayed post-test in the no gesture condition. Observing pitch gestures mitigated these sleep-dependent gains, suggesting that pitch gestures may enhance initial encoding so that overnight consolidation is less critical. These results extend research examining the impact of sleep on L2 tone acquisition to atonal L1 speakers and reveal interactions between multimodal learning and sleep-based memory consolidation.

  • Age Related Differences in Loss Aversion - A Meta-analysis

    Throughout lifetimes, individuals make decisions involving gains and losses. Loss aversion refers to the psychological tendency to overvalue losses over equivalent gains from a reference point, a concept essential for decision-making. The available literature on age and loss aversion yields mixed results. Hence, we conducted a meta-analysis by combining individual data sets. We performed a study-level random-effects meta-analysis (N = 20 data sets) and an individual-participant-level (IPD) meta-analysis, in a subset of data sets (n=1,565; 5 data sets) to model the relationship. We found a small and insignificant linear association between age and loss aversion. The individual-participant-level meta-analysis revealed a small effect characterized by a U-shaped relationship, with a decline in middle age. Incidentally, sex emerged as an explanatory variable for explaining heterogeneity in loss aversion. Overall, the meta-analysis motivates further empirical investigations on age-related changes in decision making across the lifespan.

  • Constructing a Concept of Age: Age Group Categories as Placeholder Structures

    Humans regularly use age to make predictions about other people. However, little is known about how children come to understand age. Children are sensitive to physical correlates of age already in infancy, but whether this indicates an early-emerging understanding of age as time spent alive remains an open question. We investigated the nature of children's early knowledge about age in the present study. We asked children to identify members of different age groups (e.g., "babies", "grown ups", etc.) and to order them chronologically. We found that 3- to 4-year-old children performed better on the identification task than the chronological ordering task. This suggests that children may form age group categories and map them onto labels early in development without yet understanding the chronological process of aging. These categories may initially function as placeholder structures, which children later reorganize into a chronological concept of age.

  • Modeling the Role of Cognitive Constraints and Allowances in the Evolution of Language Complexity

    The Linguistic Niche Hypothesis suggests languages adapt complexity to learner populations. However, languages also complexify under certain conditions, and the mechanisms behind re-complexification remain computationally underspecified within this framework. We propose a pressure-driven elaboration mechanism where child learners actively extend morphological patterns, with elaboration strength scaling inversely to current complexity levels. To account for demographic transitions, the model varies the proportion of child and adult learners within each cohort. Using an agent-based transmission model, we compared this mechanism against constant and no-elaboration alternatives. Results indicate that while constant elaboration leads to high complexity, pressure-driven elaboration stabilizes at intermediate levels, a pattern more consistent with documented post-creole languages. The model further predicts asymmetric dynamics, where simplification occurs significantly faster than recovery, and incomplete recovery to original substrate complexity. These findings provide a formal, quantitative framework for understanding how demographic transitions and functional pressures interact to shape linguistic structure across many generations.

  • The Cost of Precision: Memory Discrimination and its Relationship to Traumatic Memories

    Traumatic memory is marked by a paradox: after trauma exposure, individuals show overgeneralization of fear alongside vivid and over-specific sensory intrusions. Pattern Separation (PS), the orthogonalization of similar inputs into distinct memory traces, is typically considered protective against fear generalization, but its role in the formation of intrusions remains unclear. Here, we combined an Emotional Mnemonic Similarity Task (E-MST) with a Trauma Film Paradigm in healthy participants (n = 30-38 across analyses). Negative arousal significantly enhanced lure discrimination [F(1,32) = 39.66, p<.001, __p = .55], with the largest gains on visually distinct ("low-similarity") lures (d = 0.93). This effect was driven by suppression of gist responses and slowed reaction times, consistent with an adaptive vigilance mechanism. Critically, individual differences in negative-arousal precision for visually similar (hard to discriminate) lures predicted traumatic intrusion frequency, controlling for baseline neutral discrimination [partial r = .41, p = .026, 95% CI (.05, .68)]. Mnemonic flexibility, i.e., the magnitude of the negative-vs-neutral LDI shift, was associated with greater symptom recovery after extinction. These findings support a trade-off model in which negative arousal drives item-specific precision at the expense of contextual integration, generating potent yet decontextualized memory traces that subsequently lead to intrusions.

  • Normality and the Ordinary Concept of Disease

    Philosophical accounts of health and disease typically assume that disease involves deviation from normality, but disagree on whether normality is descriptive (e.g., statistical frequency) or evaluative (e.g., harm). Recent experimental work suggests that ordinary judgments of normality integrate frequency and value, raising the question of whether disease concepts inherit this mixed structure. Across two preregistered vignette-based experiments, we tested how frequency, valence, and explicitly stated normality influence judgments of disease, dysfunction, and health in both somatic and mental domains. In Study 1, both frequency and valence affected disease judgments, but frequency effects disappeared when controlling for normality. In Study 2, explicit normality information strongly shaped disease, dysfunction, and health judgments, with substantially larger effects than frequency. These findings suggest that ordinary disease concepts rely on a mixed notion of normality.

  • Predicting children's early word learning using egocentric videos of their learning environments

    How do features of children's environmental input predict their word learning? Prior research has suggested that distributional features of the language children hear (e.g., frequency and syntactic complexity) predict variation in early word learning, but these studies have typically relied on aggregated estimates from disjunct samples to calculate input features and word learning outcomes. In this study, we use child-specific input features derived from 1118h of at-home egocentric video recordings to show that word frequency, length in phonemes, and syntactic complexity predict word knowledge within individual children (N = 29). The predictive power of these within-child features was greater than features aggregated across children, supporting theories positing a direct relationship between input distributions and children's language learning. Exploratory analyses involving object frequency, word–object cooccurrences, and context distinctiveness did not reveal any additional effects of these predictors. The use of child-specific naturalistic data provides an avenue for more precise and comprehensive investigations of input–outcome relationships for word learning, allowing researchers to adjudicate among theories of language learning in young children.

  • Stage-dependent directed connectivity in a PFC–thalamus–hippocampus circuit during sleep fMRI with cues for memory consolidation

    Sleep is a dynamic neurophysiological process characterized by continuous transitions between wakefulness, non-rapid eye movement (NREM) sleep, and rapid eye movement (REM) sleep throughout the night. The cortical regulation-thalamic gating-hippocampal processing circuit, comprising the prefrontal cortex (PFC), thalamus (THAL), and bilateral hippocampus (HIP-L/HIP-R), may participate in information processing mode switching and be associated with memory consolidation. This study, based on overnight sleep fMRI, integrates data-driven staging evidence with directional modeling to quantify stage-dependent modulation across five stages within a four-node framework. Results indicate that connectivity undergoes systematic reorganization during sleep: cortical outbound regulation dominates during wakefulness and sleep onset, shifts toward enhanced thalamic-hippocampal axis activity during stable sleep, and declines overall during REM sleep while preserving hippocampal-related connections. This information flow migration aligns with thalamic gating and hippocampal-cortical interaction mechanisms required for sleep-related memory consolidation, providing insights for inferring memory-related mechanisms and clinical assessment.

  • A Cognitive Model for Personality and Interaction Based on the Laban–Malmgren System for Movement and Acting

    Existing personality models describe stable individual differences well, but often leave underspecified how such differences generate moment-to-moment expressive behaviour. Acting theory faces a complementary problem: performers must reliably externalise a character's internal state so that intentions, affect, and personality can be inferred from movement, speech, and posture alone. We address this gap by introducing a computational formalisation of the Laban–Malmgren System (LMS), a framework from movement analysis and acting pedagogy that links internal disposition to observable action. Building on Laban's analysis of movement qualities and Malmgren's extension for character work, we formalise four core internal and external dimensions of behaviour and develop a generative cognitive architecture in which: 1) stable trait parameters define priors over a latent expressive state, 2) state transitions are shaped by context and actions, and 3) a stochastic policy maps the latent state to observable behaviour. The resulting model provides a principled, action-centred link between internal dispositions, state dynamics, and expressive behaviour, making the theory accessible to simulation, quantitative analysis, and embodied artificial agents. We demonstrate its computational feasibility through a public proof-of-concept implementation, Persona: an interactive digital portrait in which an avatar conveys its internal state through detailed bodily expression and responds dynamically to viewer interaction (Saatchi Gallery, London, finalist for the 2024 Lumen Prize). By formalising a theory rooted in expert artistic practice, this work offers a complementary, action-centred perspective on personality modelling and lays the groundwork for future empirical validation.

  • Zero-Sum Structure of Everyday Causal Attribution

    Causal explanations for specific events (token causes) are widespread in daily life, from finding causes of a patient's symptoms to understanding why a policy was effective. Despite the importance of token causality in a wide-range of real-world scenarios, most of what is known about how people attribute causality is based on simple physics or highly constrained study set-ups. It is not yet known how people attribute responsibility when faced with both observed and unobserved factors. To address this, we conducted two experiments to examine how beliefs influence causal attribution and what information people infer indirectly. In both experiments, we found that people treated causality as a fixed resource to be split between observed and other factors, and that outcome valence influenced both causal and counterfactual judgments. On the other hand, prior beliefs about actions did not predict causal attributions.

  • Absences and Alternatives: Reasoning about Causal Exceptions Resembles, but is Not Explained by, Reasoning about Generics

    All causal relations are subject to exceptions. Theories of how causal cognition handles exceptions focus primarily on their frequency or probability, but does the nature of exceptions also influence reasoning? In two experiments, we presented adults with information about two candidate causes. Both produced the same effect in four of six cases. In the remaining two cases, one candidate produced no discernible effect while the other produced a similar but distinct effect. Participants were sensitive to this difference: despite the two candidates' identical probabilities of producing the majority effect, participants had strong preferences about which candidate was acceptable to cite in causal claims (Experiment 1) and employ in causal interventions (Experiment 2). These preferences held even in the absence of generic language and appear to be sensitive to complex and specific intuitions. This study offers novel insights into how causal reasoning understands and navigates exceptions in a rich, conceptual way.

  • Young children use mental simulation to reason about their performance

    Young children can use observed performance outcomes (e.g., successes and failures) to decide when to persist and which tasks to pursue. However, learners often face situations where outcomes either cannot be observed or are uninformative. Using a simple tablet game, we asked whether preschool-aged children can use mentally simulated outcomes to guide their decisions. In Experiment 1, the game froze mid-trial so the final outcome was unavailable. Children preferred to repeat the same game when their attempt would have resulted in success (versus failure). In Experiment 2, an on-screen agent intervened mid-trial, rendering the outcome uninformative about the child's performance. Children preferred to play the game without (versus with) the agent when their attempts would have been successful without the agent's intervention. These findings suggest that children can simulate alternative outcomes of their actions and use them to guide how they pursue future tasks.

  • When more precision is worse: Do people recognize inadequate scene representations in concept-based explainable AI?

    Explainable artificial intelligence (XAI) aims to uncover flaws in an AI model's internal representations. Can it help people recognize an AI's inability to distinguish between relevant and irrelevant features? In this study, a simulated AI classified images of railway trespassers as dangerous or not. To explain its decisions, similar images from the dataset were shown. These concept images varied in three relevant features (i.e., distance to tracks, direction, and action) and in an irrelevant feature (i.e., scene background). A feature used by the AI is retained in the concept images, otherwise the images randomize over it (e.g., same distance, varied backgrounds). Participants rated the AI more favourably when it retained relevant features. For the irrelevant feature, they did not mind in general, and sometimes even preferred it to be retained. This suggests that people may not recognize it when an AI model relies on irrelevant features to make its decisions.

  • How gender information influences spontaneous speech in context

    Although noun class has been thought to lack function, it is suggested that gender-marked articles might help to enhance the predictability of upcoming nouns, much like prenominal adjectives. Supporting this idea, in a written paradigm Hoppe et al. (2025) showed that German speakers produced fewer prenominal adjectives than English speakers, but only when articles provided gender information (which English articles do not provide). Using a modified paradigm to examine this phenomenon in speech, we found that spoken prenominal adjectives were consistently produced across both informative and uninformative contexts. Instead, planned articulatory analyses reveal that in uninformative contexts, German speakers enhance, producing longer, more acoustically informative noun phrases at second mention than at first mention; in informative contexts, German speakers reduce, showing the opposite pattern, as do English speakers. We discuss these results in terms of how uncertainty about signals and the messages they communicate influence articulation in different languages and contexts.

  • Decomposing Loss Aversion: Valuation vs. Attention in a Decision Field Theory Framework

    Loss aversion---the idea that losses loom larger than equivalent gains---is a cornerstone of behavioral economics. Yet estimates of the loss aversion parameter ($\lambda$) vary considerably across elicitation methods, suggesting this single parameter may conflate distinct psychological mechanisms. We test this hypothesis using a hybrid Decision Field Theory, which allows formal separation of valuation asymmetry ($\lambda$: losses weighted more heavily in utility) from attentional asymmetry ($\kappa$: losses sampled more frequently during deliberation). Fitting hierarchical Bayesian models to 10 datasets (N = 686; $\sim$ 140,000 trials), we find that attentional asymmetry alone provides superior predictive accuracy in the majority of cases. Simulation analyses further demonstrate that standard utility models absorb attentional variance into inflated $\lambda$ estimates. These findings suggest that the tendency to reject favorable mixed gambles may substantially reflect how losses are weighted during accumulation, rather than solely due to asymmetric valuation.

  • Recursive Statistical Learning in Adults, Children, and Macaques

    Humans readily learn language, mathematics, music, and a host of other sequentially structured conceptual systems beginning in early childhood. By contrast, even with extensive training, other species learn only the basics of many of these domains. One proposed explanation for this difference is that only humans are capable of using recursive operations to learn complex sequential concepts. Here, we build on recent work on recursive cognition to test the evolutionary and developmental origins of this ability. We tested adults, children, and rhesus macaques on a binary sequence learning task that can be learned using statistical learning recursively. Adults readily applied this recursive statistical learning strategy, but children did not, and macaques showed mixed evidence of doing so. We consider the possible reasons for the results of the children and macaques, as well as their possible implications for theories of cognition.

  • Fuzzy categorical perception as a buffer: Simulating how cognitive inefficiency enhances the robustness of emergent vowel systems

    Human categorical speech perception (CP) is traditionally modelled as deterministic to maximise individual efficiency, yet empirical evidence suggests it is intrinsically probabilistic or fuzzy. This study investigates this apparent cognitive inefficiency using an agent-based model of emergent vowel systems. We simulated populations playing Imitation Games under varying noise levels to compare the robustness of rigid (strong CP) versus fuzzy (weak CP) perception mechanisms. Contrary to the assumption that precision yields optimality, results demonstrate that the perceptual uncertainty of weak CP acts as a crucial buffer against environmental noise. Weak CP populations not only achieved higher communicative success but also maintained categories with significantly longer lifespans and lower cognitive maintenance costs. We argue that individual-level perceptual inefficiency is an adaptive feature essential for robust social coordination, functioning as a mechanism of loose coupling that prevents systemic collapse.

  • Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

    Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG). While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading. Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic processing.

  • From Distributional Structure to Meaning: Learning, Syntax, and the Emergence of Categorical and Graded Interpretations

    Adjectives often admit both categorical and graded interpretations, yet perceptual evidence alone typically underdetermines which interpretation is adopted during learning. We ask how learners converge on one interpretation and whether the learning task and syntactic structure play a causal role. Using novel adjectives mapped onto an arbitrary morph space with both discrete and continuous variance, we manipulate learning task (Classification vs. Comparison) and test generalization across syntactic frames in English and Mandarin. Learning task systematically shifts alignment between adjectives and perceptual dimensions. This alignment effect is amplified in Mandarin, where syntax overtly distinguishes categorical from graded predication. Critically, alignments do not reconfigure across post-learning tasks: category assignments remain stable, while graded information surfaces selectively in intensifier choice. Reaction-time asymmetries reveal directional markedness, with categorical-to-graded shifts incurring greater cost than the reverse. Together, the results show that adjectival interpretation emerges from the alignment between task structure and distributional variance, with syntax conditioning the strength of that alignment.

  • Predictability affects both pronoun production and interpretation: evidence from an interactive task

    When do speakers choose a personal pronoun (e.g., 'she') instead of a fuller referential expression (e.g., 'Rosalía' or 'the singer in white'), and how are these expressions interpreted? Prior work has identified two key factors: contextual cues affecting referent predictability and the structural properties of the antecedent (Kehler and Rohde, 2013). While structural properties are known to influence both production and interpretation, it remains unclear whether predictability plays a comparable role for speakers and listeners. Most previous studies have examined production or interpretation in isolation, often using non-interactive tasks. Here, we investigate both factors in an interactive communicative setting that places pressure on efficient communication. We find that referent predictability influences both pronoun production and interpretation, suggesting that predictability plays a stronger role in more naturalistic communicative contexts.

  • Pressure, Performance, and Sense of Agency: An Investigation of Sense of Agency and Task Accuracy Under High Stakes

    In this novel study we characterize factors that modulate sense of agency (SoA), defined as the feeling of control over one's voluntary actions and their outcomes. We assessed individuals' trait locus of control and state anxiety and manipulated task difficulty to evaluate whether internal traits are stronger modulators of SoA versus external performance pressure. We employed a novel task where participants press a button to stop a moving dot on a target. We manipulated the dot speed, time pressure to complete the task, and incentivized performance. We found that manipulating difficulty (faster dot speed) produced more failures and was predictive of lower explicit SoA ratings. Furthermore, SoA scores were boosted by receiving a high reward on success trials. This suggests a postdictive evaluation – increasing post-hoc feelings of control (not performance) when incentivized for success. Increasing anxiety was negatively correlated with SoA rating for failure in all manipulations. Keywords: Sense of Agency; Explicit Measures; Performance Pressure.

  • Not All Creative Language Is Created Equal: Subjective Perception and Aesthetic Appreciation of Novel Metaphor and Novel Verbing

    Experimental findings about linguistic creativity are typically based on a limited set of creative devices, most commonly metaphor. Here, we compare the subjective perception of novel metaphors in English with another prevalent creative device: "verbing," the derivation of new verbs from formally identical nouns (e.g., to museum the artifacts). We provide familiarity, meaningfulness, and aesthetic pleasure ratings for a large set of creative verbing and metaphor stimuli with matched controls (n = 486). Both devices were rated as relatively unfamiliar but meaningful. Only novel metaphors, but not verbing, were perceived as highly pleasurable. This partly conflicts with prior work on the relationship between creativity and aesthetic appreciation, which predicts that pleasure should increase for novel but meaningful stimuli. Our findings suggest that the link between creative cognition, as illustrated by different types of novel linguistic devices, and aesthetic appreciation is more complex than previously reported.

  • GILCID: Measuring Groupthink and Its Associated Phenomena in LLM Agent Collectives

    Large language models (LLMs) are increasingly deployed in multi-agent systems involving sustained interaction and collective decision-making. While prior work has documented LLMs' individual-level social behaviors such as conformity, it remains unclear whether collective interaction gives rise to group-level psychological phenomena. Therefore, we investigate whether groupthink, a classic group-level cognitive bias, emerges in LLM-based multi-agent systems. Grounded in classical groupthink theory and human groupthink experiments, we propose GILCID framework and five quantitative metrics for investigating groupthink and its typical associated phenomena: group polarization effects, spiral of silence, and pressure for conformity. Our results show that LLM collectives exhibit human-like groupthink dynamics, leading to systematic shifts in decisions and expressed stances. We further analyze key factors shaping groupthink dynamics, including task type, group size, and group authority, and propose two training-free mitigation strategies to reduce groupthink negative effects. This work provides insights into group-level cognition in LLM collectives.

  • Asking the right questions? What people learn about strangers in conversation

    When meeting somebody for the first time, how do we get to know them? In the current work, we investigate how people learn about others' personalities through the questions they ask in conversation. Across two studies, participants completed a personality inventory then were paired with an online partner for a ten-minute chat. They were either instructed to get to know their partner in freeform conversation or were provided questions to discuss. The questions were either informative or uninformative for getting to know a stranger. Participants completed the same personality inventory about their partner afterwards. We test whether choosing from informative questions enabled participants to form a more accurate impression of their partner. We find that freeform conversation improved personality predictions overall, but differences in the informativeness of the questions discussed had minimal effects on accuracy; deep questions may only be as good as the disclosures they elicit.

  • Quantitative reasoning is facilitated by cross-format comparison

    Proportional reasoning is a widespread form of quantitative thinking. Proportion also provides an opportunity to investigate the mechanisms that determine what algorithms we use to reason about quantity. Prior work shows that the visual features of proportional information (e.g., whether it is depicted continuously vs. discretely) influence whether thinkers adopt more optimal or suboptimal strategies. However, these studies investigated strategy selection when all proportions were presented in a single visual format. We hypothesize that comparing proportions across formats (continuous "blobs" vs. discrete dot arrays) will discourage suboptimal numerator strategies by making numerical information unavailable for comparison across proportions and encourage using more reliable part-whole relations within each proportion. Here, we report a study that tested this hypothesis. Participants completed a proportion comparison task where they compared proportions either within- or across-formats. Results support the hypothesis that cross-format comparison deters suboptimal numerator strategies and promotes more optimal part-whole quantitative reasoning strategies.

  • Twisted in Time: Interaction of Simultaneity and Causal Judgements in Multisensory Integration

    Multisensory integration tolerates limited temporal misalignment between multisensory signals, often characterised as a temporal window of integration (TWI). This study explores whether TWI size is influenced by perceived causal impressions, and whether auditory cues within the TWI can increase causal impressions when visual evidence is ambiguous. Results showed that launching and strong interaction conditions produced wider TWIs than weak conditions. Critically, this widening was directionally asymmetric: stronger causal conditions selectively expanded the post-contact integration window while leaving pre-contact tolerance unchanged, consistent with a forward causal prediction account. In the weak condition, near-synchronous sounds selectively increased causal and contact judgments. The finding of a close relationship between simultaneity, causal, and contact judgments supports a common inferential process that integrates multisensory evidence to estimate a shared cause.

  • A Script for Learning Scripts: Automatic Acquisition of Procedural Knowledge in Ontologically Grounded Agents

    Scripts describe events with component subevents. Ontologically grounded AI systems use scripts to represent the procedural knowledge necessary to execute complex tasks. In this paper, we present recent advances in dialogue-based script learning by agents configured in an ontologically grounded cognitive architecture. Specifically, we describe the procedure that details the sequence of events agents undertake when learning scripts in some subject domain or application. We demonstrate this script-learning script in a simulation system by tracing how an agent learns to replace a fuse through dialog with a human partner.

  • The Effect of Cognitive Load on Empathy Selection

    This study investigates the effect of cognitive load on the empathetic decision-making process through two experiments. Experiment 1 (N=30) employed a within-subject design via the Empathy Selection Task (Cameron et al., 2019). In the experiment, participants were required to choose between an empathy focused 'Feel' deck and a non-empathy focused 'Describe' deck under varying levels of cognitive demand, followed by a recognition task. Consistent with the cognitive cost hypothesis, high cognitive load significantly increased empathy avoidance. Experiment 2 (N=30) introduced prior stimulus exposure to the decision-making phase to examine its effects. In contrast to Experiment 1, results indicated no significant association between cognitive load and empathy selection when the target stimulus is pre-exposed. These findings suggest a dissociation between empathic capacity and empathic propensity. Although cognitive load can deter the initiation of empathy, prior exposure to the target can sustain empathic engagement. Suggesting that awareness of empathic demand may offset avoidance.

  • Selective Effects of Task Change-Driven Event Boundaries on Recognition Memory

    How do different forms of contextual change segment ongoing experience into discrete events? How do such boundaries influence memory for information encountered around them? While event boundaries are known to affect recognition memory, their impact on fine-grained mnemonic discrimination remains largely unexplored. Prior work (Morse et al., 2023) reported boundary-related memory effects when multiple sources of contextual change co-occurred, leaving open which type of change is critical. Building on this work, we used a design that separates two sources of contextual change that were previously confounded: shifts in stimulus domain and shifts in task context. We contrasted these by introducing stimulus-domain shifts (image category changes; Experiment 1) and task-context shifts (changes in the encoding judgment; Experiment 2). Following encoding, participants completed recognition and lure discrimination tests. Reliable boundary effects emerged only for task-context shifts, which produced a post-boundary decline in recognition accuracy. Stimulus-domain shifts did not yield comparable effects, and neither manipulation influenced lure discrimination. These findings indicate that shifts in task context alone are sufficient to induce boundary-related memory effects, highlighting a central role for changes in goals and attentional set in event segmentation.

  • Assessing Large-Scale Spatial Abilities in Virtual Reality: A Pilot Study on Map–Environment Transformations

    Large-scale spatial abilities rely on the capacity to transform between abstract spatial representations, such as maps, and embodied experiences of space. However, assessing these abilities remains challenging due to a trade-off between experimental control and ecological validity. This challenge has gained renewed relevance in light of increasing cognitive off-loading through digital navigation systems, underscoring the need for reliable instruments to measure and support spatial abilities over time. We introduce a virtual reality-based assessment targeting spatial transformations between maps and environments across multiple task facets, including landmark, location, direction, and route processing. A pilot study with students from grades 3 and 4 examines the feasibility of the assessment and its potential to differentiate between individuals and task difficulty levels. Initial results indicate systematic variation in item difficulty and meaningful interindividual differences, providing first evidence for the suitability of the proposed instrument for assessing large-scale spatial abilities.

  • Teaching With Examples to Communicate Conceptual Uncertainty

    When communicating our conceptual knowledge to others, we can either provide examples, or provide the defining rule of the concept directly. Here, we propose a normative computational theory of when people should use each mode of transmission. Our model is based on the idea that teachers who are uncertain about a target concept prefer to induce a similar structure of uncertainty in the learner's mind. Simulations show that example-based teaching more effectively communicates conceptual uncertainty, whereas rule-based teaching is more efficient when teachers are confident. A pre-registered study using a "number game" paradigm (n=100) supports this account: participants assigned the role of teachers made choices that align with the model predictions when deciding whether to communicate examples or rules. Subsequent participants (n=100) that learned from these messages showed uncertainty patterns that sensitive to the mode of communication, and were similar to the uncertainty patterns of their teachers.

  • Metaphors: Optional but Preferred for Conveying Inferential Meaning

    The status and function of metaphors in language and cognition have been a subject of a long-standing debate. We address this question by introducing inferential affordance, a theoretical construct that distinguishes linguistic form from the inferential meaning licensed by that form. Metaphors such as "black holes are messy eaters" convey inferences about abstract phenomena. Is metaphor necessary to express such inferences? The results of two studies show that inferential affordances often survive metaphor removal. At the same time, metaphors emerge as a preferred representational strategy. These findings challenge strong versions of Conceptual Metaphor Theory that view metaphors as vehicles for conveying inferential meaning (Lakoff & Johnson, 1980). They are not fully explained by views that treat metaphors as optional rhetorical devices (Pinker, 2007). Instead, our findings align closely with the view of metaphors as analogies (Holyoak, 2019), according to which metaphors are optional but selectively recruited to support inferential understanding.

  • Early-life environments shape learning and decision-making strategies across development

    During childhood and adolescence, sensitivity to the environment is often heightened, and individual differences in decision making emerge. However, it remains unclear which dimensions of experience produce these differences. Here, we focus on three candidate dimensions of early-life environments that theoretical work in reinforcement learning suggests should shape individuals' behavior. Environmental reward prevalence should govern learning from positive versus negative outcomes; controllability, reliance on reactive versus proactive control; and predictability, planning over a mental model. To test these predictions, we recruited a large sample of 10- to 25-year-olds to complete three reinforcement-learning tasks and a self-report measure of early-life experience. We found that behavior conformed to these predictions. Younger adolescents with less emotional support learned more slowly from positive outcomes, those with less control over their environment acted less proactively, and participants from less predictable environments planned more. Our results suggest that individual differences arise, in part, from rational adaptation to early-life environments.

  • Predicting Conceptual Concreteness from Sensorimotor Information Using Artificial Neural Networks

    The concreteness–abstractness distinction of concepts faces challenges from proposals that conceptual representations occupy a continuous, multidimensional sensorimotor space. Additionally, there are concerns regarding psycholinguistic norms, arguing that presenting words in isolation introduces ambiguity and that the tendency to have high variability in mid-range ratings complicates the interpretability of judgments. Here, an artificial neural network (ANN) was used to model the relationship between concreteness and sensorimotor information by incorporating mean sensorimotor strength and inter-rater variability across 11 perceptual and action dimensions to allow for nonlinear interactions. The ANN predicted graded concreteness, indicating that sensorimotor information encodes structured regularities irreducible to additive effects. Prediction error analysis revealed that inter-rater variability provides informative structure rather than noise, and that performance is not driven by lexical ambiguity. These findings suggest that conceptual representations rely on systematically organized sensorimotor systems, supporting the concreteness–abstractness distinction as a graded, probabilistic construct grounded in sensorimotor experience.

  • People Missallocate Risk-Reduction Investments in Decisions Under Extinction Risk

    Many real-world decisions involve small probabilities of irrecoverable loss, where failure eliminates the possibility of making future decisions. Such decisions under extinction risk differ fundamentally from standard risky choice. We study how people allocate resources to reduce extinction risks. Participants play 100 trials with variable extinction probabilities of 2.5%, 5%, and 10%. If a participant draws the extinction outcome, they lose all accumulated bonus payment and are precluded from further earnings. Participants can reduce extinction risk by investing earnings: for every –à0.02 invested, the extinction probability in the current trial is halved. We use dynamic programming to derive the optimal investment decision for each choice and compare it with participants' investments. Our results show that participants overinvest in reducing small risks, underinvest in reducing larger risks, and fail to adjust their investment appropriately across trials, leading them to go extinct at a higher rate compared to the optimal strategy.

  • Large Language Models Exhibit Left-Leaning Bias in Their Political Reasoning

    Large language models (LLMs) are increasingly used in political contexts. LLMs have been shown to exhibit biases in their knowledge representations of political topics, yet little is known about how these biases influence their reasoning about associated informal arguments. We investigated six popular LLMS using the Everyday Argument Assessment Task, which has been used extensively with human participants. LLMs first rated the veracity of 10 political claims that were either left- or right-leaning and then evaluated the quality of arguments supporting them. Most LLMs exhibited moderate left-leaning biases with regard to how they evaluated the veracity of claims and the quality of arguments. However, we did not find evidence that LLMs' veracity ratings of a claim predicted how arguments in support of that claim were rated. These findings suggest some popular LLMs exhibit a left-leaning bias towards socio-political discourse.

  • How The Bias-Variance Tradeoff Shapes Human Strategy Selection

    The extent to which biased decision-making is adaptive is one of the longest-debated questions in cognitive science. We investigate whether people adaptively select decision strategies following an optimal bias-variance trade-off: decision strategies that integrate (imperfect) estimates of many uncertain attributes may be unbiased but suffer from high variance, so simpler (biased) decision strategies can be more accurate when uncertainty is high. Using a new multi-attribute choice task, in which we manipulate uncertainty by varying the reliability of the estimates of different attributes across conditions, we provide an empirical demonstration that people adaptively shift their decision strategies in response to environmental uncertainty: participants were more likely to use single-attribute decision-making under high uncertainty and to integrate all attributes under low uncertainty. These results provide evidence that the bias-variance trade-off is a key principle shaping human strategy selection.

  • Early-emerging representation of social networks

    As adults, we mentally represent the social connections around us as a network. While imperfect, such representations help us navigate the complexity of our social environments. Yet, it remains unclear how such representations emerge in development. The current work examines the extent to which young children can accurately represent their social networks. We asked 3- to 5-year-old children to report their good friends in their preschool classroom (first-person networks) and with whom they think their classmates are good friends (third-party networks). Teachers also reported children's friendships within these classrooms. Social networks derived from children's third-party reports were significantly more aligned with their first-person and teacher- derived networks than simulated networks of a similar size, even after removing children's own friendships. Using an ecologically valid method, the current work provides evidence that humans, starting early in life, can systematically encode and represent their social networks.

  • Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive Load

    Negation instructions (e.g., "do not mention X") can paradoxi- cally increase the accessibility of X in humans—a phenomenon known as ironic rebound. We investigated whether Large Lan- guage Models (LLMs) exhibit similar failures. Across nine models, we measured the probability of forbidden tokens un- der varying cognitive loads (semantic, syntactic, repetition). We found that semantic distractors induce the strongest re- bound, while repetition aids suppression. Furthermore, models with sharper polarity discrimination (distinguishing neutral from negative framings) exhibited more persistent rebound. Circuit tracing revealed that rebound is driven by a sparse set of middle- layer attention heads that amplify forbidden tokens, overpow- ering early-layer suppression. We introduce ReboundBench, a dataset of 5,000 negation prompts, to enable further study of these cognitive-like failures in AI systems.

  • Shift costs in Reaction Time and Learning are dissociable in modality shifting task

    Attentional set shifting has been characterised using accuracy- or performance-based and reaction time–based shift costs across the literature, where "shift cost" is often treated as a unitary construct despite differences in how it is defined and measured. Using a GO/NOGO task with repeated intra- and extra-dimensional shifts between auditory and visual targets, we examined response (accuracy and reaction time, RT) dynamics for cross-modal shifts, using task difficulty to manipulate attentional load prior to shifts. Our prior work using this paradigm demonstrated that increased attentional demands prolong learning (to perform at sustained high accuracy) following modality shifts, raising the question: do RT-based shift costs show a similar sensitivity to task difficulty? The results demonstrated that RT-based shift costs were not increased following difficult tasks, but instead, reflected established sensory modality dominance effects. However, RT dynamics are not independent of learning: longer pre-shift RT predicted higher subsequent learning shift costs, revealing an interaction between processing time (attentional demand) and task difficulty in shaping learning.

  • Step, Switch, Repeat: How Mental Simulation Predicts Collisions

    Previous research suggests a capacity limit in dynamic mental imagery: people simulate the motion of only one object at a time (Balaban & Ullman, 2024). This raises an obvious question: how do people imagine interactions between multiple moving objects? Here, we used a novel paradigm to examine people's estimates of the timing and location of collision events that happened either between (1) two imagined moving entities, or (2) one imagined moving entity and a static patch. We propose and evaluate two mental simulation models: In the Parallel model, both entities are advanced simultaneously until collision. In the Switch model, the mind alternates between entities, updating one for several steps before switching to the other, and back again. We found that the Switch model better predicted participants' collision time estimates in the two-entity condition relative to the one-entity condition, whereas the Parallel model failed to capture this difference (Experiment 1 and Experiment 3). Moreover, the Switch model predicted systematic biases in reported collision positions that the Parallel model did not (Experiment 2 and Experiment 4). Together, these findings suggest that the mental simulation of interactions between multiple dynamic objects is better described by a Switch model, in which the mind alternates between entities, rather than simulating them in parallel.

  • Do Large Language Models Resolve Fairness-Efficiency Trade-offs Like People?

    With rapid improvements in reasoning, large language models (LLMs) have gained more prominent roles as decision-makers in social and organizational settings. This makes it important to understand whether they adhere to the same values as people about division of labor. Here, we test whether LLMs and humans are aligned on task allocation problems. These problems present a trade-off between efficiency (optimizing for metrics like output and completion time) and fairness (where each collaborator must complete roughly an equal or equitable share of the overall workload). In two experiments we find that people and LLMs vary in how closely their stated preferences between efficiency and fairness align. Humans and LLMs frequently diverged from their stated preferences and settled on similar allocations when actively determining an allocation on their own. Our work identifies a value-action gap in both humans and LLMs that influences the degree to which LLMs align with social preferences.

  • Wonderful is "More" than Terrible: Valence Influences Judgments of Magnitude

    Everyday decisions about quantity, size, and duration require people to make judgments about magnitude. Perceived energy/intensity influences such judgments, and here we report evaluative meaning is independently predictive. For example, ''wonderful" is reliably judged to be more than ''terrible." Results on antonym pairs show robust agreement across participants on two tasks: an explicit rating task (which is more?) and a task where pairs of adjectives are positioned on a line extending from an origin toward infinity. In both, energy ratings strongly predicted responses and valence predicted additional variance. We further find both energy and valence are reflected in distributional semantic models: FastText embeddings projected onto a direction for magnitude correlate strongly with human judgments, with comparable effects of valence and energy on the model's predictions. Thus the evaluative component of magnitude is recoverable from linguistic statistics. We discuss potential mechanisms underlying this association, including metaphorical mediation, linguistic markedness asymmetries, and cognitive biases toward simulating improvement.

  • Children successfully infer algorithmic properties based on only partial descriptions

    Much of early childhood learning involves learning how to correctly carry out a series of steps - an algorithm - to reach a goal. It is not known whether the general ability to make inferences about novel algorithms develops in childhood, alongside the ability to execute new algorithms, or if the inferential capacity develops later in life. We tested 4- to 8-year-old children using a set of storybooks depicting animals playing procedural games, each corresponding to an algorithm with a definite set of possible ending states. Critically, the stories depicted only the first few steps of each procedure. Then, participants were asked questions about the properties of those algorithms. We found that children successfully made multiple types of inferences about novel algorithms. Even the youngest children were already sensitive to algorithmic properties. These results provide the first evidence that children understand that algorithms are more than just a series of steps.

  • Metaphor is a Key to Abstraction: Modelling Figurative Abstraction with Distributional Semantics

    Metaphorical abstraction is the process wherein concrete concepts (e.g., key) develop lexicalized abstract meanings (e.g., important) as a result of frequent figurative contexts (e.g., communication is a key). In this study, we examine the semantic properties associated with metaphorical abstraction. Participants interpreted a series of low-familiar metaphors (e.g., religion is a raft), and generated their own related metaphors (e.g., spirituality is a raft). After this familiarization procedure, participants were presented with familiarized metaphors (e.g., faith is a raft) and non-familiarized metaphors (e.g., highways are snakes) and listed semantic properties associated with both the familiarized and non-familiarized metaphorical concepts (e.g., raft, snakes). The results showed that familiarized metaphorical concepts evoked semantic properties that were more semantically distant and more abstract than those evoked by non-familiarized metaphors or literal sentences. These results demonstrate how metaphorical abstraction results in the growth of novel and abstract semantic connections.

  • Humans Know More Than Exemplar Models Do

    Exemplar models are the canonical account of human category learning. We investigate the sufficiency of this framework across two experiments demonstrating that human learners spontaneously extract information that exemplar models fail to capture. In Experiment 1, we show that participants classify novel test items based on abstract domain-level regularities rather than summed similarity to stored exemplars. In Experiment 2, we test the design principle of selective attention by providing a perfect unidimensional cue while making a secondary dimension partially diagnostic. Contrary to the attention weight optimization of exemplar theory, participants included the redundant 'extra' features in their category representations as evidenced by single feature classification and generalization performance. These findings constitute an important challenge to core components of exemplar theory and suggest a generative process where learners build rich statistical models of the environment.

  • Finding the Conceptual Glue: Ad-hoc Categorization in People and LLMs

    What do "signal," "heat," "contact," and "director" have in common? Terms in electrical engineering? Things needed on a movie set? Appreciating what these concepts have in common requires coming up with a situation that provides a coherent grouping for them. These "ad-hoc'' categories are constructed on the spot, but what makes some better than others? Humans and large language models (LLMs) generated categories (e.g. "Things with corners'') for randomly grouped sets of nouns (e.g., "eye," "road," "table," "paper") and rated how difficult it was to link them. We then asked additional participants to rate applicability and specificity of these categories. LLMs produced more specific and applicable categories and the model and human performance gap increased for less semantically similar categories. Models, without human cognitive constraints, may be able to generate more candidates before selecting the best one, while people settle on what's "good enough."

  • Systematic Replication Study of Interleaved Math Practice

    Interleaved practice involves mixing problems of different types within assignments and distributing problems of the same type over time. A recent classroom randomized study conducted by Rohrer et al. (2020) reported large learning effects of interleaved practice relative to blocked practice. In the present study, we test whether we can replicate the effects of Rohrer et al. (2020) while also systematically extending the interleaved intervention to enhance its suitability for educational contexts. We report on results of two of three planned cohorts (n=813 middle-school students across 8 districts), and find that the interleaved condition outperforms a blocked condition (g=.18, p<.05). Implementation analysis and a cost analysis suggest that the intervention can be implemented with fidelity at low cost to schools. We discuss implications of these findings as well as the broad challenge of translating learning research to practice.

  • Semantics in Action: Nameability Predicts Verb Extension

    What determines whether a novel referent is judged as similar enough to extend an existing term, or different enough to coin a new one? We investigated how nameability and event structure influence verb extension. Participants learned novel words for unlexicalized actions varying in nameability, then saw variants differing in end state, process, or object and chose whether to extend the learned word or select an alternate. Nameabil- ity was the strongest predictor: more nameable actions saw more extension. For event structure, results replicated prior findings—end-state changes resisted extension most, followed by process changes, with object changes permitting the most extension. Extension was the dominant response overall, con- sistent with efficiency pressures favoring reuse of existing vo- cabulary. These findings suggest that lexical decisions reflect a tradeoff between the accessibility of existing representations and the degree of semantic difference introduced by variation.

  • Individual Tempo Asymmetries and Coupling Dynamics in Dyadic Musical Synchronization

    Interpersonal synchronization supports social interaction by enabling coordinated behavior and shared emotional experience, yet how emotions influence the mechanisms underlying synchronization remains poorly understood. We investigated how emotional feedback influences auditory–motor synchronization between dyadic partners with faster and slower Spontaneous Production Rates (SPR). Nonmusicians completed a joint tapping task with pre-feedback, feedback (matched or mismatched positive or negative), and post-feedback phases. Temporal asynchronies were modeled using a delay-coupled oscillator model to estimate participant-specific coupling parameters. Model comparisons between delay-coupled and zero-delay models showed that allowing for a delay significantly improved model fit. Slower-SPR partners coupled more strongly than faster partners, while feedback valence did not affect coupling strength. Coupling strength was positively associated with arousal in pre-feedback trials and with perceived synchronization success under mismatched feedback trials. These findings demonstrate the role of intrinsic frequency differences and shared feedback structure in organizing interpersonal synchronization and its subjective correlates.

  • Human Lexical Semantics are not Universal: Colexification Evidence from Western and Asian Signed Languages

    Word meanings shift synchronically, in the moment, or diachronically, over time. This paper addresses categorical shifts in meaning, focusing on metaphor and polysemy. For instance, belt could be used in the context of "belt on a dress" and then over time, it metaphorically extended to a geographic region in "belt of poverty." Prior work has documented systematic patterns of polysemy across unrelated spoken languages, suggesting universal lexical semantics. However, it is an open question whether these patterns generalize to sign languages. The present study sampled three unrelated sign languages (Japanese, German, and American Sign Languages) and compared their semantic patterns against 14 spoken languages across Western and Asian continents. Findings demonstrate that metaphorical polysemy is constrained by iconicity (e.g., belt cannot be colexified in sign languages). Our conceptual structures may not be universally fixed as we assume, and this work elucidates the mechanisms of grammaticalization unique to signed languages.

  • Can Large Language Models Reduce Vaccine Conspiracy Beliefs?

    Vaccine conspiracy beliefs contribute to vaccine hesitancy and resistance to public health interventions. Recent work suggests that large language models (LLMs) may reduce conspiracy beliefs through personalized, evidence-based dialogue. The present study extends this work by examining whether LLM-mediated conversations reduce vaccine conspiracy beliefs and whether effectiveness depends on conversational style in a country characterized by negative attitudes towards vaccination. Participants engaged in an interactive conversation with an LLM following open-ended elicitation of their vaccine conspiracy beliefs. The LLM used rational, emotional, or free-argumentation, or a neutral control conversation. Belief in vaccine conspiracies was assessed before and after the interaction using summaries of the participant's personally endorsed vaccine conspiracy beliefs. Results showed a reduction in vaccine conspiracy beliefs only when rational argumentation is used. However, no effect was observed for vaccine-related behavioral intentions. These findings confirm the conclusions of previous studies that belief change via AI requires rational engagement.

  • LLM-generated possibilities increase blame attribution

    Much of high-level cognition relies on identifying which possibilities are relevant in a situation, and representations of available alternative actions can predict the extent to which people attribute blame. Building on work showing that large language models (LLMs) sample option spaces differently from humans, we examine how moral evaluations change when potential alternative actions are supplied by an LLM rather than generated by a human. In Study 1, considering LLM-generated alternative options led participants to evaluate the agent's chosen action more negatively and to attribute more blame than when participants generated options themselves. In Study 2, participants rated LLM-generated options as having a higher value than human-generated ones and again attributed more blame after considering LLM-generated options. We additionally show that LLM-generated options occupied a narrower region of the semantic space than human-generated options. The implications for AI's influence on human cognition and judgment are discussed.

  • Prototype Abstraction Is Reduced by Concurrent Judgments of Learning

    Judgments of learning (JOLs) often modify memory but their impact on category learning remains poorly understood. In this study, we examined how concurrent JOLs affect prototype abstraction, a category learning task believed to rely on implicit memory. Participants studied low- and high-level distortions of a prototypical Blarg object either passively, while generating random numbers, or while making JOLs. They were then tested on their ability to identify novel Blargs, including the previously unpresented prototypes. Compared to the random number generators, participants who made JOLs while studying were less likely to endorse the prototypical Blarg. The JOL group was also less likely than the passive group to endorse foil items. Computational modeling suggested that higher learning rates, which could indicate effortful, explicit processing, reduced both prototype and foil endorsement. These findings suggest that JOLs disrupt prototype abstraction by pushing participants away from implicit learning, promoting explicit encoding strategies.

  • Metacognitive accuracy is associated with self-reported sleep outcomes and subjective sleepiness

    Previous work has shown that performance degrades following acute or chronic sleep loss as does the capacity to judge one's own ability to perform. In this online study, we investigate the extent to which individuals are aware of their own ability to perform and what sleep-specific factors may contribute to their judgments of their own ability. In a large online survey-based study with no manipulation of sleep or induced sleep loss, we find that higher self-rated sleepiness and poorer subjective sleep quality are associated with lower accuracy of metacognitive judgments. Self-reported sleep outcomes show an additional relationship with confidence judgments made throughout a task. This provides preliminary evidence that sleepiness and poor sleep quality may affect perceptions of performance in a fraction arithmetic task in addition to actual performance.

  • Are they human? Detecting large language models by probing human memory constraints

    The validity of online behavioral research relies on study participants being human rather than machine. In the past, it was possible to detect machines by posing simple challenges that were easily solved by humans but not by machines. General-purpose agents based on large language models (LLMs) can now solve many of these challenges, threatening the validity of online behavioral research. Here we explore the idea of detecting humanness by using tasks that machines can solve too well to be human. Specifically, we probe for the existence of an established human cognitive constraint: limited working memory capacity. We show that cognitive modeling on a standard serial recall task can be used to distinguish online participants from LLMs even when the latter are specifically instructed to mimic human working memory constraints. Our results demonstrate that it is viable to use well-established cognitive phenomena to distinguish LLMs from humans.

  • Feature-Based Navigation Strategy Selection in Complex Maze Environments

    Human navigation in maze-like environments can be approximated by distinct heuristic strategies, yet which strategy best reflects human behavior appears to depend on environmental structure rather than global optimality alone. We compare three families of navigation models—exit-oriented (EXO), target-angle (Angular), and target-distance (Distance)—across 25 grid mazes and multiple start locations. Model behavior is evaluated using success rate, path efficiency, and human-likeness defined as the correlation between model and human state-visit profiles. Across mazes, Angular heuristics are more human-like than EXO in the majority of environments, while hybrid combinations offer limited additional benefit. As an exploratory diagnostic test of whether maze structure contains information about relative strategy alignment, we label each maze by the EXO–Angular human-likeness gap (excluding ambiguous ties) and evaluate a simple leave-one-maze-out classifier using maze-level features. Results show weak but non-trivial signal under balanced accuracy, highlighting both the possible relevance of maze features and the current limits imposed by the small, strongly imbalanced dataset. These findings support a feature-based view of navigation strategy selection, in which environmental structure shapes the apparent human-likeness of different heuristics.

  • The Astonishing Ability of Large Language Models to Parse Jabberwockified Language

    We show that large language models (LLMs) have an astonishing ability to recover meaning from severely degraded English texts. Texts in which content words have been randomly substituted by nonsense strings, e.g., "At the ghybe of the swuint, we are haiveed to Wourge Phrear-gwurr, who sproles into an ghitch flount with his crurp'', can be translated to conventional English that is, in many cases, close to the original text, e.g., "At the start of the story, we meet a man, Chow, who moves into an apartment building with his wife.'' These results show that structural cues (e.g., morphosyntax, closed-class words) constrain lexical meaning to a much larger degree than imagined. Although the abilities of LLMs to make sense of "Jabberwockified'' English are clearly superhuman, they are highly relevant to understanding linguistic structure and suggest that efficient language processing either in biological or artificial systems likely benefits from very tight integration between syntax, lexical semantics, and general world knowledge.

  • An Episodic Memory Model Can Account for Reinforcement Learning in Humans

    Multiple models have been proposed for how people solve reinforcement learning tasks, with the general conclusion that people may implement a mixture of "model-free" and "model-based" strategies. Here, we sought to investigate whether a process-based episodic memory model (called EGO) can account for human behavior in two inference tasks, the classic and revaluation 2-step tasks. The EGO model was able to mimic the behavior of humans and multiple RL models solely through its parametrization. While the parameters to mimic the same RL models across tasks differed, those to mimic human behavior were similar. These results suggest that rather than through a mixture of distinct processes, people may instead implement a single memory-based process to solve inference problems. The default parametrization of this process may reflect a global optimum for the average statistics of daily life, and may be adjusted through experience and control.

  • When two is a crowd: Evidence for asymmetric number representations

    Research on morphological marking of number across the world's language points to a potential asymmetry: while many languages maintain a distinct form for the singular (as distinct from the plural), fewer languages maintain a distinct form for the dual, and very few distinguish trial or paucal. The presence of these patterns and their asymmetries could indicate that morphological systems are structured around a small set of primitives, including \textsc{one} and \textsc{two}, which are weighted differently, i.e., one is prioritized over the other. But how these primitives might relate to conceptual representations of numerosity, where small numbers have also been shown to have a privileged status, remains unclear. Further, there has been little work exploring whether representations of small numerosities might differ in their relative weight or importance in numerical categorization tasks. Here, we report the results from two experiments which aim to provide direct evidence for the role of primitives for \textsc{one} and \textsc{two} in non-linguistic and linguistic categorization. We test participants whose native languages differ in how they express number morphology (English, Japanese, Slovenian). We find strong evidence for a preference to treat 1 distinctly across all populations, and in both task types. The evidence for categorizing 2 as distinct only appears in linguistic tasks, and is generally less strong, even in a population whose language morphologically encodes the dual.

  • Insufficient Information To Assess Trustworthiness: Visualizations of Aggregates Rated Less Trustworthy than Visualizations of Individual Data

    Scientific results are often communicated with bar charts that display category means while omitting the underlying distribution of individual observations. Such visual simplification is commonly assumed to aid accessibility, yet direct evidence about how this design choice affects perceived trustworthiness is limited. We compared initial trustworthiness judgments for average-only bar charts versus sina plots, a data-showing alternative that displays individual observations within each category (Sidiropoulos et al., 2018). In a preregistered online study (N = 162; Prolific), participants viewed two visualizations (bar chart and sina plot) of the same textbook-sourced scientific result (order randomized), completed graph-reading estimates, rated each graph's trustworthiness on a 0–100 scale, rated six antecedents of trustworthiness (e.g., accuracy, completeness, bias), and provided free-response justifications. Sina plots were rated as more trustworthy than bar charts (d = .66), with 70% of participants favoring sina plots versus 17% favoring bar charts; the effect remained substantial when restricted to the first graph viewed (d = .45). Qualitative analysis suggested that bar charts were frequently criticized as providing insufficient information to judge trustworthiness, whereas sina plots more often elicited default trust. Together, these results suggest that showing individual data points can increase perceived trustworthiness even while increasing visual density.

  • Category-Specific Sensitivity Allows Modeling Variability Effects in a Similarity Framework

    Categorization is a core cognitive ability that enables humans to treat distinct entities as equivalent members of the same concept. Similarity-based models have been highly successful in explaining a wide range of categorization phenomena, but struggle to account for the category variability effect—the tendency to assign ambiguous items to more variable categories—because they do not explicitly represent distributional properties such as variability. We argue that explaining this effect does not require abandoning similarity-based models, but rather relaxing the assumption that model parameters are fixed across categories. Specifically, we consider two parametric extensions: allowing a bias parameter to favor high-variability categories, and allowing a sensitivity parameter to vary with each category's distributional precision. We directly compare both mechanisms in a model simulation across prototype and exemplar models, evaluating their categorization performance. Results consistently favor category-specific sensitivity over response bias, particularly for prototype models. We interpret this finding in terms of the broader cognitive science literature and extend the sensitivity account to prevalence-induced concept change, arguing that both variability and prevalence effects may reflect a common underlying mechanism in which sensitivity scales with the degree to which a category is specified. Together, these findings point towards a unified similarity-based account of distributional influences on categorization.

  • Effectiveness and Efficiency of Multimedia Worked Examples vs. Practice with Feedback on Learning Problem Solving

    Multimedia effects refer to the phenomenon that combining text with diagrams is generally more effective than text-only representations. The current study tests whether doing practice problems with multimedia feedback is more effective than multimedia worked examples. In the current preregistered study, 213 undergraduate students learned to solve conditional probability problems in one of three conditions: multimedia worked examples, practice with multimedia feedback, or combining both types of learning activities. Results showed that posttest performance did not differ significantly between the three conditions, but worked examples took less time than the practice-based conditions. Further, participants perceived the worked examples as more helpful and less difficult than practice with feedback activities, and receiving only worked examples led to overconfidence of how well they learned.

  • Is my textbook on the desk, or on the bookshelf? Impact of enactment and embodiment on forced-choice recognition of object locations

    Grounded cognition suggests that motor reactivation impacts memory for manipulable objects. According to the episodic memory literature, motor reactivation improves recall after action during learning episodes. Previously, we found evidence for both grounded and episodic effects, but no interaction despite similarities in the proposed mechanisms (motor reactivation). We hypothesized that an interaction may emerge if testing does not require motor action. Participants learned object-location associations for manipulable (tools) and non-manipulable (animals) stimuli. They either moved (images of) objects to locations, or observed another participant's movements. At test, participants chose between the correct location and a lure 40 degrees away. We replicated both main effects in recall accuracy, but again found no interaction. We suggest that the separation of semantic and episodic contributions is due to the specificity of the motor action involved in training. Motor information retrieved through simulation can support these learning benefits even without physical reactivation (movement).

  • Bright Pink Flamingos and Bright Pink Secretaries: Meaning from Unexpected Input

    Understanding how different sources of our semantic knowledge contribute to language comprehension, especially under unpredictable conditions, remains a challenge. Using vision-language CLIP, and language-only GPT2 representations, we examined whether these information sources differentially predict single-trial EEG responses to expected (i.e., high cloze probability) and unexpected (i.e., near-zero cloze probability) words in constraining sentence contexts. We also assess whether these effects vary across semantic processing by analyzing 100-ms intervals from 200-700 ms post word onset. GPT2 provided stronger fits for both expected and unexpected words from 300 to 500 ms. CLIP explained variance beyond GPT2 from 300 ms for expected, and from 400 ms for unexpected words. For the unexpected words, CLIP performed better than GPT2 in the 500-600 ms interval. Results show that both pure distributional and vision-informed language semantic information explain unique variance in EEG responses, with systematic differences due to word predictability as well as processing time.

  • You should have known: Epistemic Norms in Judgments of Legal Ignorance

    An influential legal maxim holds that ignorance of the law is no excuse. While legal theorists have long debated this principle, less is known about laypeople's judgments of ignorant legal violations and the cognitive factors driving them. Across two preregistered vignette studies, we examine how perceived public knowledge of a law shapes blame and punishment judgments toward ignorant agents. Study 1 shows that the exculpatory force of legal ignorance depends strongly on how widely known the law is perceived. The perceived moral wrongness of the prohibited action somewhat moderates this effect. Study 2 shows that ignorance is less excusable when a law is perceived as publicly known, regardless of agents' informational access or whether they are citizens or foreign visitors. Together, the findings suggest that legal ignorance is evaluated under robust epistemic norms centred on what agents are expected to know, rather than on individual differences in access to information.

  • Motivation and Motor Action: A Behavioral Test of the Sword and Shield Hypothesis

    According to the 'textbook' model of affective motivation in the brain, approach motivation is lateralized to the left cere- bral cortex and avoidance motivation to the right. By contrast, according to the Sword and Shield Hypothesis (SSH), the hemispheric laterality of motivation follows the way people perform approach actions (typically with the dominant hand) and avoidance actions (typically with the nondominant hand). As predicted by the SSH, neuroimaging and neurostimulation studies have shown that the cortical laterality of approach motivation reverses between right- and left-handers. The present study used a bimanual reaction time task to test a fundamental assumption of the SSH. Across two experiments, right- and left handers were faster to use their dominant hand for approach responses and their nondominant hand for avoidance responses, supporting the SSH: Hemispheric specialization for approach and avoidance motivation corresponds to manual specialization for performing approach and avoidance actions.

  • The Influence of Gender on Accuracy in Estimating Moral Behavior

    Accurate beliefs about the social world are essential to adaptive social functioning. However, people's perceptions of the social world are often inaccurate. Across three experiments (N = 991), we examine people's estimates of the prevalence of ethical and unethical behaviors for men and women. In Study 1, participants reported their own engagement in common unethical behaviors (e.g., cheating on a romantic partner), and then estimated how often men and women engaged in those same behaviors. As predicted, people overestimated the frequency of unethical behavior—an effect which was amplified for men. In Study 2, participants again overestimated the frequency of unethical behaviors, but also underestimated the frequency of ethical behaviors. These effects were again amplified for estimates of men's behaviors. Study 3 replicated these findings using hypothetical scenarios. Overall, we find consistent evidence of moral cynicism concerning the behaviors of others. People systematically assume more wrongdoing and less virtue than actually exists.

  • Differential Effects of Top-Down and Bottom-Up Attention in Category Learning

    The ability to selectively attend to diagnostic stimulus features plays a critical role in category learning. However, attention can be selectively allocated via different mechanisms: Top-Down and Bottom-Up attention. Although the two mechanisms have differential effects on multiple perceptual phenomena, their differential effects on category learning have not been established. To differentially examine the roles of Top-Down and Bottom-up attentional mechanisms in category learning, we conducted a study in which we varied the way in which attention to diagnostic features was allocated. Adult participants learned to categorize novel stimuli in one of four experimental conditions: Baseline, Top-Down, Bottom-Up, and Top-Down + Bottom-Up. Results indicate that Drawing attention to the D-feature via Top-Down and Bottom-Up attentional mechanisms leads to differential effects in category learning and generalization. Moreover, simultaneously engaging Top-Down and Bottom-Up attentional mechanisms during training leads to impaired category learning, suggesting that the mechanisms have opposing effect in learning.

  • Epistemic Extrapolation: Inferring the Fine-Grained Structure of Others' Knowledge From Minimal Evidence

    Successful communication depends on rapidly forming expectations about what others are likely to know, yet little is understood about how these expectations are constructed and updated. We introduce a novel paradigm for studying epistemic extrapolation: the ability to infer another person's knowledge from limited evidence. Across two studies, participants completed a trivia-style recognition task involving public figures and judged what a hypothetical other person was likely to know. Study 1 (n = 100) measured baseline expectations without person-specific information, assessing how closely participants' judgments tracked population-level knowledge patterns. Study 2 (n = 250) tested epistemic extrapolation across 10 conditions by providing sparse knowledge profiles (another player's response to 1-2 items) and asking participants to infer item-level knowledge. Participants showed accurate general expectations about others' knowledge and systematically updated these expectations from minimal evidence. These findings suggest that people infer others' knowledge by reasoning about the organization of knowledge within a domain.

  • The social life of STEM creativity: Social interaction and the birth and death of scientific ideas

    Modern scientific research combines exceptional innovation with increasing rates of collaboration. Creativity in STEM fields thus offers a model system for investigating collective creativity. Much of what we know about STEM innovation and collaboration comes from analyses of published articles, but this ignores the rich, hidden "backstage" of the scientific process --- the informal interactions, initial hunches, and abandoned ideas that drive STEM creativity. Here, we investigate this hidden social life of STEM creativity, focusing on how social interaction shapes the generation and abandonment of research ideas. In a survey of PhD-level STEM researchers ($N = 150$), we find that idea generation and abandonment are associated with different styles and dynamics of social interaction. We introduce a minimal mathematical model of scientific collaboration that can reproduce these empirical patterns. Understanding the informal social interactions that drive scientific progress sheds light more generally on the human capacity for collective creativity.

  • A bias-tracking model of rational political polarization.

    Evidence suggests the US population has polarized over politics in recent decades, as the issue positions of America's ideological sub-groups have moved apart. I use simulations to demonstrate a plausible mechanism by which political polarization can occur among `fully' Bayesian agents. The agents learn from the testimony of competing sources using an empirically-validated Bayesian cognitive model, and do not communicate. Polarization occurs under conditions resembling real-world political debate, and the polarization produced shares several characteristics of real-world US polarization. The core of the mechanism is that agents attempt to infer and account for the bias of the information sources, but doing so creates a feedback spiral where polarized source perceptions drive polarized beliefs and vice versa, unless there is strong disambiguating evidence in one direction or another.

  • What Makes an LLM Response Read as Empathic? A Multi-Dimensional Study of Empathic Alignment

    As large language models (LLMs) are increasingly deployed in emotional support and counseling-style dialogue, what makes their responses read as empathic remains an open question. Drawing on prior work in empathic communication, we examine four communicative dimensions of perceived empathy in LLM-generated responses: specificity, emotional reflection, affective word choice, and diversity. Across five instruction-tuned LLMs, emotional reflection (explicitly acknowledging and mirroring feelings) was the primary bottleneck across these metrics, while the other three dimensions clustered near ceiling. A preference-based learning approach that targeted reflection improved metric-based empathic quality on the four-dimensional composite. However, human raters showed no reliable preference between baseline and DPO-optimized responses, suggesting that these metric-based gains preserved, rather than enhanced, perceived response quality. We read this gap as a scope characterization: this metric set captures optimizable aspects of empathic response generation, but does not exhaust the cues humans use when judging empathy.

  • Measuring conscious contents using EEG complexity

    Measuring consciousness has been a longstanding problem. Although behavioral responses are commonly used, converging evidence indicates that they dissociate from consciousness per se. Measures of complexity applied to brain activity, such as Lempel-Ziv complexity and the perturbational complexity index, have been shown to discriminate between levels of consciousness, but less of this work has been done in the context of conscious contents. To address many of the limitations of previous work, in this study we measure participants' neurophysiological (EEG complexity), subjective, and behavioral responses in states of normal wakefulness to visual and auditory stimuli that vary in granularity of subjective characteristics, such as meaningfulness. In addition, we investigate if any of five dimensions of subjective ratings correlate with any differences in EEG complexity. This study advances our understanding of consciousness by clarifying the relationship between stimulus complexity and measures of brain complexity and phenomenology.

  • Animals don't count: Clearing up confusion on subitizing

    This paper aims to clear up terminological confusion in the study of numerical cognition, regarding subitizing (Kaufman, Lord, Reese, & Volkmann, 1949). While the term 'subitizing' was originally coined to describe the behaviour of numerate humans, it has since been applied to preverbal infants and animals. I first explore what this more inclusive interpretation entails, including the possibility of seeing subitizing as a cognitive system, and why this leads to confusion. I then argue that framing subitizing as something that is possible in subjects for which there is no conclusive evidence of there being systems that produce numerical content dilutes the concept of number. I end by discussing how the promiscuous interpretation applies to discontinuities in behaviour and in the ontogeny of numerical cognition. In each case, I show that this interpretation leads to confusion about numbers, and forces number nativism where more neutral takes are preferrable.

  • Causal language about social interactions

    Causal language is central to our understanding of social interactions—whether someone "caused" or "allowed" another's action shifts our impression of what happened. Yet models of causal language use have largely focused on physical events (e.g., billiard balls), ignoring the beliefs and preferences implicated in human action. We present a computational framework and three experiments investigating how people use causal expressions ("caused," "enabled," "allowed," "made no difference") across physical, epistemic, and preference-based interventions between agents. We find that people prefer different causal expressions across these intervention types: they describe removing a physical obstacle as a different form of facilitation than providing information. We capture people's language use with a model that selects utterances based on counterfactual simulations of events, inferences about agents' mental states, and utterance informativity. This model explains human judgments better than baseline models, suggesting that describing social influence involves reasoning about mental states, alternative actions, and alternative utterances.

  • Leveraging Speech to Identify Signatures of Insight and Transfer in Problem Solving

    Many problems seem to require a flash of insight to solve. What form do these sudden insights take, and what impact do they have on how people approach similar problems in the future? In this work, we prompted participants (N = 189) to talk aloud as they attempted to solve a sequence of five "matchstick-arithmetic" problems. These problems either all relied on the same kind of non-obvious solution (Same group) or a different kind each time (Different group). We found that Same participants improved more rapidly than Different participants, and as they improved, they talked more and talked about different things when solving later problems. Specifically, they were more likely to spontaneously categorize the problem they were working on. Taken together, these findings suggest that a hallmark of transferable insights is their accessibility for verbal report, even if the underlying precursors of insight remain difficult to articulate.

  • Mapping Cognitive-Attentional Profiles in Chinese Dyslexia: The Impact of ADHD and its Medical Treatment on Literacy Outcomes

    Chinese dyslexia frequently co_occurs with ADHD, creating complex literacy and attentional difficulties. This project examined (1) cognitive profile types in Chinese dyslexia, (2) their links to attention, and (3) whether ADHD and its treatment alter literacy_related performance. Study 1 assessed 150 children in Grades 1–3 on Chinese reading and related skills, comparing typical, dyslexia_only, and dyslexia+ADHD groups. Study 2 related attention scores to cognitive profiles within the dyslexia_only group. Study 3 tested medication effects on reading and cognitive performance in children with dyslexia+ADHD. Pure dyslexia was marked by rapid automatized naming (RAN), morphological, and orthographic deficits, whereas comorbid dyslexia+ADHD showed a consistent RAN deficit only. Within dyslexia, attentional problems were elevated mainly in children with RAN or phonological deficits. Medication improved attention, RAN, and phonological awareness but not literacy, underscoring the value of profile_based assessment, systematic ADHD screening, and integrating attention management with Chinese literacy intervention.

  • Mapping the Folk Concept of Health

    Health is widely treated as multidimensional, yet little is known about how these dimensions are structured in lay thinking or how this structure guides health-related judgments. We used a conceptual scaling approach to derive participant-specific conceptual maps positioning the term unhealthy relative to three clusters reflecting Disease, Lifestyle, and Functional Ability aspects of health. Participants' conceptual understanding of unhealthy was most closely aligned with a Lifestyle interpretation of health. We also observed substantial inter-individual differences in the degree to which participants' understanding of health was pluralistic. Alignment in participants' conceptual maps predicted how they applied the concept in a subsequent vignette task, suggesting that the structure of lay concepts constrains concept application. These findings may inform psychological theories of health, philosophical debates about its nature, and have implications for effective health communication.

  • Social-Cognitive Bases of Uniform Information Density: Evidence from Speaker Trait Variation

    According to the Uniform Information Density (UID) hypothesis, speakers prefer utterances in which information is spread evenly across the linguistic signal. Despite abundant empirical evidence for UID across multiple levels of linguistic representations, the reasons for why UID effects arise are less well understood. Here we explore a listener-oriented hypothesis that a communicative need to facilitate listener comprehension contributes to UID effects. To assess this possibility, we test whether and how information distribution patterns in naturalistic dialogues are shaped by individual differences in speakers' socio-cognitive traits. We hypothesize that if UID effects stem at least partly from speakers taking listeners' needs into account, then speakers with greater socio-cognitive abilities should exhibit stronger UID effects. Using corpus data, we show that UID is modulated by the speaker's ability to infer others' mental states, supporting the listener-oriented hypothesis of UID. We discuss the implications of the current work for refining theories of UID and, more broadly, for incorporating social cognition into accounts of the cognitive mechanisms underlying language processing.

  • A Self-directed Expanded Judgment Paradigm: Isolating the Pairwise Mechanism of the Attraction Effect

    The attraction effect—where a decoy option increases preference for a dominating target—is a cornerstone of context-dependent choice, yet it is paradoxically fragile. Sequential accounts propose that context effects depend on which pairwise comparisons are emphasized during deliberation, but existing tests often confound comparison availability with memory/recency. We introduce a Self-directed Expanded Judgment paradigm in which participants repeatedly unblur and judge available option pairs, allowing information search to be observed and constrained. In Experiment 1, we validate the method by replicating a positive attraction effect. In Experiment 2, we causally manipulate comparison availability by disabling specific pairwise links while equalizing cumulative stimulus exposure at the decision stage (top-up control). Consistent with the preregistered hypothesis, the target's relative share (RSTew) differed reliably between conditions, with the disabled Competitor-Decoy pair producing a positive attraction effect. In exploratory analyses, the disabled Target-Decoy pair produced negative attraction (repulsion). These findings provide causal evidence that the accessibility of specific pairwise comparisons can modulate—and under some conditions reverse—context effects.

  • Invisible walls: how pedestrians navigate interactional territories in public spaces

    Public space simultaneously supports many activities, including conversation, free speech, and transit. However, individuals may have competing desires to use the same parcel of public space. Spatial coordination problems caused by these conflicts are partially solved with proxemic norms – tacit rules which govern the social use of space. But, while prior work has focused on how pedestrians orient to personal space, little is known about how pedestrians orient to interactional territory. In a field experiment with 1,138 participants, we show that pedestrians are acutely aware of others' body orientation and avoid walking between people who are approximately facing each other. We also show evidence of collective norm violations: pedestrians are more than twice as likely to breach a proxemic norm when preceding pedestrians have done so. This study shows how spatial claims in public space are coordinated and offers a promising framework for investigating the psychology of norms.

  • Cognitive Debiasing via Disentangled Pre-training for Cross-Subject EEG Emotion Recognition

    Capturing shared cognitive processes across individuals is crucial for cross-subject EEG emotion recognition. Existing studies overlook the subjective experiential component introduced by individual cognitive biases, causing redundant information to be entangled with the extracted emotion-related features. To this end, we propose a cross-subject EEG emotion recognition method named CDDP (Cognitive Debiasing via Disentangled Pre-training). Specifically, during pre-training, cognitive disentanglement and contrastive learning are leveraged to extract subject-invariant intrinsic features (essentially emotion-relevant features) and subject-specific bias features (subjective experiential knowledge) from EEG signals. Meanwhile, a variational autoencoder (VAE) is introduced to generate diverse bias features in the latent space, alleviating overfitting caused by limited source domain EEG data. During fine-tuning, an emotion classifier is jointly trained with the pre-trained subject-invariant intrinsic feature encoder to capture discriminative emotion representations. CDDP achieves accuracies of 88.62% and 75.38% on the SEED and SEED-IV datasets, respectively, demonstrating state-of-the-art performance.

  • Cinematic Techniques for Controlling Attentional Focus Enhance Emotional Engagement in Interactive Narratives

    To enhance emotional engagement in interactive narratives, we propose an Arousal Induction Model that integrates cinematic techniques with the psychological theory of misattribution of arousal. The model combines attentional focus control (using over-the-shoulder shots and match-cuts) with arousal induction via auditory and visual effects.We evaluated this model through an RPG-based experiment ($N=49$). Results showed that the arousal-induced staging significantly improved participants' alignment with the protagonist's interpersonal preferences and increased subjective story evaluations. Physiological analysis (CSI) confirmed successful sympathetic activation during staged events.These findings suggest that simultaneously intervening in cognitive attention and physiological arousal can effectively deepen emotional engagement and improve experience quality. This research provides a framework for designing more moving and immersive interactive narrative experiences.

  • Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality

    Human reasoning is characterized by the rational use of resources to optimize performance under constraints. Recently, inference-time scaling has improved the reasoning performance of Large Language Models by increasing test-time computation. Instruction-tuned (IT) models explicitly generate long reasoning traces, whereas Large Reasoning Models (LRMs) are trained via reinforcement learning to discover reasoning paths that maximize accuracy. However, it remains unclear whether resource-rationality can emerge from such scaling without explicit rewards related to computational costs. We introduce a Variable Attribution Task (VAT) in which models infer which variables determine outcomes given candidate variables, input–output trials, and predefined logical functions. By varying the number of candidate variables and trials, we systematically manipulate task complexity. Both models exhibit a transition from brute-force to analytic strategies as complexity increases. IT models degrade on XOR and XNOR functions, whereas LRMs remain robust. These results suggest that resource rationality can emerge from inference-time scaling itself, even without explicit cost-based rewards.

  • Syntactic Prominence and Pragmatic Bias in Turkish Subject Anaphora: Humans and Large Language Models

    We examined whether large language models (LLMs) responded to syntactic prominence of discourse entities and pragmatic biases in Turkish subject anaphora resolution similar to human judgments. Using an offline comprehension task with native speakers, we showed that null and overt pronouns responds to syntactic prominence and pragmatic bias. We then evaluated several autoregressive LLMs on the same materials. While GPT-4o correlated with human responses, LLaMA-4 more closely approximated the interaction between syntactic prominence and pragmatic biases observed in human data, raising questions about the relationship between model performance, scale, and human fit. We also found that all models mostly differed from humans in response variability, with model responses tending to be more deterministic. Finally, the results indicated partial but limited alignment between human and model anaphora resolution in Turkish.

  • Intuitive Judgement with Analytical Oversight: A Dual-Process Architecture for Complex Relational Understanding

    Complex relational understanding tasks such as Document-level Relation Extraction require resolving semantic ambiguity amid a quadratic explosion of entity pairs, causing severe class imbalance and false negatives. Although large language models show strong reasoning ability, directly applying them to DocRE is inefficient and prone to hallucination. Inspired by dual-process theories of human cognition, we propose DocRE-Thinker, a hybrid framework integrating fast intuition with controlled deliberation. A fine-tuned backbone serves as System 1, efficiently generating candidate relations while estimating uncertainty. A frozen large language model acts as System 2 and is invoked only for high-uncertainty cases, performing uncertainty arbitration and attribute-based rule induction to justify relations. These symbolic rules are converted into supervision through a text-gradient feedback mechanism, enabling System 1 to internalize analytical reasoning in a rule-aware manner. Experiments on DocRED and Re-DocRED show state-of-the-art performance with significant recall gains across benchmarks, validating robust and efficient cognitively grounded relational reasoning.

  • Primates solve risky decision-making problems via spatial computations

    What are the neural mechanisms underlying risky decision making? How do the implementation details of these mechanisms impact observable behaviour? Here we hypothesise that non-human primates (NHPs) are re-using circuits for spatial navigation to solve decision making problems. This is motivated by recent findings of grid coding of probability and magnitude in NHP frontal cortex (Bongioanni et al., 2021; Veselic et al., 2025). Grid cells are an optimal code for space (Dorrell et al., 2023; Fiete et al., 2008; Sorscher et al., 2023) but this form of representation is not normatively optimal for risky-decision making. However, this might be a small price for efficiency: by embedding problems in space, NHPs can recycle spatial solutions to quickly solve new problems. First, we show that this spatial framework can replicate classical prospect theoretic findings (Kahneman & Tversky, 1979). Then, we show that the framework also predicts unique biases in the behaviour of NHPs, which we find in two NHP datasets. Finally, we find that a spatial model best fits individual subjects' behaviour in both NHP and human data. In total, we provide a mechanistic implementation for risky decision making, which accurately predicts systematic biases. This proposal makes novel predictions in terms of both behaviour and neuronal activity. Keywords: Decision making; prospect theory; grid cells; frontal cortex; computational modelling; resource rationality.

  • Stats Wars: Return of the Bayesians

    In view of the ongoing replication crisis and various frustrations about classical null-hypothesis significance testing, calls for a statistical reform have increased. Often, it has been argued that the Bayesian approach offers a promising alternative to classical statistics, but at the same time, both approaches are sometimes hard to compare. This paper offers some basic steps for establishing a more decisive common ground and take steps towards resolving some fundamental conceptual issues, by introducing the frameworks of epistemic accuracy and efficient experimentation. Efficiency (i.e. the ratio between the accuracy of conclusions, and the experimental resources needed) is of central importance to good scientific inquiry. This paper shows how considerations of efficiency and accuracy motivate a broadly Bayesian approach, which however can also make room for frequentist tenets, and thereby push us towards a more unified and powerful framework for scientific inference.

  • The Role of Comparison and Category Confusability on Concept Acquisition

    Comparison can support category learning by highlighting shared similarities or diagnostic differences between categories. We examined how comparison type (match vs. contrast vs. single-item control) influences featural category learning and whether these effects depend on category confusability (low vs. high). On training trials, the match condition was shown co-presented exemplars from the same category, whereas those in the contrast condition were shown co-presented exemplars from different categories; the control condition was shown individual exemplars. Participants learned artificial categories in a mixed design with comparison type manipulated between subjects and confusability within subjects. Learning was assessed via training performance, embedded endorsement, and posttest endorsement measures. Across all measures, contrast outperformed match, whereas control performed comparably to contrast. Notably, even with fewer exposures, single-item presentation still outperformed match presentation, suggesting that co-presenting exemplars does not always lead to better learning.

  • Valence is more prominent than truth conditional meaning

    Valence, the positive/pleasant or negative/unpleasant value of information, grounds our experience with the world. What specific role does valence play in constructing meaning? Some have downplayed it, centering the referential relationship between words and the world, "truth-conditions". Others have taken valence to be central to meaning. Empirically, the jury is still out. Yet, few have examined how these two types of information compete for prominence within the same word (the Valence Prominence Hypothesis, that valence usually wins). In a first set of studies, we find that valence is significantly faster than truth-conditional categorization. In the second set of studies, we find that valence and not truth-conditional similarity drive judgments. Further, we tested valence against 4 different informational domains fundamental to development and cognition. Animacy stood out as the domain which most closely parallels Valence in its prominence, suggesting a link between meaning and cognitive fundamentality.

  • The Role of Holes in Whack-a-Mole: Investigating Procedural Learning with Graded Statistical Structures Across State Space Sizes

    Procedural learning---the implicit acquisition of motor and cognitive skills through repeated practice---has been proposed to support abilities from typing to language acquisition. Traditional serial reaction time paradigms study procedural learning using stimulus sequences with simple underlying statistical structures: deterministic or with binary probability levels, involving only 4 possible states. We introduce a novel paradigm where participants learn sequences with graded probability levels, generated from transition matrices with a continuous gradient of transition probabilities. We found robust effects of both surprisal (transition probability from previous state) and entropy (uncertainty of next state) on reaction times across 5 state space sizes (4-8), with no interaction between surprisal and entropy, indicating learners tracked specific transition probabilities even at high-uncertainty states. These findings demonstrate that procedural learning is sensitive to fine-grained statistical structure and scales beyond traditional 4-state paradigms, supporting theoretical proposals linking procedural learning mechanisms to skill acquisition in complex domains like language.

  • DEGSleepNet: A Dual Evolving Graph Network for EEG-Based Single-Channel Automatic Sleep Staging

    The development of brain–computer interfaces (BCI) has provided a solid data foundation for sleep staging. Transformer-based methods have achieved moderate sleep staging by capturing temporal dependencies within time-frequency representations. However, Transformers with substantial computational overhead are limited to capturing pairwise temporal dependencies rather than group-wise temporal dependencies. To address these issues, we propose a Dual Evolving Graph Network (DEGSleepNet) for sleep staging. DEGSleepNet consists of multiple Mamba with one-dimension inverse discrete cosine transform (MambaIDCT) blocks and dual evolving graph (DEvoGraph) blocks. DEvoGraph consists of a time EvoGraph and a frequency EvoGraph. Time EvoGraph sequentially captures both group-wise and pairwise temporal dependencies, while frequency EvoGraph does the same for frequency dependencies. Extensive experiments demonstrate DEGSleepNet achieves the best performance on public sleep datasets. Notably, DEGSleepNet has few model parameters, making it suitable for sleep monitoring on wearable devices.