Thinking Through Action: Prediction, Planning, and Metacognition in Problem-Solving
Skip to main content
eScholarship
Open Access Publications from the University of California

UC San Diego

UC San Diego Electronic Theses and Dissertations bannerUC San Diego

Thinking Through Action: Prediction, Planning, and Metacognition in Problem-Solving

Abstract

Cognitive science has traditionally studied perception, planning, and action as separate processes. However, real-world problem-solving rarely follows this neat separation, instead involving complex interactions among all three components.How do these processes work together when humans solve problems in the physical world? Rather than viewing physical problem-solving as a linear sequence—as the classical view suggests—this dissertation explores how perception, planning, and action are deeply intertwined and mutually supportive. I examine three key interactions in physical problem-solving: visual task decomposition, perceptual encoding for prediction, and metacognitive self-knowledge that orchestrates how we perceive, plan, and act.

In Chapter 1, I investigate how perception and planning interact through visual subgoals in physical assembly tasks. I show that humans select subgoals that strategically balance progress toward goals with planning effort. Through experiments comparing human behavior against computational models, I find that people consider both immediate and future planning costs, demonstrating how perceptual decomposition can make planning more tractable.

In Chapter 2, I examine the perceptual foundations of physical reasoning using the Physion benchmark of simulated physical interactions. Comparing humans against AI architectures reveals that a significant bottleneck in AI physical understanding lies in the initial perceptual encoding of scenes. While some models approach human-level accuracy with idealized inputs, they lack the flexibility humans demonstrate with raw visual data, underscoring the importance of robust scene understanding for physical prediction.

Chapter 3 explores metacognition through introspective self-prediction in large language models (LLMs). I investigate whether these systems can predict their own responses to hypothetical scenarios—a fundamental capacity for orchestrating problem-solving processes. I show that LLMs can learn to forecast their own behavior more accurately than even more powerful external models, suggesting potential for developing metacognitive abilities that could support coordination of perception, planning, and action.

Through examining these key interactions—task decomposition, perceptual encoding, and self-knowledge—in both humans and AI systems, this dissertation illustrates how perception, planning, and action work together to enable efficient problem-solving. These findings contribute to our understanding of human physical problem-solving and provide insights for designing artificial systems that can interact with the physical world with the robustness and resourcefulness humans routinely demonstrate.