From 3D Reconstruction to Spatial Intelligence: Bridging Vision Foundation Models with the 3D World
Skip to main content
eScholarship
Open Access Publications from the University of California

UC San Diego

UC San Diego Electronic Theses and Dissertations bannerUC San Diego

From 3D Reconstruction to Spatial Intelligence: Bridging Vision Foundation Models with the 3D World

Abstract

Understanding the 3D world from visual observations is a fundamental challenge in computer vision, with broad implications for robotics, autonomous driving, and embodied AI. Despite remarkable progress in 2D vision, a critical disconnect persists: existing foundation models are anchored to the 2D image plane and lack the geometric grounding needed to reason about 3D environments. At the root of this disconnect lies a data problem — 3D data is expensive, difficult to scale, and insufficient to support the kind of large-scale training that has driven 2D progress. This thesis addresses these challenges across three interconnected thrusts: scalable 3D data curation, spatial reasoning, and real-world 3D perception. First, I tackle the 3D data bottleneck from two complementary perspectives. From a reconstruction standpoint, I present COLMAP-Free 3D Gaussian Splatting, which eliminates the dependence on structure-from-motion preprocessing, enabling scalable 3D reconstruction from casually captured video. From a generative standpoint, I present HOIDifusion, which leverages diffusion models to generate realistic 3D hand-object interaction data, addressing a data gap that is particularly critical for embodied AI. Second, I investigate how to bridge the gap between 2D foundation models and the 3D world. I propose a feature distillation framework that transfers 3D structural knowledge into 2D vision-language models, equipping them with spatial reasoning capabilities without requiring full 3D supervision at inference time. Third, I push 3D perception toward robust real-world deployment. I address category-level 6D object pose estimation in the wild through semi-supervised learning, and demonstrate how scene reconstruction can serve as structural mapping priors to improve 3D object detection for autonomous driving. Together, these contributions form a cohesive path from raw visual data to semantically rich, spatially aware intelligence — advancing the long-term goal of machines that can perceive, reason about, and act within the 3D world.