- Main
From 3D Reconstruction to Spatial Intelligence: Bridging Vision Foundation Models with the 3D World
- Fu, Yang
- Advisor(s): Wang, Xiaolong
Abstract
Understanding the 3D world from visual observations is a fundamental challenge in computer vision, with broad implications for robotics, autonomous driving, and embodied AI. Despite remarkable progress in 2D vision, a critical disconnect persists: existing foundation models are anchored to the 2D image plane and lack the geometric grounding needed to reason about 3D environments. At the root of this disconnect lies a data problem — 3D data is expensive, difficult to scale, and insufficient to support the kind of large-scale training that has driven 2D progress. This thesis addresses these challenges across three interconnected thrusts: scalable 3D data curation, spatial reasoning, and real-world 3D perception. First, I tackle the 3D data bottleneck from two complementary perspectives. From a reconstruction standpoint, I present COLMAP-Free 3D Gaussian Splatting, which eliminates the dependence on structure-from-motion preprocessing, enabling scalable 3D reconstruction from casually captured video. From a generative standpoint, I present HOIDifusion, which leverages diffusion models to generate realistic 3D hand-object interaction data, addressing a data gap that is particularly critical for embodied AI. Second, I investigate how to bridge the gap between 2D foundation models and the 3D world. I propose a feature distillation framework that transfers 3D structural knowledge into 2D vision-language models, equipping them with spatial reasoning capabilities without requiring full 3D supervision at inference time. Third, I push 3D perception toward robust real-world deployment. I address category-level 6D object pose estimation in the wild through semi-supervised learning, and demonstrate how scene reconstruction can serve as structural mapping priors to improve 3D object detection for autonomous driving. Together, these contributions form a cohesive path from raw visual data to semantically rich, spatially aware intelligence — advancing the long-term goal of machines that can perceive, reason about, and act within the 3D world.