Toward Multimodal Representation Learning for Visual Understanding and Healthcare
Skip to main content
eScholarship
Open Access Publications from the University of California

UC Berkeley

UC Berkeley Electronic Theses and Dissertations bannerUC Berkeley

Toward Multimodal Representation Learning for Visual Understanding and Healthcare

Abstract

This dissertation investigates fundamental challenges in multimodal representation learning, focusing on the gap between visual perception and structured reasoning in both computer vision and healthcare domains. Despite remarkable advances in vision-language models, current systems can recognize visual patterns with high accuracy yet struggle to reason about the underlying physical structure of what they see. The work is presented in two main parts. The first part examines the spatial reasoning capabilities of vision-language models through three interconnected studies. We first introduce X-Planner, an MLLM-based planning agent that decomposes complex image editing instructions into executable sub-tasks with spatial grounding, revealing that current models are spatially grounded yet geometrically blind. We then present All-Angles-Bench, a comprehensive benchmark for multi-view understanding that systematically diagnoses this geometric blindness across state-of-the-art models, identifying cross-view object mismatch and spatial misalignment as root failure modes. Finally, we propose GASP, a framework that injects geometric-aware spatial priors directly into VLM architectures through visual correspondence and depth consistency supervision, achieving genuine 3D understanding that generalizes across benchmarks without requiring 3D question-answering data. Recognizing that the same perception-to-reasoning gap arises in clinical settings, the second part applies multimodal learning to ocular surface disease diagnosis. We first develop a standardized AI-driven pipeline for meibography image analysis that bridges the gap between curated research data and real-world clinical data through automated morphological feature extraction. Building on these quantified visual features, we introduce MDPipe, a multi-modal diagnostic pipeline that combines visual morphology data with clinical metadata through LLM-based reasoning, providing both accurate diagnoses and clinically sound rationales preferred by human clinicians. In summary, this dissertation demonstrates that factoring multimodal intelligence into visual perception, structured quantification, and language-based reasoning provides a principled and effective paradigm for building reliable AI systems across diverse domains.

Main Content

This item is under embargo until August 31, 2027.