Skip to main content
eScholarship
Open Access Publications from the University of California

UC Berkeley

UC Berkeley Electronic Theses and Dissertations bannerUC Berkeley

Data Augmentation and Generation for Improving Learning-based Autonomous Driving Models

Abstract

Learning-based models have evolved to become the main-stream in autonomous driving research and industry recent years, where data serve an irreplaceable role. However, real-world driving data are costly, requiring expensive data collection instruments and substantial human labors. Therefore, there is a critical need to maximize the informational yield of available datasets and investigate synthetic data generation as a scalable alternative, in order to enable the scaling of learning-based driving models while circumventing the prohibitive costs associated with massive real-world data acquisition.In this thesis, we first introduce heuristic approaches for the direct augmentation and synthesis of driving data in Chapter 2, aiming to empirically validate the significance of data scaling for autonomous driving models. Specifically, we employ prescriptive planning models to synthesize physically plausible vehicle trajectories and apply geometric transformations to augment existing map topologies. These operations, while simple enough and involve no additional data collections, introduce critical diversity to the motion data, both for the trajectory side and for the map side. By using these augmented data to pre-train a motion prediction model, we observe substantial gains in the motion prediction accuracy.Beyond direct geometric augmentations, enriching the semantic modalities of existing datasets serves as a critical strategy to extract deeper insights from available driving logs and therefore to maximize data efficiency. In Chapter 3, we augment motion data with natural language descriptors that characterize complex agent interactions. These interactions, especially those governed by intricate traffic rules and subtle human intentions, are often notoriously difficult to distill from raw numerical trajectories alone. Natural language, however, provides a powerful medium to render these latent constraints explicit. This research led to the development of WOMD-Reasoning, which, at the time of its release, constituted the most extensive language-annotated dataset for real-world driving interactions. Our results demonstrate that this multimodal enrichment not only catalyzes motion prediction accuracy but also provides a necessary layer of interpret-ability, allowing models to better utilize the causal logic embedded within existing driving scenarios.The advent of generative AI has established a new paradigm for synthesizing photorealistic, vision-based driving data, which is imperative for training robust end-to-end autonomous driving systems. While these generative frameworks offer the theoretical capability to produce high-fidelity scenarios, their practical utility is severely constrained by prohibitive training costs. Consequently, achieving training efficiency is a prerequisite for scalable data generation. To circumvent this computational bottleneck, Chapter 4 introduces Immiscible Diffusion, a novel image-noise assignment strategy that fundamentally accelerates the convergence of diffusion models, achieving up to a 3× efficiency gain. Building upon this foundation, Chapter 5 provides a rigorous analysis of the feature-level suboptimalities inherent in standard diffusion processes, while elucidating the underlying mechanisms that drive the efficiency enhancements of our approach. Guided by these insights, we propose generalized implementations of Immiscible Diffusion that further elevate training efficiency to > 4×, establishing a versatile framework adaptable to diverse model architectures and training environments. Collectively, these algorithmic innovations enable cost-effective, high-fidelity data synthesis, serving as the computational backbone for scalable vision-based driving data synthesis and closed-loop driving simulation.In summary, this dissertation establishes a cohesive framework for data-centric autonomous driving, progressing from heuristic data augmentation and semantic data enrichment to fundamental optimizations in generative modeling for data generation. By systematically resolving bottlenecks in data utilization and generation, this work unlocks the potential of synthetic data as a scalable engine for autonomous driving systems. These contributions lay the essential groundwork for a future where robust driving models are trained within high-fidelity, closed-loop simulations, minimizing the reliance on labor-intensive real-world data collection.