Skip to main content
eScholarship
Open Access Publications from the University of California

UC Berkeley

UC Berkeley Electronic Theses and Dissertations bannerUC Berkeley

Learning to Solve Long-Horizon Tasks with Formal Logic and Structured Feedback

Abstract

Recent advancements in policy learning have enabled impressive results in robotics and cyber-physical systems (CPS), ranging from dexterous manipulation to bipedal locomotion. Despite these successes, learned robot policies still struggle to integrate high-level reasoning and feedback with low-level control, limiting their efficacy in complex, long-horizon tasks. This dissertation introduces methodologies that enable learned policies to leverage the same skills that humans use to reliably achieve complex goals in the real world: devising high-level plans, learning from partial successes, creating cooperative strategies, and dynamically adjusting behavior through trial-and-error.The central insight across the methods contained in this thesis is to use formal, symbolic structure to express both complex tasks and their feedback mechanisms. In doing so, policies can learn high-level behavior from precise, structured plans while maintaining the expressivity of their underlying (deep) parameterization to learn low-level control. We explore this insight across a number of settings. We first demonstrate how exploiting the under- lying structure of tasks specified via formal logic allows reinforcement learning agents to learn complex, long-horizon tasks in both the single and multi-agent setting. The precise, compositional nature of formal logic enables a dense, incremental reward structure in other- wise prohibitively sparse reward signals. We then take a step back and explore how formal task specifications can be learned from data and provide a novel method for learning these specifications from user-provided preferences. Last, we discuss policy learning in the era of foundation models, and resolve the single-shot nature of existing robot foundation models (RFMs) by devising a structured feedback mechanism that allows RFMs to iteratively learn complex tasks through trial-and-error.As a whole, this dissertation advocates for a principled integration of symbolic structure into end-to-end policy learning techniques to enable complex, long-horizon behaviors. We conclude by discussing how a foundation model-centric field can still benefit from structured task specifications, and the role formal logic can play in an increasingly neural landscape.