- Main
Structured Models for Sample-Efficient Learning and Control in Energy Systems
- Bian, Yuexin
- Advisor(s): Shi, Yuanyuan
Abstract
Modern energy systems are increasingly shaped by flexible demand, distributed storage, and building infrastructure. These systems generate rich data and require fast, adaptive decisions, but they are also governed by physical laws, safety constraints, and strategic behavior that are difficult to capture with generic black-box learning. This creates a disconnect between the structure-rich models used in classical energy systems analysis and the flexible learning pipelines favored in modern data-driven control. Methods that ignore known structure often suffer from poor sample efficiency, weak interpretability, and unreliable constraint handling, while purely model-based approaches struggle with hidden objectives and computational scale. The goal of this thesis is to bridge this disconnect by developing structured learning and control methods that integrate optimization, physics, and control-theoretic structure directly into trainable decision-making architectures.First, the dissertation develops differentiable optimization models for forecasting strategic demand response behavior from historical observations, allowing latent agent objectives and constraints to be learned from price-responsive actions. Second, it develops a PDE-constrained framework for building ventilation and temperature control based on the Navier--Stokes and convection--diffusion equations, enabling energy-efficient control while explicitly enforcing air-quality and thermal constraints under the modeled dynamics. Third, it addresses the computational burden of PDE-based control by introducing an ensemble neural operator surrogate trained on CFD simulations of a real classroom, yielding over five orders-of-magnitude speedup while preserving the spatial structure needed for closed-loop optimization. Fourth, it extends structured modeling to reinforcement learning through DiffOP, in which the policy is defined by a parameterized optimal control problem and trained end-to-end from reward feedback. Finally, it shows that structure can also enter through policy representation, where discretized categorical actors with regularized networks improve stability and performance in on-policy continuous control.Taken together, these results show that structured learning offers a systematic way to combine model-based rigor with learning-based adaptability, yielding methods that are more sample-efficient, interpretable, and reliable for modern energy and control systems.