American Option Pricing in Continuous Time via Reinforcement Learning
Skip to main content
eScholarship
Open Access Publications from the University of California

UC Santa Barbara

UC Santa Barbara Electronic Theses and Dissertations bannerUC Santa Barbara

American Option Pricing in Continuous Time via Reinforcement Learning

Abstract

American options are ubiquitous in financial markets, and many numerical methods price them by approximating their continuous-time exercise feature with a discrete sequence of stopping opportunities. Under classical dynamic programming, a coarse grid with only a few stopping opportunities can materially undervalue the option, whereas on a very fine grid, approximation errors accumulate through the backward recursion and may ultimately break down the computation. We develop a least squares Monte Carlo (LSMC)-inspired algorithm that prices high-dimensional American options in continuous time, in settings where alternative approaches are impractical. Building on the LSMC framework, we initialize an aggregate deep neural network from regression-based timing-value data and then improve it through reinforcement-learning iterations. Starting from a coarse grid, we progressively increase the density of stopping opportunities to limit concept drift while updating the timing-value estimates, thereby enabling stopping decisions at an arbitrarily fine time resolution. We concentrate the training effort near the stopping boundary, where small changes in timing values can flip the stop/continue decisions. This allows us to price previously intractable high-dimensional American options while remaining computationally efficient.