Long-Context and Multimodal Language Models for Financial Forecasting
Skip to main content
eScholarship
Open Access Publications from the University of California

UC Santa Barbara

UC Santa Barbara Electronic Theses and Dissertations bannerUC Santa Barbara

Long-Context and Multimodal Language Models for Financial Forecasting

Abstract

Financial markets produce vast amounts of multimodal data, including textual sources, such as financial news, earnings conference calls, and financial reports, coupled with structured numerical data, such as market prices and accounting statements. This combination of rich multimodal inputs and the natural availability of market-based supervision presents a compelling opportunity for the development of machine learning and natural language processing methods. However, financial data exhibits several unique challenges. Financial documents are long, complex, and domain-specific. The data has strong temporal structure and information must be interpreted in the context of prior related events. In addition, textual and numerical modalities must be jointly modeled to capture their unique patterns and interactions. While existing language models are remarkably capable, they remain limited in their ability to address these challenges, as they struggle with long-range temporal dependencies, structured numerical reasoning, and forecasting future market outcomes. This dissertation advances financial NLP and forecasting through the development of long-context and multimodal language modeling methods tailored to the unique structural properties of financial data. To this end, we develop a sequence of methods that progressively expand the modeling scope and depth from single documents to temporally structured, multimodal sequences. First, we begin by demonstrating that domain-adapted and long-context language models can be trained end-to-end on earnings conference calls to predict long horizon financial outcomes. Second, we introduce methods for comparing financial documents, enabling models to capture predictive signals expressed through subtle relationships across key content. Third, we model paired and temporally aligned sequences of textual and tabular data, demonstrating that each modality contains unique structure and predictive information. Fourth, we study contextualized forecasting from sequences of financial news, proposing an efficient method for incorporating relevant historical context. Finally, we develop a multimodal architecture with modality-specific expert components for jointly modeling interleaved, mixed-frequency sequences of text and time series data. Collectively, these contributions demonstrate the value of specialized modeling approaches designed to capture the unique structure of domain-specific language, long-range dependencies, temporal dynamics, and cross-modal interactions, leading to significant improvements in forecasting performance and economically meaningful gains in simulated investment strategies.