Skip to main content
eScholarship
Open Access Publications from the University of California

UC Berkeley

UC Berkeley Electronic Theses and Dissertations bannerUC Berkeley

AI Personalization in Data-Scarce Markets

Abstract

This dissertation studies the effectiveness of general-purpose artificial intelligence systems, such as GPT and Gemini, for marketing personalization in data-scarce markets. While these systems are trained on large-scale public web data, their performance may be limited in settings where relevant users are underrepresented or where observed preference signals are systematically noisy. This dissertation develops a framework for understanding these limitations and evaluates strategies for improving personalization in such environments. I develop a theoretical framework centered on two primitive dimensions of data scarcity: representativeness and precision. Representativeness gaps arise when relevant user segments are missing non-randomly from training data, while precision gaps arise when observed data provide noisy or distorted signals of underlying preferences. The framework shows how these forms of scarcity interact with the tendency of general-purpose AI systems to rely on dominant patterns in their training data, thereby reducing personalization effectiveness. I evaluate these predictions in a large-scale field experiment involving more than 30,000 retail merchants in India. I find that merchants do adopt AI-generated sales recommendations, but sales gains are concentrated among those facing relatively limited data gaps. For these merchants, AI recommendations increase sales by approximately 10 percent. As data gaps intensify, however, the AI system increasingly defaults to broadly appealing generic content rather than tailored content. In particular, merchants facing representation gaps receive poorly matched recommendations because the system relies on patterns learned from populations that are not aligned with their preferences.Finally, I examine two targeted interventions designed to mitigate these failures: incorporating local preference information from pilot data and enabling merchant agency through menu-based choice. The results show that these interventions address different types of scarcity. Data augmentation is effective in mitigating representation gaps, whereas precision gaps are less responsive to AI-only correction and instead require greater reliance on merchant choice. Overall, the dissertation demonstrates that the effectiveness of AI personalization in data-scarce markets depends not only on model capability, but also on the specific form of data scarcity and the design of complementary human and data interventions.

Main Content

This item is under embargo until August 31, 2028.