Foundations of Neuron-level Interpretability: Automation, Evaluation and Training
Skip to main content
eScholarship
Open Access Publications from the University of California

UC San Diego

UC San Diego Electronic Theses and Dissertations bannerUC San Diego

Foundations of Neuron-level Interpretability: Automation, Evaluation and Training

Abstract

Despite the impressive performance of deep neural networks, our understanding of how they work is very limited. To address this, the field of mechanistic interpretability aims to build a bottom-up understanding of the mechanisms that a neural network uses to calculate its output. An important building block for mechanistic interpretability is interpreting the roles of individual neurons inside the network. However, interpreting individual neurons poses many challenges. First, a single network may have thousands or millions of neurons, so it is not feasible for humans to manually interpret each neuron. Second, the field lacks clear standardized definitions of what makes a good explanation, making it difficult to compare competing methods or measure progress. Finally, individual neurons can be polysemantic, i.e., they activate on multiple unrelated concepts, which makes interpreting them difficult. In this dissertation, I discuss solutions to these three essential problems in neuron-level interpretability. First, in Chapters 1 and 2, I introduce CLIP-Dissect and Linear Explanations, two methods that leverage pre-trained multimodal models to automatically explain all neurons of a vision model. Second, in Chapters 3 and 4, I discuss our two recent papers addressing the lack of standardized evaluation, where we unify diverse evaluations from existing studies under the same mathematical framework and find that most methods don’t pass simple sanity checks, as well as propose ways to more efficiently evaluate these explanations. Finally, in Chapter 5, I discuss a solution to the challenge of polysemanticity in the form of Label-free Concept Bottleneck Models, an automated method for training models to have interpretable individual neurons.