arXiv · 2609.21142
Exploring Text Classification Models with Sparse Autoencoders
Abstract
As language models (LMs) rise in prominence, there is interest in making them more transparent in order to better understand their internal behavior. Recent interpretability work has focused on using sparse autoencoders (SAEs) to break down neuron activations at a given layer in the LM into human-understandable features, where each feature represents a concept that the model has learned. In this paper, we share work on using SAEs to analyze the behavior of text classification LMs. We present techniques for exploring the relationships between the SAE's features and the model's predictions and errors. We integrate these techniques into SAEfarer, a tool for analyzing concepts learned by text classification LMs. We assess SAEfarer in an expert pilot evaluation with five Ph.D. students.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daniel Kerrigan, Brian Barr, Enrico Bertini. 2026-09-17. Exploring Text Classification Models with Sparse Autoencoders. https://arxiv.org/abs/2609.21142
Cite the original work for its findings. Save a collection to share your selection of sources.