arXiv · 2609.32355
Gradients for Interventions and Activations for Detection: Targeted Feature Learning in Language Models
Abstract
Model-internal features can be studied through both their ability to identify a specified concept and their causal effect when manipulated, e.g., through steering or weight editing. A prominent approach to feature learning is Sparse Autoencoders (SAEs), which learn broad feature dictionaries whose relation to particular concepts is typically identified post hoc. However, many interpretability questions are instead hypothesis-driven and concern a concept specified in advance. We study this setting as targeted feature learning, where a single feature is constructed for such a predefined concept. We present a controlled comparison across three model signals (activation values, activation gradients, and parameter gradients) and two estimators (contrastive mean and a learned one-dimensional encoder-decoder), yielding six targeted methods, with CAA and GRADIEND as existing instances and four new methods covering the remaining combinations. We compare these methods against pretrained SAEs across 15 tasks and three language models, evaluating both detection and causal intervention. Across models, the strongest detection performance is achieved by contrastive activation value methods, whereas the strongest intervention performance is achieved by gradient-based methods. Overall, our results show that targeted feature quality depends jointly on the model signal and estimator, with detection and intervention capturing complementary properties.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jonathan Drechsel, Steffen Herbold. 2026-09-26. Gradients for Interventions and Activations for Detection: Targeted Feature Learning in Language Models. https://arxiv.org/abs/2609.32355
Cite the original work for its findings. Save a collection to share your selection of sources.