arXiv · 2609.23892
Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs
Abstract
Mechanistic interpretability defines features as the fundamental units of a neural network and circuits as the weighted subgraphs that carry out its computation. Because individual neurons are polysemantic, Cross-Layer Transcoders (CLTs) were introduced as a way to approximate a model's circuits by generating an attribution graph. The nodes of that graph, however, are unlabeled features: reading a graph means pruning it and then working out by hand what each surviving node means. To make CLTs easier to use for circuit discovery, we introduce Circuit-Diff, which intervenes on the model itself with a low-rank factual edit and takes the features whose role in the attribution graph changes under that edit as related to the edited knowledge. On the edits we examine, the flagged nodes are not only detectors of the object token: read off the CLT's released feature dashboards, they include features for the history, geography and associations surrounding the old and new objects. We formalize the method, measure how reliable a frozen CLT remains after a factual edit, test the selected nodes causally by patching them on up to 24 CounterFact edits, give a case study, and release an open-source implementation built on the circuit-tracer package, together with two further tools (multi-prompt aggregation and rule-based supernode labeling).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Edward G. Friedman, Xiangchen Song. 2026-09-20. Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs. https://arxiv.org/abs/2609.23892
Cite the original work for its findings. Save a collection to share your selection of sources.