arXiv · 2609.13529
Generative Interpretability via Scalable Neuro-Symbolic Models
Abstract
As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward \emph{generative interpretability}, an architectural property under which a model's inference pass natively exposes semantically meaningful checkpoints that are human-understandable and amenable to causal intervention. We show the merits of generative interpretability as comparison to other interpretability research paradigms, and propose Neuro-Symbolic Models as a concrete instantiation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaocong Yang. 2026-09-11. Generative Interpretability via Scalable Neuro-Symbolic Models. https://doi.org/10.1145/3806096.3844850
Cite the original work for its findings. Save a collection to share your selection of sources.