arXiv · 1909.11275
Switched linear projections for neural network interpretability
Abstract
We introduce switched linear projections for expressing the activity of a neuron in a deep neural network in terms of a single linear projection in the input space. The method works by isolating the active subnetwork, a series of linear transformations, that determine the entire computation of the network for a given input instance. With these projections we can decompose activity in any hidden layer into patterns detected in a given input instance. We also propose that in ReLU networks it is instructive and meaningful to examine patterns that deactivate the neurons in a hidden layer, something that is implicitly ignored by the existing interpretability methods tracking solely the active aspect of the network's computation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lech Szymanski, Brendan McCane, Craig Atkinson. 2020-02-06. Switched linear projections for neural network interpretability. https://arxiv.org/abs/1909.11275
Cite the original work for its findings. Save a collection to share your selection of sources.