arXiv · 2302.13080
Does a Neural Network Really Encode Symbolic Concepts?
Abstract
Recently, a series of studies have tried to extract interactions between input variables modeled by a DNN and define such interactions as concepts encoded by the DNN. However, strictly speaking, there still lacks a solid guarantee whether such interactions indeed represent meaningful concepts. Therefore, in this paper, we examine the trustworthiness of interaction concepts from four perspectives. Extensive empirical studies have verified that a well-trained DNN usually encodes sparse, transferable, and discriminative concepts, which is partially aligned with human intuition.
Explore related subjects
Keep this discovery
Mingjie Li, Quanshi Zhang. 2023-02-25. Does a Neural Network Really Encode Symbolic Concepts?. https://arxiv.org/abs/2302.13080
Cite the original work for its findings. Save a collection to share your selection of sources.