arXiv · 2609.33481
What Does It Mean to Forget a Person? Individual-Level Unlearning in Vision-Language Models
Abstract
Erasing individual identities from Vision-Language Models (VLMs) is uniquely challenging because personal data is entangled across modalities rather than stored as isolated attributes. However, existing multimodal unlearning benchmarks primarily evaluate attribute-centric forgetting, overlooking the more critical objective of individual-level unlearning: eliminating a model's ability to access, link, and reconstruct target-related information across modalities. To address this gap, we propose IDUnlearn-Bench, the first benchmark for individual-level multimodal unlearning in VLMs. It represents each individual as connected multimodal evidence and evaluates four task families: attribute access, identity access, identity binding, and identity reconstruction. Experiments on representative VLMs and unlearning methods show that successful attribute-centric forgetting often leaves substantial identity-level knowledge intact and can be non-monotonic: reducing one form of risk may amplify another. Models may suppress selected information while still identifying the target, linking records, or reconstructing the individual. These findings reveal a fundamental gap between forgetting information about a person and forgetting the person as a whole.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiongtao Sun, Hui Li, Tiantong Wu, Jiaming Zhang, Fuyao Zhang, Wen Jun Tan. 2026-09-27. What Does It Mean to Forget a Person? Individual-Level Unlearning in Vision-Language Models. https://arxiv.org/abs/2609.33481
Cite the original work for its findings. Save a collection to share your selection of sources.