arXiv · 2609.37638
Targeted Visual Counterfactual Explanations for Contrastive Vision-Language Model
Abstract
Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbf{M}ask-guided \textbf{A}daptive \textbf{C}ounterfactual \textbf{E}xplanations (\mace), a targeted visual counterfactual method designed specifically for CLIP zero-shot classification. \mace constructs an editable region from either source attribution or source--target attribution differences and expands the mask only when needed to reach a specified target class. A latent diffusion inpainting model then modifies the selected region, while a frozen CLIP model provides modification guidance and anchors the remaining image content to the original input. We evaluate \mace on ImageNet, Food-101, Oxford Pets, and CUB-200. The source-mask variant achieves the highest target top-1 success rate across all four datasets, while the difference-mask variant produces the smallest pixel-level and perceptual changes and the best realism scores. Both variants improve proximity and realism over a Stable Diffusion-only baseline using the same generative backbone. These results show that adaptive mask-guided editing produces effective CLIP counterfactuals. They further reveal a tradeoff between counterfactual validity and source-image preservation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Van Bach Nguyen, Jörg Schlötterer, Christin Seifer. 2026-09-29. Targeted Visual Counterfactual Explanations for Contrastive Vision-Language Model. https://arxiv.org/abs/2609.37638
Cite the original work for its findings. Save a collection to share your selection of sources.