arXiv · 2606.10904
When the Defense Writes the Refusal: Auditing Keyword-Scored Evaluation of Inference-Time Defenses for Multimodal Large Language Models
Abstract
Prompt-based inference-time defenses for multimodal large language models (MLLMs) emit safety language into the very answers they defend, and a lexical keyword scorer whose refusal list contains that language then counts defended answers as refusals. We audit one such evaluator, the legacy keyword rule of our own experimental archive, over three executed defenses: static defensive prompting, single-pass rationale-conditioned prompting, and a five-sample Gaussian perturb-and-vote image adapter, evaluated across eight InternVL and Qwen-VL models. On the canonical 28,000-output benign grid, 7,532 outputs carry a keyword flag, a separate LLM-judge estimator attributes about 138.1 refusals, and the pooled refusal point estimate is 0.52%: the keyword-flag count was 54.6 times the estimated refusal count, and most keyword-positive answered cases echo the defense's own safety scaffold. The misfires concentrate in the rationale-wrapper arms; this defense-evaluator interaction is therefore arm-dependent, not a constant offset, and invalidates keyword-scored comparison of prompt-based inference-time MLLM defenses. We formalize the audit with three claim-admissibility levels (input-valid, output-bound, construct-validated) under which three of seven planned benchmark branches are excluded. The retained pipeline profiles vary across models and corpora and set priorities for paired semantic re-evaluation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bulat Nutfullin, Vladimir Evgrafov, Dmitry Namiot. 2026-06-09. When the Defense Writes the Refusal: Auditing Keyword-Scored Evaluation of Inference-Time Defenses for Multimodal Large Language Models. https://arxiv.org/abs/2606.10904
Cite the original work for its findings. Save a collection to share your selection of sources.