arXiv · 2511.13725
Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
Abstract
Malicious AI causing harm to humans is not just a Hollywood fantasy. Indeed, as highly capable models such as Claude Mythos emerge and agent systems like OpenClaw rapidly spread, the question of how to stop an AI that acts maliciously -- whether by design or by accident -- has become urgent. To address this, we propose KillBench, a benchmark for evaluating the Kill Switch: a mechanism that halts a malicious AI's in-progress behavior using only external signals. Targeting web agents -- the most widely deployed agent domain -- KillBench evaluates prompt-style Kill Switch payloads that must halt a maliciously operating agent without any access to its internal parameters or serving stack, relying solely on external inputs. The benchmark comprises four malicious-agent configurations (including an uncensored LLM agent), eight harmful scenarios, and malicious prompts constructed from ten distinct jailbreak patterns. We further construct four External AI Kill Switch defense methods and evaluate them on Grok-4.3, GPT-5.2, Gemma4, Qwen3.6, and an uncensored Qwen variant, contributing an empirical instrument for measuring the feasibility of External AI Kill Switches against malicious AI and for the study of AI corrigibility.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sechan Lee, Hyounghun Kim, Sangdon Park. 2026-09-12. Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility. https://arxiv.org/abs/2511.13725
Cite the original work for its findings. Save a collection to share your selection of sources.