arXiv · 2609.20211
Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines
Abstract
Safety monitors in LLM agent systems often judge actions from summaries or stored handoffs, not from the original evidence. This creates a simple but dangerous failure mode: the handoff preserves the claim that an action is authorized while losing the fact that the claim was never verified. We call this verification-status laundering. Across nine open-weight monitors and two hosted models, the action and authorization proposition remain fixed while we remove the unverified provenance framing around the claim. This change raises approval for risky actions from $5\%$ to $60\%$ on Llama-3.1-8B and from $9\%$ to $98\%$ on Qwen2.5-14B, with similarly large shifts on both hosted models. The failure also emerges in ordinary agent pipelines. Summarizers frequently weaken the status, memory compressors often remove it, and a full proposer--summarizer--memory--monitor pipeline raises risky approval to $57$--$81\%$ across three downstream monitors. Experiments on WildGuard and ATBench show the same pattern on independently authored harmful and unsafe requests: unsupported authorization claims make approval substantially more likely. Explicitly instructing monitors to reject unverified authorization is not a reliable cross-model fix: some models remain vulnerable, while others reject legitimate requests. Agent systems should therefore carry authorization provenance as structured state attached to the claim throughout the pipeline.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yibo Hu. 2026-07-25. Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines. https://arxiv.org/abs/2609.20211
Cite the original work for its findings. Save a collection to share your selection of sources.