arXiv · 2610.08951
ASPIRE: Agentic Safety & Prompt Injection Red-teaming Engine
Abstract
LLM agents retrieve untrusted content and act through tools, creating indirect prompt-injection risks that can cause unauthorized actions or persistent state changes. Existing automated red-teaming largely optimizes payloads for pre-specified scenarios, leaving latent vulnerabilities across the agent's behavior space unexplored. We present ASPIRE, an Agentic Safety & Prompt Injection Red-teaming Engine for open-ended, behavior-level vulnerability discovery. ASPIRE maintains an evolving Agent Security Behavior Graph and uses complementary Explore and Exploit experts to discover, verify, and generalize consequence-centric tests. Trajectory evidence updates the graph and diagnoses partial or failed attempts, while cross-run memory transfers useful red-team strategies. Experiments on various benchmarks show that ASPIRE substantially expands coverage across consequences, injection methods, environments, and behavior paths while maintaining strong attack success.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pengfei He, Deep Mitra, Vishesh Sharma, Jiliang Tang, Vinay S Rao, Tomas Pfister, Long T. Le. 2026-10-06. ASPIRE: Agentic Safety & Prompt Injection Red-teaming Engine. https://arxiv.org/abs/2610.08951
Cite the original work for its findings. Save a collection to share your selection of sources.