arXiv · 2610.00797
Sapien: A Stateful Policy Engine for Autonomous AI Agents
Abstract
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Corinn Tiffany, Wen Zhang, Eugene Bagdasarian, Lillian Tsai. 2026-09-30. Sapien: A Stateful Policy Engine for Autonomous AI Agents. https://arxiv.org/abs/2610.00797
Cite the original work for its findings. Save a collection to share your selection of sources.