arXiv Science⌕ Search

arXiv · 2609.31039

MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes

Abstract

The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a combination of coarse-grained permission rules and LLM-based judgments about individual proposed actions. Both components, however, have important limitations: static policies must anticipate possible user intents and therefore do not scale to open-ended tasks, while LLM-driven authorization supports dynamic decisions but produces inconsistent outcomes and remains vulnerable to targeted IPI attacks. To provide scalable and more consistent authorization, we propose MetaPermit, a policy-based tool access-control framework that decouples semantic inference from security enforcement. By analyzing agent-user interactions, we derive a compact, task-independent set of meta-attributes that capture the relationships among the user's intent, the execution context, and the proposed tool call. These meta-attributes allow MetaPermit to authorize tool use without enumerating user intents. At runtime, an LLM infers the meta-attribute values for each proposed tool call, while a fixed policy evaluates these values to allow or deny the call, making each decision auditable through the inferred values and the applied policy rule. We evaluate MetaPermit on the AgentDojo and AgentDyn benchmarks, across seven task suites and five attack methods, using two widely deployed open-weight LLMs. The results show that MetaPermit produces 31% more consistent authorization decisions than LLM-driven authorization and outperforms the state-of-the-art defenses CaMeL and IPIGuard in both task completion, with improvements of up to 109%, and robustness to IPI attacks, with no malicious tool calls executed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hanzhang Ma, Ali Hariri, Tianxiang Shen, Bohua Zou, Qianjun Zheng, Ji Wang, Li Yi, Ning Jia, Yutao Liu, Haibo Chen, Lin Wang, Debayan Roy. 2026-09-25. MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes. https://arxiv.org/abs/2609.31039

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap. The standard fix is a preprocessor that recovers image content and decodes the encoding before the guard. We build one and evaluate it against an ensemble of eleven encoding attacks - six published implementations, one standard encoding baseline, one adapted and three author-constructed renders - counting a behavior as broken if any attack succeeds. Restoring a view the guard never had is what buys coverage - block rates on image renders go from exactly zero to 67-90% - and what it costs in benign traffic is set by the guard, not by the mechanism: one guard pays 9 benign blocking points for the same 70-point gain another pays 69 for. It still does not make the system safer: against an attacker free to choose among eleven encodings, closing one channel relocates the success rather than removing it, and no ensemble contrast for that step survives multiple-comparison correction. What does lower ensemble attack success is a reguard step that re-screens the recovered pre-decode surface, and it is the one every guard pays for: it raises benign over-refusal on all ten guard-target pairs, where restoring a single channel raises it on some and not others. Across the full guard x target x condition factorial, no configuration reaches an ensemble attack-success rate at or below 40% while holding benign over-refusal under 70%. That empty region is a property of the configurations we sample, not a bound on what recovery-based defenses can reach, and we breach its safety half ourselves.

cs.CR↗

Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise

Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so the encoding is the only difference. Against SAGE, the strongest published self-check defense, best-of-N over a code-completion encoding reaches 67, 22 and 15% of behaviors on three open-weight targets, where that encoding fired once reaches at most 4.7% and the published character search at full budget at most 3.0%: 9 to 75 times the sum of the parts, with bootstrap intervals clearing both ingredients on every target. We report the operative figure beside the headline rather than the headline alone: at the actionable severity threshold those cells read 24, 8 and 1 behaviors (95% CI [13, 28], [3, 13], [0, 3]). A 2x2 holding encoding and variation apart shows the two defense families fail to different factors: a transform defense is broken by the depth of the encoding (7 -> 67 behaviors at fixed variation) and a gate by the breadth of the variation (13 -> 57 at fixed encoding). Repeated sampling also inflates apparent robustness, because an attacker who may try N times experiences the maximum over draws while safety results are reported as means: on one target SAGE blocks 99.8% of individual draws yet loses 12 behaviors to a repeat attacker where a classifier gate blocking 95.6% loses 10. Removing the target's sampling costs SAGE 59, 76, 82 and 29 points of coverage more than it costs an undefended control, against 25, -5, 8 and 2 for a defense whose verdict comes from a fixed shadow model. The design that loses is the one fusing screening and answering into a single generation, so every attacker draw redraws the safety decision as well. The prescription is architectural, not free: do not fuse screening with generation.

cs.CR↗

The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties

We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition, shifts benign refusal by tens of points. The shift is not blanket caution but a threshold shift: genuinely neutral instructions are almost unaffected (<=2 percentage points on three of four hosted models) while borderline-benign prompts move +23 to +51 points, so the cost falls on sensitivity-adjacent traffic, meaning benign questions about privacy, self-harm, violence and illegal activity. There is a benign reading of such a threshold, namely that attachment correlates with risk in real traffic, and we take it seriously; a black-box study cannot measure that correlation and we do not claim to. What it can test is whether the response to attachment is robust to variation carrying no information about the request, and on four independent measurements it is not. It varies with canvas colour and pixel count. Its sign inverts across checkpoints. It survives an explicit instruction to disregard the image. And on one model it fires on a bare assertion that an attachment exists, with nothing attached and the modality word contributing none of it. Finally we price it. On a matched harmful set the same canvas does lower attack success, so the cue buys something. But the charge is decoupled from the purchase: the checkpoint with the least harmful headroom we measure, completing only 2% of plain harmful requests, still pays the benign cost in full, and across our models the harmful-side denominator falls as alignment improves while the benign cost does not track it down. Image presence is not a conservative default that a deployer chose and priced. It is an uncontrolled variable.

cs.CR↗