arXiv Science⌕ Search

arXiv · 2610.03434

Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems

Abstract

Large language model-based agents are increasingly deployed to perform domain-specific tasks by interacting with enterprise knowledge, tools, and external services. Existing runtime guardrails primarily target prompt injection and other attack-specific behaviors under a black-box threat model, but provide limited guarantees that agents operate within their intended functionality. As a result, production agents remain vulnerable to malicious requests and out-of-domain queries that existing defenses often fail to distinguish. We present Persona Guardrail, a production-grade runtime defense framework that enforces explicit functional boundaries for customer-facing agentic AI systems through synchronous input and output validation driven by semantic allowlist and blocklist specifications. We also introduce PAGE (Persona-Aware Guardrail Evaluation), a benchmark for evaluating function-specific guardrails across benign, adversarial, and out-of-domain interactions on both user and agent turns. Compared with a generic LLM-based guardrail, Persona Guardrail improves overall accuracy from 85.7% to 95.9%, increases out-of-domain detection from 57.3% to 93.5%, and reduces the false-approved rate from 25.0% to 4.7%. Currently deployed in production, Persona Guardrail meets its latency budget while sustaining a very low false-block and false-allow rate under realistic production workloads. These results demonstrate that Persona Guardrail provides a practical, scalable, and production-ready foundation for securing agentic AI systems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bijeeta Pal, Sridhar Reddy Maddireddy, Muhaimin Bin Munir, Zoltan Puha, Max Zhurovich, Adi Raghavendra, Sean Tout. 2026-10-02. Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems. https://arxiv.org/abs/2610.03434

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Extended Differential Cryptanalysis of Kuznyechik

We study the input-side relation $F(cx\oplus a)\oplus F(x)=b$ as an extended differential for block-cipher analysis. Because multiplication occurs before the first S-box, fixing that S-box output leaves an ordinary XOR difference for subsequent binary linear layers and common key additions. For permutations, the extended and outer $c$-differential tables are related by inversion. For the Kuznyechik S-box, the exceptional inverse classes $\mathtt{02}/\mathtt{e1}$, $\mathtt{04}/\mathtt{91}$, and $\mathtt{03}/\mathtt{be}$ have exact extended differential uniformities $64$, $33$, and $21$. We connect them to the pseudo-exponential structure of Perrin and Udovenko through a cross-field transfer theorem. In the common polynomial basis, the disagreement rank between multiplication maps in the hidden and Kuznyechik fields equals the degree of the multiplier polynomial, yielding structural lower bounds $60$, $30$, and $16$. We also analyze the complete extended differential table: its normalization is doubly stochastic, its energy is an exact multiplicative correlation of ordinary DDT rows, and transferred hidden fibers give explicit Pearson and Shannon-information bounds. Whitening selects the effective row $a\oplus(1\oplus c)K$. For two-round standard Kuznyechik we derive the exact fixed-key channel $Q_{c,\ell}=W_cP_\ell W_1$ and a two-stage maximum-likelihood attack. With $c=\mathtt{04}$ it recovers the full $256$-bit master key with failure probability at most $0.01$, using $226{,}512$ chosen-plaintext pairs and at most $240{,}669$ encryption queries. Restricted product-model searches give gains of $5.2$, $4.6$, and about $1.6$ bits over the corresponding classical families at two, three, and four rounds. A separate nine-round no-pre-whitening Monte Carlo study is reported only as exploratory evidence because its final configurations followed preliminary screening.

cs.CR↗

Mitigating Watermark Forgery in Generative Models via Randomized Key Selection

Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, whose presence can be detected using a secret watermark key. A core security threat are forgery attacks, where adversaries insert the provider's watermark into content \emph{not} produced by the provider, potentially damaging their reputation and undermining trust. Existing defenses resist forgery by embedding many watermarks with multiple keys into the same content, which can degrade model utility. However, forgery remains a threat when attackers can collect sufficiently many watermarked samples. We propose a defense with a sample-count-independent upper bound on forgery success for blind attackers, conditional on key-symmetric, independent detector outcomes. Our scheme does not further degrade model utility. We randomize the watermark key selection for each query and accept content as genuine only if a watermark is detected by \emph{exactly} one key. Unlike cryptographic watermarks that rely on computational hardness assumptions and require designing new watermarking schemes from scratch, our method can be applied to any existing watermarking method to improve its forgery resistance. We focus on text watermarking, but our defense is modality-agnostic, since it treats the underlying watermarking method as a black-box. To show this, we include a preliminary study on image watermarking using Tree-Ring. Separately from this conditional guarantee, we empirically observe that, at $r=4$ keys, harmful-text forgery success drops from as high as $87\%$ with a single key to as low as $1\%$ against the adaptive blind attackers that we evaluate, at negligible computational overhead; a preliminary image study shows a reduction from $100\%$ to $2\%$.

cs.CR↗

Understanding Gaps in LLM Pipelines Towards Scalable Fuzzing Harness Generation: An Empirical Study and Enhancement

Large language model (LLM)-based techniques have achieved notable progress in fuzz harness generation. However, applying them to arbitrary functions \textit{at scale} remains difficult---generated harnesses often fail to compile or, worse, compile but remain logically ineffective. What factors drive success and what limitations hinder current methods remain unclear. To answer these questions, we conduct an empirical study on state-of-the-art LLM-based harness generation frameworks across 29 OSS-Fuzz projects. Our study establishes two pillars of success: contextual information to guide generation and pre-structured pipelines that ensure workflow stability. However, significant gaps remain: (1) existing context retrieval methods lack the robustness to reliably obtain necessary information across diverse projects; (2) current validation mechanisms fail to detect logically ineffective harnesses, such as those containing fabricated definitions; and (3) compilation workflows are brittle, unable to distinguish harness-level errors from build-configuration issues. We demonstrate the utility of these findings by enhancing existing techniques with a hybrid tool pool for robust context retrieval, an enhanced validation pipeline, and a compilation-error triage strategy. Evaluated on 243 OSS-Fuzz projects (65 C and 178 C++), the enhanced approach improves the three-shot success rate by approximately 20\% over state-of-the-art techniques, reaching 87\% for C and 81\% for C++. Our one-hour fuzzing results show that more than 75\% of the generated harnesses increase target-function coverage, surpassing baselines by over 10\%. In addition, the enhanced approach identified 15 new vulnerabilities across 11 real-world projects.

cs.CR↗