arXiv Science⌕ Search

arXiv subjects

Qishuai Jing

Publications and source records attributed to Qishuai Jing.

2 recordsLinked to original sources

What a Policy Gate Can and Cannot Know: Measured Boundaries of Cross-Platform Command Adjudication

Gateways that adjudicate an agent's actions before they execute are only as good as their understanding of the action. We study a policy gate that never parses shell syntax: it consumes a typed, realised action (verb, operands, resolved zones, program-object identity) and decides ALLOW, ASK or DENY. Working on a Linux twin of the Windows benchmark of our previous study [1], we ask how faithfully it adjudicates, what survives translation, and whether deciding stays affordable as the system is used. A frozen 61-case table scores 61/61 in two rounds with no false allow; a 50-operator mutation campaign kills 46 of 50 mutants (92.0%), with all four surviving mutants classified. An 82-row audit yields 43 re-expressions, 21 carrier differences, 16 study-specific inapplicable rows and two unresolved cases; of 25 rows labelled "no counterpart," four remain unmapped here. Consulted by the executor, 25 escapes become zero with no benign payload blocked. In an exploratory one-gateway snapshot, three unauthenticated endpoint labels produce case-level non-refusal majorities of 81.8-98.0%, but immediate-execution majorities of 4.0-52.5%. Adjudication reads no accumulating state; credential verification does, scanning its whole ledger. We report that cost and the fix we would make.

cs.CR↗

Separation of Duties for Privileged LLM Agents: A Governed Execution Architecture with Measured Security-Utility Trade-offs

Large language model agents are increasingly granted real privileges (executing commands, modifying files, calling APIs), so an agent that errs has already acted. Existing defences concentrate on the agent's inputs, while the path from a candidate action to privileged side effects remains less directly studied. We argue that this path must be governed outside the model, and study an architecture interposing four roles (planner, policy gate, executor, auditor) between agent and operating system. Two choices are central: actions arrive as structured intents, so adjudication never parses shell syntax; and approval is a one-shot credential bound to the exact bytes that will run. We evaluate on a 313-case benchmark across eight variants, with prompt-only baselines from three hosted LLMs on 150 stratified cases, five repetitions (2,250 attempted calls; 2,249 completed). Effective attack success falls from 98.3% under direct execution to 7.7% deployed. Re-execution against the real implementation yields a similar aggregate rate (7.6% over 66 sandbox-evaluable payloads) but substantial case-level disagreement, at a corrected false-denial rate of 11.1%. A substantial final-stage reduction (from 30.8% to 7.7%) is attributable to the operating-system sandbox, and the benchmark found four implementation defects, none by design review.

cs.CR↗