arXiv Science⌕ Search

arXiv · 2609.35234

Poster: Towards ProofWeave: A Privacy-Minimised, Integrity-Anchored Evidence Plane for Continuous Agentic Assurance

Abstract

Agentic AI systems increasingly act via tools, memory, delegation, and external services. Existing observability and provenance mechanisms can reconstruct events post hoc, but they rarely show, at the time of the record, whether each policy-relevant action was checked by the intended control before execution. This leaves a trust-observability gap for continuous monitoring, detection, and response: later assurance may rest on evidence that is incomplete, privacy-leaking, mutable, or detached from the policy context that governed the event. What's missing in the literature is contemporaneous, policy-bound evidence that the intended control was evaluated under the policy in force at the time. We introduce ProofWeave, a record-time chain-of-evidence concept for agentic AI assurance. At each policy-relevant action boundary, ProofWeave generates a privacy-minimised and integrity-anchored evidence transaction that binds (i) agent intent or action, (ii) control response, and (iii) a policy-at-time snapshot. Each transaction is committed to an append-only ledger and materialised into a derived proof graph. A bounded Weaver Agent translates policy intent into proof obligations, while deterministic validators check evidence completeness, privacy minimisation, policy binding, and integrity. In the minimal scenario, an agent attempts to transmit a secret to an unapproved external sink. The audit compares a logs-only correlation baseline with ProofWeave across verdict latency, join ambiguity, privacy exposure, tamper detection, and resistance to graph-only proof injection. ProofWeave reduces candidate bindings per verdict from up to `10,201` to one, validation operations from up to `10,201` to approximately `26`, and assurance evidence storage from `0.79`MiB to `0.15`MiB per project.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Guy Lupo, Nguyen Hung Nguyen, Viet Vo, Chamikara M. A. P., Guangdong Bai. 2026-09-28. Poster: Towards ProofWeave: A Privacy-Minimised, Integrity-Anchored Evidence Plane for Continuous Agentic Assurance. https://doi.org/10.1145/3830454.3846451

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

SoK: Cryptocurrency Mixing and Anonymity - Architectures, Threat Models, Operational Aspects and Security

Public blockchains record transaction histories that enable address clustering, taint analysis, and cross-service attribution, thereby motivating the development of mixers and privacy layers. Our work presents a structured scoping review of 22 representative systems, defining a common unlinkability objective and five adversary archetypes. We evaluate these systems against a taxonomy of attack surfaces, including chain analysis, timing inference, custodial compromise, coordination abuse, network metadata, and trusted execution compromise. While nominal anonymity-set size and cryptographic strength characterize privacy in theory, effective anonymity in practice depends on transaction denominations, cover traffic, relayer behavior, and compliance-interface design. Distinguishing nominal from effective anonymity, we derive four core lessons: (1) Privacy is strongest when integrated into everyday transactions, since standalone mixing creates an easily profiled user subset; (2) Trust points, including operators, peer quorums, and hardware enclaves, must be explicit so users know who can break privacy; (3) Network metadata, including gas funding and timing, must be treated formally as protocol data in privacy evaluations; and (4) Compliance should use auditable cryptographic predicates for selective disclosure rather than broad operator discretion. Ultimately, our systematization clarifies the strengths, failures, and future requirements of blockchain privacy architectures.

cs.CR↗

Decoding One Safety Trigger Token for Balancing Safety and Usability in Large Language Models

Large Language Models (LLMs) have been extensively used across diverse domains, including virtual assistants, automated code generation, and scientific research. However, they remain vulnerable to jailbreak attacks, which manipulate the models into generating harmful responses despite safety alignment. Recent studies have shown that current safety-aligned LLMs undergo shallow safety alignment. In this work, we conduct an in-depth investigation into the underlying mechanism of this phenomenon and reveal that it manifests through learned ''safety trigger tokens'' that activate the model's safety patterns when paired with the specific input. Through both analysis and empirical verification, we further demonstrate the high similarity of the safety trigger tokens across different harmful inputs. Accordingly, we propose D-STT, a simple yet effective defense algorithm that identifies and explicitly decodes safety trigger tokens of the given safety-aligned LLM to activate the model's learned safety patterns. In this process, the safety trigger is constrained to a single token, which effectively preserves model usability by introducing minimum intervention in the decoding process. Extensive experiments across diverse jailbreak attacks and benign prompts demonstrate that D-STT significantly reduces output harmfulness while preserving model usability and incurring negligible response time overhead, outperforming ten baseline methods.

cs.CR↗

A traffic analysis attack against Introduction Protocol and Onion Services

Tor onion services rely on long-lived introduction circuits to support anonymous rendezvous between clients and services. Although Tor incorporates defenses against traffic analysis, the introduction protocol retains deterministic routing structure that can be exploited by an adversary. We present a practical intersection attack against Tor introduction circuits that over repeated interactions can identify each hop from the introduction point toward the onion service while requiring observation at only one relay per stage. The attack repeatedly probes the target service and intersects sets of destination IP addresses observed within narrowly bounded INTRODUCE1-RENDEZVOUS2 intervals, without assuming global visibility or access to packet payloads. Our traffic-analysis technique identifies with certainty the next relay in the path to target at each stage, thereby revealing a gap in Tor's privacy model, which is intended to resist traffic-analysis attacks in which an adversary uses traffic patterns to determine which points in the network to observe or attack. We evaluate the attack's feasibility through live-network experiments using a self-operated onion service and relays. To support data minimization, we implement a Tor-compatible plugin that computes intersections online over pseudonymized data retained only in volatile memory. Our experiments show reliable convergence in practice, with convergence rate influenced by relay consensus weight and time-varying background traffic. We further assess practicality under a partial-global adversary model and discuss the implications of geographic concentration in Tor relay selection weight across cooperating jurisdictions.

cs.CR↗