arXiv Science⌕ Search

arXiv · 2609.28205

GUIAuditor: Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices

Abstract

The proliferation of smart devices exposes children to online risks like grooming and financial scams that are deeply embedded within legitimate applications. Current approaches rely on automated prevention and detection, a paradigm that is fundamentally limited by its inherent fallibility. Whether rule-based or AI-driven, they inevitably produce false positives and negatives, failing to provide reliable protection. In this paper, we argue for a complementary, human-in-the-loop, post-hoc forensic paradigm. We present GUIAuditor, the first system designed to realize this vision by creating GUI Provenance: a queryable, semantic record of a child's interaction sequence. To generate this, GUIAuditor leverages a Multimodal Large Language Model (MLLM) to translate the temporal sequence of GUI events into a human-understandable narrative. To make this practical on mobile devices, a novel evidence distillation pipeline reduces the data requiring analysis by over 89.2% compared to periodic sampling approaches adopted by industry standards, with negligible impact on accuracy. On a new dataset of 295 interaction clips, GUIAuditor achieves a 95.23% Macro-F1 Score in logging significant events and, crucially, its two-stage forensic query engine successfully retrieves the correct evidence as the top result for over 90.20% of natural language questions. An end-to-end evaluation on three modern smartphones shows that the full pipeline, including on-device MLLM inference, adds 2.1W of power draw and 7.4s of per-event latency, with a peak memory footprint of ${\sim}$3.1GB. These results show that post-hoc GUI forensics can run on modern mobile devices and provide useful context for guardian-led safety review.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Junlin Liu, Yifeng Cai, Shuai Wang, Zhineng Zhong, Shaofei Li, Jiacheng Liu, Yuanchun Li, Ziqi Zhang, Xiangqun Chen, Ding Li, Yao Guo. 2026-09-23. GUIAuditor: Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices. https://doi.org/10.1145/3831978

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The shifted-prime Erdős-Wintner law for primitive-root determinant densities: extremal order, dimension zero, and Fourier decay

For a prime $p$, let $c(p)=\frac{φ(p-1)}{p-1}\prod_{j\ge1}(1-p^{-j})$, the limiting density of matrices over $\mathbb F_p$ with primitive-root determinant. Its limiting law over the primes is the classical continuous shifted-totient law on $[0,1/2]$. We prove Hausdorff dimension zero and vanishing lower and upper dyadic $L^q$ dimensions for $q>1$. Its image $μ_f$ under $x\mapsto-\log x$ is Rajchman. As $T\to\infty$, for $U_T$ uniform on $[0,T]$, $\log|\widehat{μ_f}(U_T)|/\log\log T\to-1$ in probability. For every $A>0$, $|\widehat{μ_f}(τ)|\le(\log\log T)^4/\log T$ outside a subset of $[0,T]$ of relative measure $O_A((\log T)^{-A})$. As $h\downarrow0$, $\sup_aμ_f([a,a+h])=\mathfrak S_2e^{-γ}/\log(1/h)+O(\log^{-2}(1/h))$, where $\mathfrak S_2$ is the twin-prime singular series; maximizing left endpoints lie within $h$ of $\log3$ for small $h$. We prove $\min_{p\le x}c(p)\sim e^{-γ}/\log\log x$ and $\limsup_{p\to\infty}(c(p)\log\log p)^{-1}=e^γ$. The limiting law of $\log(φ(p+1)/φ(p-1))$ has support $\mathbb R$ and Hausdorff dimension zero. For the classical law of $σ(p-1)/(p-1)$ on $[3/2,\infty)$, we prove dimension zero, a sharp left-endpoint asymptotic, and a Rajchman logarithmic image. Its odd-prime component has an entire Mellin transform of order one. Partial-factorization bounds yield certified asymptotic searches for fully splitting negacyclic number-theoretic transform primes with prescribed reciprocal-density bounds at fixed power-of-two length. We determine the second distinct squared norm of $A_{n_1}\otimes\cdots\otimes A_{n_k}$ for $k,n_i\ge2$, yielding exact cyclotomic codifferent shell gaps and a uniform smoothing asymptotic at $ε=2^{-cφ(m)}$ for $c>2\log_2(1+\sqrt6)$. These results are unconditional. An explicit unproved exponent-pair hypothesis yields $|\widehat{μ_f}(τ)|=O(1/\log\log|τ|)$.

cs.CR↗

Studying Detection Rule Generation as a Unified Task

Security systems use detection rules to identify suspicious activity. Existing studies often investigate rule generation for specific security systems, devoting substantial effort to developing dedicated methods and evaluation setups. Such customization contributes to fragmented research, limiting method reuse and result comparability across systems. We therefore study detection rule generation as a unified task across diverse natural language inputs and rule languages. To support method reuse, we propose UniRule, which abstracts diverse rules into shared natural language representations for retrieval. To enable consistent evaluation, we introduce a protocol that compares rules under shared criteria and aggregates the results into method scores. Experiments demonstrate the effectiveness of UniRule and the reliability of the evaluation protocol. They also show that method performance in one setting can be predicted from results in others, with average error close to that obtained using that setting's own data. These findings support studying detection rule generation as a unified task.

cs.CR↗

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation. This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the attacker's automated judge. Our analysis shows that conventional detect-and-block defenses can allow attacker success rate (ASR) to approach one as the query budget grows, since predictable refusals provide useful feedback to automated search. We then examine detect-and-misdirect, where detected malicious interactions receive controlled, non-operational responses designed to induce false-positive errors in the attacker's judge. This strategy reduces the positive predictive value of attacker-selected candidates and yields a bounded asymptotic ASR. We evaluate a proof-of-concept realization of this strategy through Contextual Misdirection via Progressive Engagement (CMPE), a lightweight conversational misdirection method designed to replace predictable refusal text with safe but strategically misleading responses in automated jailbreak settings. On jailbreak benchmarks, CMPE reduces estimated ASR upper bounds by up to two orders of magnitude and nearly eliminates verified attack success in end-to-end experiments with PAIR, GPTFuzz, and AutoDAN-Turbo.

cs.CR↗