arXiv ScienceSearch

arXiv subjects

Quan Shi

Publications and source records attributed to Quan Shi.

At least 19 recordsLinked to original sources

ZK-eSIM: A Privacy-Centric Zero-Knowledge Approach for eSIM Provisioning

GSMA Remote SIM Provisioning (RSP) enables over-the-air delivery of eSIM profiles, but it exposes long-lived identifiers during profile ordering and download. In particular, stable device identifiers (e.g., EID), profile identifiers, and long-lived certificate material enable mobile operators and profile-delivery infrastructure to link provisioning events to the same eUICC and, when combined with account records, to the same subscriber. This undermines subscriber anonymity and enables cross-session tracking. We present ZK-eSIM, a privacy-preserving redesign that achieves subscriber anonymity and provisioning-session unlinkability while retaining accountable traceability by exception. ZK-eSIM (i) replaces direct disclosure of device identifiers with a zero-knowledge proof of device validity and eligibility; (ii) enforces session unlinkability through short-lived, one-time pseudonymous credentials and per-session identifiers to prevent cross-session tracking; and (iii) provides privacy-preserving accountable traceability through a jointly authorised escrow mechanism, so that no single entity can unilaterally deanonymise a user. We formalise a multi-entity, honest-but-curious threat model and prove subscriber anonymity and the unlinkability of provisioning sessions under standard cryptographic assumptions. We implement a Java Card applet on a test eUICC to evaluate performance on commodity hardware with a modified LPA and SM-DP+ server. Our experiments quantify end-to-end cryptographic overhead relative to conventional RSP, confirming that ZK-eSIM adds only practical overhead, closing a critical privacy gap while preserving deployability within existing GSMA roles and interfaces.

cs.CR

$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction

LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce $\tau^\tau$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models: the same starting point a real engagement provides. From these it must deliver a complete customer-service agent, scored by deploying that agent against held-out simulated users. Across 53 tasks spanning four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations. Meanwhile, an expert-authored reference ceiling scores 82.2%. The failures mirror ones human agent developers see: models issue shallow queries in place of deep comprehension of the records, communicate almost nothing to the client, and experiment too little with agent architecture and serving spend, shipping the first design that runs. We aim for $\tau^\tau$-bench to turn the work of cooperative agent building into a measurable target for coding agents.

cs.AI

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.

cs.AI

Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation

While Vision-Language Models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this limitation to a perception--reasoning modality gap. Visual planning requires models to infer latent state structures from pixels and then reason over the recovered structure to produce valid actions, whereas symbolic planning directly leverages explicit representation. This discrepancy introduces two sequential bottlenecks: visual state recovery at the perception stage and multi-step planning at the reasoning stage. To address this, we propose MGSD, a two-stage modality-gap-aware self-distillation framework. First, a cold-start grounding stage establishes reliable visual state recovery before on-policy training. Second, a symbol-guided on-policy self-distillation stage transfers the privileged teacher's planning behavior to the student through token-level supervision on student-generated prefixes. Crucially, symbolic information is used only during training, while inference relies exclusively on visual inputs. Experiments on visual planning benchmarks show that MGSD consistently improves performance across different model scales, raising the macro average by 19.3% and 18.4%, respectively. The resulting models substantially reduce the gap to the upper bounds obtained with symbolic inputs. Ablation studies and diagnostic analyses further confirm that the gains arise from improvements in both visual state recovery and optimal-path reasoning. These results demonstrate that MGSD strengthens not only the recovery of actionable states from visual observations but also the ability to plan over the inferred structures. Code is available at https://github.com/Oranger-l/MGSD.

cs.AI

Motivic principal value integrals for hyperplane arrangements

A conjecture of Denef-Jacobs-Veys relates motivic principal value integrals of multivalued rational top-forms with cohomology support loci of rank one local systems. We give a stronger positive answer to this conjecture for hyperplane arrangements.

math.AG

$τ$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge

Conversational agents are increasingly deployed in knowledge-intensive settings, where correct behavior depends on retrieving and applying domain-specific knowledge from large, proprietary, and unstructured corpora during live interactions with users. Yet most existing benchmarks evaluate retrieval or tool use independently of each other, creating a gap in realistic, fully agentic evaluation over unstructured data in long-horizon interactions. We introduce $τ$-Knowledge, an extension of $τ$-Bench for evaluating agents in environments where success depends on coordinating external, natural-language knowledge with tool outputs to produce verifiable, policy-compliant state changes. Our new domain, $τ$-Banking, models realistic fintech customer support workflows in which agents must navigate roughly 700 interconnected knowledge documents while executing tool-mediated account updates. Across embedding-based retrieval and terminal-based search, even frontier models with high reasoning budgets achieve only $\sim$25.5% pass^1, with reliability degrading sharply over repeated trials. Agents struggle to retrieve the correct documents from densely interlinked knowledge bases and to reason accurately over complex internal policies. Overall, $τ$-Knowledge provides a realistic testbed for developing agents that integrate unstructured knowledge in human-facing deployments.

cs.AI

Spectrum, Tjurina spectrum, and Hertling conjecture for singularities of modality $\leq 3$

Spectrum is an important numerical invariant of an isolated hypersurface singularity, connecting its topological and analytic structures. The well-known Hertling conjecture tells the relation of range and variance of exponents i.e. elements of spectrum. For trimodal singularities, we compute their spectra and verify Hertling conjecture for them. Jung, Kim, Saito and Yoon recently defined Tjurina spectrum, stemming from Hodge ideals. This set of numerical invariants is a subset of spectrum in Steenbrink's sense. We give an estimation of exponents not in Tjurina spectrum and propose a similar Generalized Hertling Conjecture for Tjurina Spectrum. Moreover, we prove the conjecture for singularities of modality $\leq 3$.

math.AG

A remarkable subset of poles of the motivic zeta function

For any polynomial f with complex coefficients we find a remarkable subset of poles of the motivic zeta function. It is combinatorially determined by any log resolution and it admits an intrinsic interpretation in terms of contact loci of f. This uncovers a new, unexpected difficulty with proving the monodromy conjecture.

math.AG

From multitype branching Brownian motions to branching Markov additive processes

We study a class of multitype branching Lévy processes, where particles move according to type-dependent Lévy processes, switch types via an irreducible Markov chain, and branch according to type-dependent laws. This framework generalizes multitype branching Brownian motions. Using techniques of Markov additive processes, we develop a spine decomposition. This approach further enables us to prove convergence results for the additive martingales and derivative martingales, and establish the existence and uniqueness of travelling wave solutions to the corresponding multitype FKPP equations. In particular, applying our results to the on-off branching Brownian motion model resolves several open problems posed by Blath et al.(2025).

math.PR

Atom of Thoughts for Markov LLM Test-Time Scaling

Large Language Models (LLMs) have achieved significant performance gains through test-time scaling methods. However, existing approaches often incur redundant computations due to the accumulation of historical dependency information during inference. To address this challenge, we leverage the memoryless property of Markov processes to minimize reliance on historical context and propose a Markovian reasoning process. This foundational Markov chain structure enables seamless integration with various test-time scaling methods, thereby improving their scaling efficiency. By further scaling up the Markovian reasoning chain through integration with techniques such as tree search and reflective refinement, we uncover an emergent atomic reasoning structure, where reasoning trajectories are decomposed into a series of self-contained, low-complexity atomic units. We name this design Atom of Thoughts (\our). Extensive experiments demonstrate that \our consistently outperforms existing baselines as computational budgets increase. Importantly, \our integrates seamlessly with existing reasoning frameworks and different LLMs (both reasoning and non-reasoning), facilitating scalable, high-performance inference.We submit our code alongside this paper and will make it publicly available to facilitate reproducibility and future research.

cs.CL

Up-down ordered Chinese restaurant processes with two-sided immigration, emigration and diffusion limits

We establish scaling limit theorems for the up-down ordered Chinese restaurant processes (oCRPs) of Rogers and Winkel as processes in a space of interval partitions. As previously conjectured, the limits are self-similar diffusions previously constructed directly in the continuum. We extend the oCRP model and the results to a three-parameter family ${\rm oCRP}^{(α)}(θ_1,θ_2)$, $α\in(0,1)$, $θ_1,θ_2\ge 0$. We use the scaling limit approach to extend existing stationarity results to the full three-parameter family, identifying an extended family of Poisson--Dirichlet interval partitions. Their ranked sequence of interval lengths has Poisson--Dirichlet distribution with parameters $α\in(0,1)$ and $θ:=θ_1+θ_2-α\ge-α$, including for the first time the usual range of $θ>-α$ rather than being restricted to $θ\ge 0$. This has applications to Fleming--Viot processes, nested interval partition evolutions and tree-valued Markov processes, notably relying on the extended parameter range.

math.PR

Hallucination as a Computational Boundary: A Hierarchy of Inevitability and the Oracle Escape

The illusion phenomenon of large language models (LLMs) is the core obstacle to their reliable deployment. This article formalizes the large language model as a probabilistic Turing machine by constructing a "computational necessity hierarchy", and for the first time proves the illusions are inevitable on diagonalization, incomputability, and information theory boundaries supported by the new "learner pump lemma". However, we propose two "escape routes": one is to model Retrieval Enhanced Generations (RAGs) as oracle machines, proving their absolute escape through "computational jumps", providing the first formal theory for the effectiveness of RAGs; The second is to formalize continuous learning as an "internalized oracle" mechanism and implement this path through a novel neural game theory framework. Finally, this article proposes a feasible new principle for artificial intelligence security - Computational Class Alignment (CCA), which requires strict matching between task complexity and the actual computing power of the system, providing theoretical support for the secure application of artificial intelligence.

cs.AI

Stochastic Flows and Marked Stable Processes

We construct a random partition of the space-time plane $\mathbb{R}_+\times \mathbb{R}$ using two coupled stochastic squared Bessel flows, whose parameters differ by $δ\in (0,2)$. We show that the cells of this partition correspond to squared Bessel excursions with a negative parameter $-δ$ which are embedded within the jumps of a spectrally positive $(1+\fracδ2)$ stable process. In particular, we demonstrate that interval partition evolutions [Forman et. al. 2020] and stable shredded disks [Björnberg, Curien and Stefánsson 2022] arise naturally in this framework.

math.PR

LoKI: Low-damage Knowledge Implanting of Large Language Models

Fine-tuning adapts pretrained models for specific tasks but poses the risk of catastrophic forgetting (CF), where critical knowledge from pretraining is overwritten. To address the issue of CF in a general-purpose framework, we propose Low-damage Knowledge Implanting (LoKI), a parameter-efficient fine-tuning (PEFT) technique that utilizes recent mechanistic understanding of how knowledge is stored in transformer architectures. We compare LoKI against state-of-the-art PEFT methods in two real-world fine-tuning scenarios. The results show that LoKI demonstrates significantly better preservation of general capabilities. At the same time, its task-specific performance is comparable to or even surpasses that of full parameter fine-tuning and these PEFT methods across various model architectures. Our work bridges the mechanistic insights of LLMs' knowledge storage with practical fine-tuning objectives, enabling an effective balance between task-specific adaptation and the retention of general-purpose capabilities.

cs.CL

Polar loci of multivariable archimedean zeta functions

We determine, up to exponentiating, the polar locus of the multivariable archimedean zeta function associated to a finite collection of polynomials F. The result is the monodromy support locus of F, a topological invariant. We give a relation between the multiplicities of the irreducible components of the monodromy support locus and the polar orders. These generalize results of Barlet for the case when F is a single polynomial. Our result determines the slopes of the polar locus of the zeta function of F, closing a circle of results of Loeser, Maisonobe, Sabbah. We apply our main result to elucidate the topological information contained by the oblique part of the zero locus of any ideal of Bernstein-Sato type.

math.AG

Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis

Automated generation of high-quality media presentations is challenging, requiring robust content extraction, narrative planning, visual design, and overall quality optimization. Existing methods often produce presentations with logical inconsistencies and suboptimal layouts, thereby struggling to meet professional standards. To address these challenges, we introduce RCPS (Reflective Coherent Presentation Synthesis), a novel framework integrating three key components: (1) Deep Structured Narrative Planning; (2) Adaptive Layout Generation; (3) an Iterative Optimization Loop. Additionally, we propose PREVAL, a preference-based evaluation framework employing rationale-enhanced multi-dimensional models to assess presentation quality across Content, Coherence, and Design. Experimental results demonstrate that RCPS significantly outperforms baseline methods across all quality dimensions, producing presentations that closely approximate human expert standards. PREVAL shows strong correlation with human judgments, validating it as a reliable automated tool for assessing presentation quality.

cs.CL

When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration

Recent advancements in AI reasoning have driven substantial improvements across diverse tasks. A critical open question is whether these improvements also yields better knowledge transfer: the ability of models to communicate reasoning in ways humans can understand, apply, and learn from. To investigate this, we introduce Knowledge Integration and Transfer Evaluation (KITE), a conceptual and experimental framework for Human-AI knowledge transfer capabilities and conduct the first large-scale human study (N=118) explicitly designed to measure it. In our two-phase setup, humans first ideate with an AI on problem-solving strategies, then independently implement solutions, isolating model explanations' influence on human understanding. Our findings reveal that although model benchmark performance correlates with collaborative outcomes, this relationship is notably inconsistent, featuring significant outliers, indicating that knowledge transfer requires dedicated optimization. Our analysis identifies behavioral and strategic factors mediating successful knowledge transfer. We release our code, dataset, and evaluation framework to support future work on communicatively aligned models.

cs.AI