arXiv ScienceSearch

arXiv subjects

Yao Li

Publications and source records attributed to Yao Li.

At least 19 recordsLinked to original sources

QART: A Quantum-Classical Hybrid Architecture for Long-Horizon Reasoning -- Exploring a Conditional Path toward Quantum Scaling

Long-horizon reasoning is vulnerable to early errors that compromise later decisions. We present QART, the Quantum-Augmented Reasoning Transformer, a quantum--classical hybrid architecture combining a backbone language model with quantum encoding, CIM-based QUBO optimization, and quantum decoding. Semantic information can come from hidden representations or model-generated text; detailed encoding and optimization procedures remain proprietary. Under explicit assumptions, we establish a conditional asymptotic reliability separation from single-trajectory autoregressive LLMs. For a common task family with aligned optimality and acceptance criteria, autoregressive acceptance probability tends to zero when cumulative conditional risk of irreversible errors diverges. QART's task-optimal-path recovery probability remains bounded away from zero if conditional probabilities for optimal-path coverage and semantic fidelity, spectral certification, dynamical reachability, and faithful readout remain uniformly positive under a specified resource schedule. The architecture alone does not imply these bounds. Paired measurements on six long-horizon benchmarks using DeepSeek V4 Flash, GLM-5.3, and GPT-5.5 xhigh in a Codex agent environment favor QART in 14 of 15 backbone--benchmark pairs. Relative gains reach 84.0% on SciCode, 47.6% on $τ^3$-Bench, and 44.4% on Terminal-Bench 4.0; the DeepSeek V4 Flash configuration regresses by 7.8% on DeepSWE. These results do not directly validate the asymptotic separation. Potential quantum scaling laws are formulated as conditional hypotheses. A quantum-advantage interpretation requires a demonstrated CIM quantum advantage over strong classical solvers and its transfer to end-to-end reasoning after all system overheads.

cs.AI

SAR-FAH: A Frequency-Adaptive Hybrid Network based on Neural ODEs for Structural-Preserving SAR Despeckling

Synthetic Aperture Radar (SAR) images are inherently degraded by speckle noise that severely limits their reliability in high-precision applications. As a signal-dependent multiplicative noise, speckle noise exhibits distinct statistical properties in homogeneous and heterogeneous regions of SAR images, which are spatially coupled. Nevertheless, existing deep learning despeckling methods operate directly in the spatial domain overlooking this statistical difference. It imposes a suboptimal trade-off between noise suppression and structure preservation, inevitably leading to artifacts, edge blurring, and texture distortion. To address these limitations, we propose a Frequency-Adaptive Hybrid model based on Neural Ordinary Differential Equations (NODEs) for SAR despeckling, termed SAR-FAH. It is a novel divide-and-conquer architecture that performs despeckling in the frequency domain to achieve improved structural preservation. We first fully decouple homogeneous and heterogeneous regions in the frequency domain via wavelet transform according to their local spatial characteristics and then revisit the statistical characteristics of speckle noise. Guided by the distinct properties of each sub-band, we design specialized sub-networks for frequency-specific restoration. Specifically, based on the smoothness of the low-frequency sub-band, the low-frequency denoising process is controlled by the module based on NODEs to ensure sufficient smoothness without artifacts, while the high-frequency sub-bands are processed by enhanced U-Net by incorporating deformable convolutions to better suppress noise and preserve edges and textures. Extensive experiments on both synthetic and real SAR images demonstrate that the proposed SAR-FAH outperforms the state-of-the-art methods both quantitatively and qualitatively.

cs.CV

On the Frankl--Tokushige conjecture and almost complete $r$-cross $t$-intersection theorems for vector spaces

Let $r\geq3$ and $k_1\geq k_2\geq\cdots\geq k_r\geq t$. Let $\mathcal{F}_1,\mathcal{F}_2,\ldots,\mathcal{F}_r$ be families of subspaces, of respective dimensions $k_1,k_2,\ldots,k_r$, in an $n$-dimensional vector space over the finite field $\mathbb{F}_q$. The $r$ families are called $r$-cross $t$-intersecting if $\dim \left(F_{1} \cap F_{2} \cap \cdots \cap F_{r}\right) \geq t$ for all $F_{i} \in \mathcal{F}_{i}, i = 1,2,\dots,r$. In 2016, Frankl and Tokushige conjectured that $\prod_{i=1}^{r}|\mathcal{F}_i|\leq\prod_{i=1}^{r}{n-1\brack k_i-1}$ for $t=1$ and $n\geq rk_1/(r-1)$. The appealing conjecture suggests establishing intersection theorems for $n\sim ck_1$ with $c=c(r)\in(1,2)$, a direction that has long been challenging. In this paper, we overcome this barrier by proving that $$\prod_{i=1}^{r}|\mathcal{F}_i|\leq\prod_{i=1}^{r}{n-t\brack k_i-t}\;\;\mbox{for all}\;\;t\geq1\;\mbox{and}\;n\geq rk_1/(r-1)+C(t,r),$$ where $C(t,r)=rt/(r-1)+1$. This proves the Frankl--Tokushige conjecture except for at most three values of $n$, and establishes an Erdős--Ko--Rado type theorem for almost all values of parameters. Furthermore, we characterize all extremal configurations. Our proof is purely combinatorial and based on the $t$-cover method, with several essential refinements. We also obtain almost complete intersection theorems for $r$-wise $t$-intersecting families and non-trivial $r$-cross $t$-intersecting families.

math.CO

The Late-time Ramp of the Double-Scaled SYK Model from the large-$n$ tail of Cactus Diagram

We derive the late-time ramp of the finite-temperature spectral form factor in the double-scaled SYK model directly from cactus diagrams, which is introduced as a multi-trace generalization of chord diagram in Ref \cite{Berkooz:2020fvm}. Although single-trace observables admit an exact chord-diagram description, multi-trace sums are obstructed by the chord-intersection weights $q_{IJ}$, which cannot be averaged independently to $q$. Our key idea is that the non-analytic contribution responsible for the ramp is controlled by the large-order tail of the Cactus-diagram expansion and is insensitive to finite changes in its low-order analytic terms. Decomposing the contribution with $n$-cross-trace-pairings Cactus diagram into two buds $B_n$ and a kernel $\mathcal K_n$, we determine their large-$n$ asymptotics with fixed $0\leq q<1$ and resum the resulting tail. For $β_L=β+it$ and $β_R=β-it$, we obtain \begin{equation*} Z_s^{\mathrm{sing}}(β+it,β-it) =s_p c_N \frac{|t|}{2π} \int_{E_{min}}^{E_{max}} e^{-2βE} dE+O(1), \qquad s_p=\begin{cases}2,&4\mid p,\\1,&4\nmid p.\end{cases} \end{equation*} This result reproduces the linear ramp predicted by random matrix theory, including its temperature dependence and symmetry factor, and agrees with the semiclassical predictions. It provides a direct microscopic origin of the ramp within the general $q$-deformed quantum algebra. While the plateau lies beyond the scope of the present analysis and calculation.

hep-th

Referenced internal-line decomposition of one-loop string integrand and application to gauge and gravitational beta functions

Following the ideas of Refs.[1,2], we systematically reformulate and generalize the de- composition of one-loop string integrands according to the internal states propagating in a chosen loop channel, extending the construction to both bosonic and heterotic string theories. Using the chiral-splitting formalism, we make manifest a double-copy structure between the left- and right-moving internal loop states. As an application, we analyze the one-loop beta functions of gauge and gravitational couplings from the low-energy field- theory limit of heterotic one-loop three-point amplitudes under a naive $T^6$ compactification. Although maximal supersymmetry forces these beta functions to vanish, our internal-state decomposition provides model-independent expressions for the contributions of different internal loop states. We further show explicitly that both the gravitational beta function and the gravitational correction to the gauge beta function vanish.

hep-th

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We present FinixDoc, an end-to-end agentic parsing system for real-world financial documents, with FinixDoc-VL, a 4B-scale vision-language model built on Qwen3-VL-4B, as its core parser. To characterize the gap between benchmark and deployment performance, we introduce a Document Parsing Capability Matrix organized along two practical axes: visual quality and document scale. Guided by this matrix, FinixDoc-VL is trained with a domain-adapted recipe combining homoglyph-aware contrastive learning and multi-stage reinforcement learning with composite domain-specific rewards. To better leverage our accumulated advantage in low-quality financial-document data and support large-scale, high-quality data production, we further build a human-in-the-loop Data Factory pipeline with confidence-aware expert review. For evaluation, we construct FinixDocBench, a financial-domain evaluation suite covering digital-native, camera-captured, ultra-large-page, and internal-workflow scenarios, with a compliance-reviewed subset released alongside this technical report. On its main subsets, FinixDoc-VL achieves the highest overall score (81.43) among evaluated baselines, outperforming the next-best open-source model by 5.13 points, with the largest gains on internal financial workflows (FinixInner: 84.08 vs. 78.73).

cs.AI

ReTouch: Empowering Contact-Rich Dexterous Manipulation with Online-Refined Tactile Prediction

Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effectively integrating tactile feedback into dexterous manipulation remains underexplored. In this work, we introduce ReTouch, a vision-language-action model (VLA) that supports contact-rich dexterous manipulation through tactile predictions continually refined online using execution-time feedback. ReTouch builds on two main innovations for tactile representation and closed-loop action generation. First, its Tactile-Patch Encoder represents tactile observations as structured tactile patch features that preserve finger identity and local contact structure, providing contact cues for fine-grained dexterous control. Second, its high-frequency action module jointly predicts future tactile states and action chunks and refines both using incoming tactile feedback during execution. This closed-loop refinement keeps tactile predictions aligned with evolving physical interactions, enabling responsive action correction and improving robustness to contact changes and execution errors. We further introduce XHT-Dataset, comprising 900 real-world demonstrations across seven contact-rich tasks collected on an XHand--UR7e platform, and evaluate ReTouch through closed-loop real-robot experiments. ReTouch surpasses the strongest baseline by 18.4 and 23.8 percentage points in average success rate under standard and challenging conditions, respectively, demonstrating its effectiveness and robustness.

cs.RO

Model Card for OpenAI Privacy Filter

OpenAI Privacy Filter is a compact, bidirectional token-classification model for detecting and redacting personally identifiable information (PII) and secrets in unstructured text. The model is derived from an autoregressively pretrained checkpoint and converted into a bidirectional, banded-attention classifier that labels an input sequence in a single forward pass. A constrained Viterbi decoder produces coherent spans across eight privacy categories and exposes configurable operating points for precision-recall tradeoffs. Privacy Filter has 1.5 billion total parameters, 50 million active parameters per token, and a 128,000-token context window. It is designed for efficient local deployment and domain-specific fine-tuning. Privacy Filter is intended as a configurable data-minimization component within layered privacy workflows, not as an anonymization or compliance guarantee.

cs.CR

A Generalized Block Circulant Preconditioner for Crank-Nicolson All-at-Once Systems with Applications to Option Pricing PDEs

The Crank--Nicolson (CN) method is a widely used time integration scheme for evolutionary partial differential equations (PDEs) arising in various scientific and engineering disciplines. Since the numerical solution at each time level depends on the solution at the previous time level, the resulting discretization is inherently sequential and therefore difficult to parallelize in time. In this paper, we develop an all-at-once formulation of the CN discretization together with a generalized block circulant preconditioner that enables an efficient parallel-in-time solution within a Krylov subspace framework. We establish a detailed spectral analysis of the preconditioned system, proving that most eigenvalues are equal to $1$, while the remaining eigenvalues are confined to the annulus: \begin{equation*} \left\{ z\in\mathbb{C}: \frac{1}{1+α}<|z|<\frac{1}{1-α}, \ \Re(z)>0 \right\}, \end{equation*} where $0<α<1$ is a free parameter. Besides, the efficient implementation of the proposed preconditioner is described. Given certain conditions, we prove that the preconditioned GMRES($m$) method achieves a fast convergence rate independent of discretization stepsizes from the residual point of view. Finally, we verify both theoretical findings and the efficacy of the proposed preconditioner via numerical experiments on financial option pricing PDEs (even with variable coefficients).

math.NA

Personalizing Privacy Protection With Individuals' Regulatory Focus: Would You Preserve or Enhance Your Information Privacy?

In this study, we explore the effectiveness of persuasive messages endorsing the adoption of a privacy protection technology (IoT Inspector) tailored to individuals' regulatory focus (promotion or prevention). We explore if and how regulatory fit (i.e., tuning the goal-pursuit mechanism to individuals' internal regulatory focus) can increase persuasion and adoption. We conducted a between-subject experiment (N = 236) presenting participants with the IoT Inspector in gain ("Privacy Enhancing Technology" -- PET) or loss ("Privacy Preserving Technology" -- PPT) framing. Results show that the effect of regulatory fit on adoption is mediated by trust and privacy calculus processes: prevention-focused users who read the PPT message trust the tool more. Furthermore, privacy calculus favors using the tool when promotion-focused individuals read the PET message. We discuss the contribution of understanding the cognitive mechanisms behind regulatory fit in privacy decision-making to support privacy protection.

cs.HC

ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models

LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically relies on fusing bird's-eye-view (BEV) features across agents. However, current BEV representations, typically extracted by LiDAR backbones trained from scratch, are geometry-dominated and lack general semantic priors, inherently limiting the efficacy of feature-level collaboration. Meanwhile, vision foundation models (VFMs) pretrained on large-scale image data have demonstrated strong capability in learning general-purpose and informative visual representations for 2D tasks, and have the potential to enhance agent-wise LiDAR BEV representations for collaboration. Despite this potential, adapting VFMs to LiDAR-based 3D detection remains challenging due to the substantial image-point cloud modality gap. To bridge this gap, we propose ViCo3D, a collaborative 3D object detection framework powered by VFMs. Specifically, ViCo3D adapts VFMs to LiDAR-based collaborative perception from three aspects: First, ViCo3D projects point clouds onto the BEV plane as three-channel images, enabling DINOv2 to extract BEV-space visual features from LiDAR inputs. Besides, to effectively integrate these DINOv2-derived features with LiDAR geometric features, ViCo3D introduces a multi-scale BEV fusion module within the single-agent encoder. In addition, ViCo3D adopts an ego-centric cross-agent fusion strategy to aggregate complementary information from multiple agents. Experiments on DAIR-V2X and V2XSet demonstrate that ViCo3D achieves state-of-the-art 3D detection performance. Remarkably, it delivers up to 1.8x greater collaborative gains than prior methods on DAIR-V2X. The code will be made public available for future investigation.

cs.CV

Primary ICD Category Prediction using LLM-based Probing

Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables. We evaluated whether frozen medical large language model (LLM) representations can serve as a shared embedding space for multimodal primary diagnosis category prediction. Materials and Methods: We constructed a MIMIC-IV cohort of 13,645 admissions from the 10 most frequent primary ICD-10 codes, consolidated into seven categories. Structured variables were serialized into clinical narratives and combined with leakage-pruned discharge notes. Using a frozen MedFound-Llama3-8B-finetuned backbone, we extracted hidden states from five transformer layers and trained linear probes for structured-only, unstructured-only, and combined inputs, comparing against XGBoost and information-matched PLM-ICD baselines and evaluating MIMIC-III adaptation with a compact bottleneck adapter. Results: The combined probe performed best on MIMIC-IV (87.69% strict; 91.45% medical accuracy), exceeding both single-modality probes and baselines. The structured-only probe outperformed its standard baseline by 6.19 points in medical accuracy. Diagnostic information became increasingly linearly separable in deeper layers, and a 2M-parameter adapter restored cross-dataset transfer to MIMIC-III using only 5% of target labels. Discussion: LLM embeddings can unify structured and narrative EHR information for multimodal diagnosis prediction, supporting efficient reuse of clinical representations across modalities and datasets through a small representation-level module. Conclusion: Multimodal probing of frozen medical LLM representations provides a practical approach for studying EHR modalities and adapting clinical representations across datasets.

cs.AI

Mechanical response of quasi-two-dimensional colloidal clusters under uniaxial tension

Despite extensive studies of equilibrium conformations of colloidal clusters, little is known about their mechanical response. Here, we investigate the tensile behavior of a quasi-two-dimensional colloidal cluster subjected to uniaxial tension up to fracture. The sample is a ribbon-shaped assembly of 16 colloidal beads bound by short-range depletion attraction. Using multiple optical tweezers, we clamp the cluster at both ends and perform a tensile test along its long axis. Combining video microscopy with particle tracking, we measure the tensile stress, strain, and particle configurations during deformation. We observe diverse mechanical response behaviors, including elastic, plastic, and soft-mode deformation, with fracture occurring at a strain near 10\%. To explain these behaviors, we construct a spring-mass frame model with breakable elastic bonds. We perform canonical Monte Carlo simulations on the full model with 32 degrees of freedom and compute the statistical distributions of mechanical observables using a simplified model with only 7 degrees of freedom. Both the simulations and the theoretical calculations accurately reproduce the experimental stress--strain curves. Moreover, the configuration distributions predicted by the simplified model agree well with both experiment and simulation in the elastic and soft-mode regimes, with only minor discrepancies in the plastic regime. This work demonstrates that the simplified spring-mass model captures the essential physics governing the rich tensile response behavior of the colloidal cluster.

cond-mat.soft

From Clicks to Intent: Cross-Platform Session Embeddings with LLM-Distilled Taxonomy for Financial Services Recommendations

Sequential user behavior modeling is widely adopted in industrial recommender systems; however, significant gaps remain in financial services, where pre-login web interactions and authenticated in-app experiences differ drastically. Specifically, pre-login web users typically explore new products, whereas logged-in app users focus on account servicing. Due to the challenge of cross-channel entity resolution (e.g., matching anonymous web sessions to authenticated mobile accounts), web-based intent signals remain underutilized for post-authentication personalization. Existing methods for capturing web-based intent are often ad-hoc and narrow, lacking the flexibility to support both quantitative downstream recommendations and qualitative understanding at scale. In this work, we propose a scalable and dual-purpose intent prediction framework for web-based interactions and demonstrate its applicability for personalization. Our approach transforms raw web clickstreams into two outputs: a self-supervised Transformer encodes multi-modal clickstreams into a compact session embedding, while an LLM-based taxonomy generation and distillation pipeline produces interpretable intent labels. Our system demonstrates that self-supervised clickstream representations combined with LLM-distilled taxonomies can jointly serve quantitative tasks and qualitative understanding in production: on the mobile homepage tile ranking task, the session embedding improves macro Recall@1 by 1.88% and reduces Log Loss by 13.38% over production baselines. On the user conversion prediction task, the embedding outperforms the LLM labels by 4.3% on micro F1, while the distillation layer delivers interpretable labels at ultra-low latency with only a 7% performance drop.

cs.IR

Chiral Packings in Cylinders are Ultrasensitive to Confinement Deformation

Sphere packings in circular cylinders have attracted substantial research interest, among which the discovery of chiral helical structures is the most iconic. However, recent experimental results on zebrafish do not match the known packing structures in circular cylinders. To account for the inherent imperfections of biological tubes, we take elliptic cylinders as the canonical deformation of circular cylinders and investigate the densest packings of hard spheres in them using simulation, theory, and experiments. Starting from the chiral structures in circular cylinders, we demonstrate that even a weak cross-sectional deformation can trigger entirely new phases, including ones that either eliminate global chirality or significantly complicate the chiral structures. This reveals the significant effect of cylindrical anisotropy. The new helical phases under anisotropic confinement remain chiral and develop hierarchical periodic structures, which are difficult to obtain by simulations but are predicted by our newly developed theory for helical phases in elliptic cylinders. The theory also predicts double oscillated-chain phases without chirality, which perfectly match the simulations. Our work offers fresh insights into understanding packings in anisotropic cylinders, which will help researchers to design new materials and to understand many living systems.

cond-mat.soft

Latent Geometric Chords for Query-Efficient Decision-Based Adversarial Attacks

While decision-based black-box adversarial attacks present a severe security threat, current methodologies suffer from fundamental limitations. Pixel-wise attacks frequently introduce unnatural, high-frequency visual artifacts, while latent-space frameworks are confined by the limited search space of low-dimensional manifolds and inherent reconstruction flaws. To resolve these limitations, we propose Latent Geometric Chords (LGC) for Query-Efficient Decision-Based Adversarial Attacks alongside a variant, LGC-H. At its core, LGC navigates decision boundaries by executing a curvature-aware geometric search within a compressed semantic manifold. To guarantee high visual fidelity and circumvent dimensionality bottlenecks, we introduce a Residual-based Adversarial Generation (RAG) mechanism. RAG isolates semantic perturbations as geometric chords and superimposes them directly onto the original source image. RAG substantially resolves baseline reconstruction flaws and effectively doubles the permissible search space dimensions. Experimental results demonstrate that LGC achieves robust cross-dataset transferability and substantially outperforms state-of-the-art baselines. Notably, our method, LGC, minimizes perturbation magnitudes while achieving state-of-the-art visual fidelity--with a Structural Similarity Index Measure (SSIM) exceeding 0.99 and a Learned Perceptual Image Patch Similarity (LPIPS) below 0.01 at 5000 queries--and sustaining high attack success rates under stringent perceptual constraints, successfully compromising adversarially trained robust models. The source code is available at: https://github.com/eihmuekhine/Latent-Geometric-Chords.

cs.CV

Density-aware Sample-specific Attack

Despite recent progress in backdoor attacks, existing methods remain susceptible to post-training defenses that erase the backdoor through fine-tuning or pruning. We revisit the core objectives of backdoor attacks and derive principled criteria characterizing optimal sample-specific trigger construction under a Bayes-optimal model of the victim's training. Our analysis reveals that both attack success and clean-accuracy preservation are simultaneously optimized when triggered samples are steered into low-density regions of the clean data distribution, a distributional condition that controls all moments of the poisoned distribution at once rather than a handful of input-space summary statistics. We introduce a bilevel optimization framework that estimates density ratios via conditional time-score matching and optimizes a mixture-model objective to place triggered samples in these sparse regions. Extensive evaluations on MNIST, CIFAR-10, GTSRB, and TinyImageNet demonstrate that our method achieves above 99\% attack success rate before defense and retains 50--85 percentage points higher post-defense ASR than the strongest baselines under fine-tuning defenses. Against neuron-pruning defenses, the method exhibits complete immunity, with zero neurons identified for removal across all pruning thresholds. These results expose a fundamental gap in current defense paradigms and underscore the need for defenses that operate beyond the support of the clean distribution.

cs.LG

The Last Human-Written Paper: Agent-Native Research Artifacts

Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with a machine-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities. Our code is open-sourced at https://github.com/Orchestra-Research/Agent-Native-Research-Artifact.

cs.LG