arXiv ScienceSearch

arXiv subjects

Shi Wang

Publications and source records attributed to Shi Wang.

At least 19 recordsLinked to original sources

QTrans: A Quantum Transformer for Sentiment Classification

In small-scale binary sentiment classification scenarios, factors such as negation, contrastive shifts, and cross-word dependencies lead to the non-linear coupling of sentiment cues, making it difficult for conventional lightweight models to fully capture the contextual relationships between tokens. To address this issue, we propose a model named QTrans, which uses parameterized quantum circuits to construct query, key, and value features and derives attention coefficients from Gaussian distances between quantum measurements. By further integrating a quantum feed-forward neural network, residual connections, and layer normalization, the model establishes an end-to-end trainable quantum-classical hybrid framework for sentiment classification. Experimental results on the MR, CR, and MPQA datasets show that QTrans achieves test accuracies of 72.13\%, 69.51\%, and 63.45\%, respectively, representing improvements of 2.88, 3.17, and 3.79 percentage points over the best-performing classical baselines for each dataset. Overall, QTrans expands the application of parameterized quantum circuits in lightweight sentiment analysis and lays an experimental foundation for further research into quantum multi-head self-attention for modeling textual relationships.

cs.LG

Actions on CAT(-1) spaces with critical exponent less than 1

We show that for a discrete isometry subgroup acting on a proper CAT(-1) space X, if the critical exponent is less than $1$, then the critical exponent equals the Hausdorff dimension of the entire limit set. Consequently, the limit set must be a Cantor set. As an application, we prove that any finitely generated, torsion-free discrete subgroup in Isom(X) with critical exponent less than one must be geometrically finite and free. This answers a question of Kapovich.

math.GT

Homological dimension of discrete subgroups in higher rank simple Lie groups

We investigate the homological dimensions of discrete subgroups of non-compact simple Lie groups using a flow recently introduced by the authors. Using new estimates on the critical index for a broad class of discrete subgroups, we prove that the homological dimension of non-lattice, discrete, Zariski-dense subgroups is bounded above by $n-αr$ where $n$ is the dimension of the associated symmetric space, $r$ is the real rank, and $α\geq \frac{1}{8}$. Under the assumption that the discrete group has regular limit cone, our bound improves to $\frac{5}{6}n+1$. Additionally, we make some conjectures on the possible homological dimensions for non-lattice discrete subgroups and pose a few questions in the more general semisimple setting.

math.GT

Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration

Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, educational misuse, and platform governance. These concerns make robust detection of machine-generated text increasingly necessary. Recent zero-shot detectors mainly exploit probability-based statistical discrepancies, but they do not explicitly account for the training process of LLMs, which leaves a distinct generation mechanism insufficiently modeled and limits detection robustness. To address this issue, we propose EchoPrompt, a training-free detector based on latent prompt restoration. Our key intuition is that machine-generated text is typically produced conditioned on an upstream prompt, and this hidden dependency can be partially reactivated by prepending a unified generic prefix. Specifically, EchoPrompt restores a generic assistant-response context, measures the induced likelihood gain with an instruction-tuned model, calibrates it against the corresponding base model, and aggregates the resulting differences into a score that quantifies latent prompt dependency. Extensive experiments show that EchoPrompt achieves state-of-the-art performance among zero-shot detectors while maintaining strong robustness across challenging evaluation settings.

cs.CL

Evaluating noises of fast-simulated boson sampling with statistical benchmark methods

It is important to know noise levels of boson sampling in order to cautiously demonstrate the quantum computational advantage or realize certain tasks. Based on those statistical benchmark methods such as the correlators and clouds, which are initially proposed to discriminate boson sampling and other mockups, we quantificationally evaluate noises of photon partial distinguishability and photon loss compensated by dark counts. This is feasible owing to the fact that the output distribution unbalances are suppressed by noises, which are actually results of multi-photon interferences. This is why the evaluation performance is better when high order correlators or correspondent clouds are employed. Our results indicate that the statistical benchmark methods can also work in the task of evaluating noises of boson sampling. An effective scheme is also introduced to fast simulate noisy samples, especially those with photon partial distinguishability.

quant-ph

Average signature of geodesic paths in compact Lie groups

For any compact connected Lie group $G$, we introduce a novel notion of average signature $\mathbb A(G)$ valued in its tensor Lie algebra, by taking the average value of the signature of the unique length-minimizing geodesics between all pairs of generic points in $G$. we prove that using the average signature together with the trace operation with respect to the given bi-invariant Riemannian metric on $G$, one can recover certain geometric quantities of $G$, including the dimension, the diameter, the volume and the scalar curvature.

math.DG

Experimental Advances on Light Baryon Spectroscopy at BESIII Experiment

The BESIII experiment is currently the world's only electron-positron collider operating in the tau-charm physical energy region. Since starting data taking in 2009, BESIII has accumulated the world's largest data set in the center-of-mass energy range of 1.84-4.95 GeV, including approximately 10 billion $J/ψ$ events and 3 billion $ψ(3686)$ events, together with extensive data on open-charm hadron pair production near threshold regions. These unique datasets, characterized by high statistics and low background, provide unprecedented experimental conditions for studying light baryon spectroscopy. This article systematically reviews the progress made by BESIII in baryon spectroscopy, with a focus on recent breakthrough achievements, including the discovery of excited nucleon states, $Λ$ hyperon states, $Σ$ hyperon states, $Ξ$ hyperon states and $Ω^{-}$ hyperon states. These results expand the spectrum of baryon excited states and provide crucial experimental support for understanding non-perturbative QCD and resolving the ``missing baryon resonances'' problem.

hep-ex

Extended validations on photon number resolving detector based Gaussian boson sampling with low noises

Gaussian boson sampling (GBS) is a variety of boson sampling overcoming the stable single-photon preparation difficulty of the later. However, like those in the original version, noises in GBS will also result in the deviation of output patterns and the reduction of classical simulation complexity. We extend the pattern recognition validation, together with the correlation approach as a comparison, on GBS using photon number resolving detectors, with noises of both photon loss and distinguishability, to quantificationally evaluate noise levels. As for the classical simulation with noises to be used during validations, it is actually a simulation of mixed states where we employ an existing photon-pair strategy to realize polynomial speedup locally. Furthermore, we use an output-binning strategy to realize validation speedup. Our simulation indicates that the pattern recognition protocol is useful for noise evaluations of GBS even when noises are sufficiently low.

quant-ph

DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack

With the rapid development of cloud-based services, large language models have become increasingly accessible through various web platforms. However, this accessibility has also led to growing risks of model abuse. LLM watermarking has emerged as an effective approach to mitigate such misuse and protect intellectual property. Existing watermarking algorithms, however, primarily focus on defending against paraphrase attacks while overlooking piggyback spoofing attacks, which can inject harmful content, compromise watermark reliability, and undermine trust in attribution. To address this limitation, we propose DualGuard, the first watermarking algorithm capable of defending against both paraphrase and spoofing attacks. DualGuard employs the adaptive dual-stream watermarking mechanism, in which two complementary watermark signals are dynamically injected based on the semantic content. This design enables DualGuard not only to detect but also to trace spoofing attacks, thereby ensuring reliable and trustworthy watermark detection. Extensive experiments conducted across multiple datasets and language models demonstrate that DualGuard achieves excellent detectability, robustness, traceability, and text quality, effectively advancing the state of LLM watermarking for real-world applications.

cs.CR

The natural flow and the critical exponent

Inspired by work of Besson-Courtois-Gallot, we construct a flow called the natural flow on a non-positively curved Riemannian manifold $M$. As with the natural map, the $k$-Jacobian of the natural flow is directly related to the critical exponent $δ$ of the fundamental group. There are several applications of the natural flow that connect dynamical, geometrical, and topological invariants of the manifold. First, we give $k$-dimensional linear isoperimetric inequalities when $k > δ$. This, in turn, produces lower bounds on the Cheeger constant. We resolve a recent conjecture of Dey-Kapovich on the non-existence of $k$-dimensional compact, complex subvarieties of complex hyperbolic manifolds with $2k > δ$. We also provide upper bounds on the homological dimension, generalizing work of Kapovich and work of Farb with the first two authors. Using the natural flow together with Morse theory, we also give upper bounds on the cohomological dimension, which partially resolve a conjecture of Kapovich. Finally, we introduce a new growth condition on the Bowen-Margulis measure that we call uniformly exponentially bounded that we connect to the cohomological dimension and which could be of independent interest.

math.DG

Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text Detection

The rapid advancement of large language models has increasingly blurred the boundary between human-written and AI-generated text, raising societal risks such as misinformation dissemination, authorship ambiguity, and threats to intellectual property rights. These concerns highlight the urgent need for effective and reliable detection methods. While existing training-free approaches often achieve strong performance by aggregating token-level signals into a global score, they typically assume uniform token contributions, making them less robust under short sequences or localized token modifications. To address these limitations, we propose Exons-Detect, a training-free method for AI-generated text detection based on an exon-aware token reweighting perspective. Exons-Detect identifies and amplifies informative exonic tokens by measuring hidden-state discrepancy under a dual-model setting, and computes an interpretable translation score from the resulting importance-weighted token sequence. Empirical evaluations demonstrate that Exons-Detect achieves state-of-the-art detection performance and exhibits strong robustness to adversarial attacks and varying input lengths. In particular, it attains a 2.2\% relative improvement in average AUROC over the strongest prior baseline on DetectRL.

cs.CL

Signature inversion of $C^1-$axial linear curves

We introduce a signature inversion scheme for $C^1$-axial linear curves which are widely used in various areas. We show that in the presence of a linear coordinate function, the derivatives of the underlying curve at any point $x$ can be recovered by tracking the signature coefficients $S_{k,l}$ with $\frac{k}{k+l} \to x$. We furthermore give a quantitative estimates for the convergence rate in this inversion scheme and establish a modulus of continuity of the signature inverse $S^{-1}$ under different topologies by using this inversion procedure.

math.FA

Noise-Resistant Feature-Aware Attack Detection Using Quantum Machine Learning

Continuous-variable quantum key distribution (CV-QKD) is a quantum communication technology that offers an unconditional security guarantee. However, the practical deployment of CV-QKD systems remains vulnerable to various quantum attacks. In this paper, we propose a quantum machine learning (QML)-based attack detection framework (QML-ADF) that safeguards the security of high-rate CV-QKD systems. In particular, two alternative QML models -- quantum support vector machines (QSVM) and quantum neural networks (QNN) -- are developed to perform noise-resistant and feature-aware attack detection before conventional data postprocessing. Leveraging feature-rich quantum data from Gaussian modulation and homodyne detection, the QML-ADF effectively detects quantum attacks, including both known and unknown types defined by these distinctive features. The results indicate that all twelve distinct QML variants for both QSVM and QNN exhibit remarkable performance in detecting both known and previously undiscovered quantum attacks, with the best-performing QSVM variant outperforming the top QNN counterpart. Furthermore, we systematically evaluate the performance of the QML-ADF under various physically interpretable noise backends, demonstrating its strong robustness and superior detection performance. We anticipate that the QML-ADF will not only enable robust detection of quantum attacks under realistic deployment conditions but also strengthen the practical security of quantum communication systems.

quant-ph

On the Jacobian of the Douady-Earle extension

Given an isotopy class between two closed hyperbolic surfaces, the Douady--Earle extension provides a unique analytic diffeomorphism representative. In this paper we investigate the Jacobian of the Douady--Earle extension map $F$. We prove that $|\operatorname{Jac} F| \equiv 1$ precisely when $F$ is an isometry. Moreover, we construct a sequence of hyperbolic surfaces $\{Σ_i\}$ together with a fixed domain surface $Σ_0$ for which the Douady--Earle extension maps $F_i:Σ_0\toΣ_i$ satisfy $\max_{x\inΣ_0} \operatorname{Jac} F_i \to +\infty$.

math.GT

PathMind: A Retrieve-Prioritize-Reason Framework for Knowledge Graph Reasoning with Large Language Models

Knowledge graph reasoning (KGR) is the task of inferring new knowledge by performing logical deductions on knowledge graphs. Recently, large language models (LLMs) have demonstrated remarkable performance in complex reasoning tasks. Despite promising success, current LLM-based KGR methods still face two critical limitations. First, existing methods often extract reasoning paths indiscriminately, without assessing their different importance, which may introduce irrelevant noise that misleads LLMs. Second, while many methods leverage LLMs to dynamically explore potential reasoning paths, they require high retrieval demands and frequent LLM calls. To address these limitations, we propose PathMind, a novel framework designed to enhance faithful and interpretable reasoning by selectively guiding LLMs with important reasoning paths. Specifically, PathMind follows a "Retrieve-Prioritize-Reason" paradigm. First, it retrieves a query subgraph from KG through the retrieval module. Next, it introduces a path prioritization mechanism that identifies important reasoning paths using a semantic-aware path priority function, which simultaneously considers the accumulative cost and the estimated future cost for reaching the target. Finally, PathMind generates accurate and logically consistent responses via a dual-phase training strategy, including task-specific instruction tuning and path-wise preference alignment. Extensive experiments on benchmark datasets demonstrate that PathMind consistently outperforms competitive baselines, particularly on complex reasoning tasks with fewer input tokens, by identifying essential reasoning paths.

cs.AI

DNA-DetectLLM: Unveiling AI-Generated Text via a DNA-Inspired Mutation-Repair Paradigm

The rapid advancement of large language models (LLMs) has blurred the line between AI-generated and human-written text. This progress brings societal risks such as misinformation, authorship ambiguity, and intellectual property concerns, highlighting the urgent need for reliable AI-generated text detection methods. However, recent advances in generative language modeling have resulted in significant overlap between the feature distributions of human-written and AI-generated text, blurring classification boundaries and making accurate detection increasingly challenging. To address the above challenges, we propose a DNA-inspired perspective, leveraging a repair-based process to directly and interpretably capture the intrinsic differences between human-written and AI-generated text. Building on this perspective, we introduce DNA-DetectLLM, a zero-shot detection method for distinguishing AI-generated and human-written text. The method constructs an ideal AI-generated sequence for each input, iteratively repairs non-optimal tokens, and quantifies the cumulative repair effort as an interpretable detection signal. Empirical evaluations demonstrate that our method achieves state-of-the-art detection performance and exhibits strong robustness against various adversarial attacks and input lengths. Specifically, DNA-DetectLLM achieves relative improvements of 5.55% in AUROC and 2.08% in F1 score across multiple public benchmark datasets. Code and data are available at https://github.com/Xiaoweizhu57/DNA-DetectLLM.

cs.CL

Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning

Knowledge graph completion (KGC) aims to infer new knowledge and make predictions from knowledge graphs. Recently, large language models (LLMs) have exhibited remarkable reasoning capabilities. LLM-enhanced KGC methods primarily focus on designing task-specific instructions, achieving promising advancements. However, there are still two critical challenges. First, existing methods often ignore the inconsistent representation spaces between natural language and graph structures. Second, most approaches design separate instructions for different KGC tasks, leading to duplicate works and time-consuming processes. To address these challenges, we propose SAT, a novel framework that enhances LLMs for KGC via structure-aware alignment-tuning. Specifically, we first introduce hierarchical knowledge alignment to align graph embeddings with the natural language space through multi-task contrastive learning. Then, we propose structural instruction tuning to guide LLMs in performing structure-aware reasoning over KGs, using a unified graph instruction combined with a lightweight knowledge adapter. Experimental results on two KGC tasks across four benchmark datasets demonstrate that SAT significantly outperforms state-of-the-art methods, especially in the link prediction task with improvements ranging from 8.7% to 29.8%.

cs.CL

ABKD: Pursuing a Proper Allocation of the Probability Mass in Knowledge Distillation via $α$-$β$-Divergence

Knowledge Distillation (KD) transfers knowledge from a large teacher model to a smaller student model by minimizing the divergence between their output distributions, typically using forward Kullback-Leibler divergence (FKLD) or reverse KLD (RKLD). It has become an effective training paradigm due to the broader supervision information provided by the teacher distribution compared to one-hot labels. We identify that the core challenge in KD lies in balancing two mode-concentration effects: the \textbf{\textit{Hardness-Concentration}} effect, which refers to focusing on modes with large errors, and the \textbf{\textit{Confidence-Concentration}} effect, which refers to focusing on modes with high student confidence. Through an analysis of how probabilities are reassigned during gradient updates, we observe that these two effects are entangled in FKLD and RKLD, but in extreme forms. Specifically, both are too weak in FKLD, causing the student to fail to concentrate on the target class. In contrast, both are too strong in RKLD, causing the student to overly emphasize the target class while ignoring the broader distributional information from the teacher. To address this imbalance, we propose ABKD, a generic framework with $α$-$β$-divergence. Our theoretical results show that ABKD offers a smooth interpolation between FKLD and RKLD, achieving an effective trade-off between these effects. Extensive experiments on 17 language/vision datasets with 12 teacher-student settings confirm its efficacy. The code is available at https://github.com/ghwang-s/abkd.

cs.LG