arXiv ScienceSearch

arXiv subjects

Yu Tong

Publications and source records attributed to Yu Tong.

At least 19 recordsLinked to original sources

VERPO: Verified Evidence Regularized Policy Optimization

Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts that do not support task success. We introduce VERPO, a Verified Evidence Regularized Policy Optimization framework that treats evidence as a proposal for policy correction while retaining the outcome objective. It separates evidence-free reference restoration from signed token-level evidence corrections. Fisher Evidence Contrast attenuates corrections along an estimated evidence-presence direction. A stopped token-wise ZPD controller scales acceptance according to local reward alignment and Fisher movement cost, while the reference channel remains independent of acceptance. Across five scientific-reasoning and tool-use tasks, the best variant on each backbone exceeds the strongest compared baseline in average score. The averages rise from 0.6826 to 0.6857 on Qwen3-4B, from 0.6895 to 0.7058 on Qwen3-8B, and from 0.4751 to 0.5657 on Llama-3.2-1B.

cs.LG

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.

cs.AI

Efficient Classical Simulation of Weakly Interacting Fermion Dynamics

We consider the task of simulating the real-time dynamics of weakly interacting fermionic systems. In particular, we focus on computing the expectation value of a local observable $A$ at time $t$. By analyzing the convergence of the perturbative expansion in the interaction strength $\lambda$ for the Heisenberg-picture observable, we propose a polynomial-time algorithm for estimating this expectation value in the weakly interacting regime $\lambda |t|^{2D+1}=\mathcal{O}(1)$, when the Hamiltonian is geometrically local on a $D$-dimensional lattice. Importantly, this condition is independent of the system size. If the goal is instead to approximate the time-evolved observable in normalized Frobenius norm, we extend the convergence regime to $\lambda |t|=\mathcal{O}(1)$ with quasi-polynomial runtime. When the non-interacting part exhibits Anderson localization, our polynomial-time algorithm can be extended up to $\lambda |t|=\mathcal{O}(1)$, modulo polylogarithmic factors. Our algorithm brings together ideas from continuous-time QMC, diagrammatic QMC, and Majorana Propagation, but with a new Heisenberg-picture operator-growth analysis that makes the sampling complexity rigorously controllable. This leads to provably efficient classical algorithms in regimes where the interaction is weak enough that the sampling variance remains bounded independently of system size. Together, these results identify broad regimes in which weak interactions, locality, and localization can be leveraged to make real-time fermionic dynamics classically tractable.

quant-ph

From Memorization to Creation: Evaluating the Cognitive Depth of LLM-Generated Educational Questions

While LLMs show promise in automating educational content creation, their ability to generate questions that stimulate higher-order thinking remains understudied. This work evaluates six widely-used LLMs through a Bloom's Taxonomy lens, focusing on their capacity to transcend rote memorization and achieve cognitive leaps. Using a hybrid human--AI evaluation protocol, we generate and analyze 20{,}700 questions across computer science, K--12 math, and social-science domains. Key contributions include: (1) a fine-grained prompting strategy that reduces question repetitiveness by 24.45\% for Qwen2.5-7B-Instruct, and increases the proportion of higher-order cognitive level outputs by 11.53\% for InternLM3-8B-Instruct; (2) quantitative metrics for cognitive shift intensity (CogShift) and category drift, revealing InternLM3's superior performance in multi-level transitions; (3) an interpretability analysis revealing metric-level correlations that enhance the transparency of Chain-of-Thought prompting. Our findings highlight the importance of cognitive-aware prompt design and provide benchmarks for deploying LLMs in personalized learning systems.

cs.HC

Behavioral Systems Theory Meets Machine Learning: Control-Aware Learning of the Intrinsic Behavior from Big Data

The abundance of process operating data in modern industries, along with the rapid advancement of learning techniques, has led to a paradigm shift towards data-centric analysis and control. However, integrating machine learning with control theory for big data-driven control of nonlinear systems remains a challenging open problem. This is because the state-based, model-centric, and causal framework of classical control theory fundamentally contradicts the trajectory-based, set-theoretic, and causality-absent rationale of big data-based learning approaches. Using the behavioral framework, we show that dynamical systems possess an intrinsic state variable that encodes the system behavior in a bijective and causality-free manner, and control design can be carried out entirely within the state space. This approach not only resolves the aforementioned conflict but also complements machine learning techniques well, leading to a neural network architecture that is capable of learning the behavior representation well-suited for control design.

eess.SY

Quantum Gibbs sampling through the detectability lemma

Gibbs state preparation is an important subroutine in quantum computing. In this work we use the detectability lemma to improve Gibbs state preparation. Specifically, we design new Gibbs state preparation methods that do not rely on simulating Lindbladian evolution, thus avoiding the overhead from it. For local Lindbladians consisting of $M$ terms, this approach reduces the cost by a factor of $O(M)$. We also combine the detectability lemma operator and quantum singular value transformation to implement ground state projection operators of frustration-free Hamiltonians, resulting in a quadratic speedup in the spectral gap dependence. Applying this method to Lindbladians for the Gibbs state of local commuting Hamiltonians, we achieve quadratically better dependence on the Lindbladian spectral gap.

quant-ph

Autonomous Hamiltonian certification and changepoint detection

Modern quantum devices require high-precision Hamiltonian dynamics, but environmental noise can cause calibrated Hamiltonian parameters to drift over time, necessitating expensive recalibration. Detecting when recalibration is needed is challenging, especially since the very gates required for sophisticated verification protocols may themselves be miscalibrated. While cloud quantum computing services implement heuristic routines for triggering recalibration, the fundamental limits of optimal recalibration are not yet known. We develop efficient Hamiltonian certification and changepoint detection protocols in the autonomous setting, where we cannot rely on an external noiseless device and use only single-qubit gates and measurements, making the protocols robust to the calibration issues for multi-qubit operations they aim to detect. For unknown $n$-qubit Hamiltonians $H$ and $H_0$ with operator norm bounded by $M$, our certification protocol distinguishes whether $\|H-H_0\|_F\geq\epsilon$ or $\|H-H_0\|_F\leq O(\epsilon/\sqrt{n})$ with sample complexity $O(nM^2\ln(1/\delta)/\epsilon^2)$ and total evolution time $O(nM\ln(1/\delta)/\epsilon^2)$. We achieve this by evolving random stabilizer product states and performing adaptive single-qubit measurements based on a classically simulable hypothesis state. Extending this to continuous monitoring, we develop an online changepoint detection algorithm using the CUSUM procedure that achieves a detection delay time bound of $O(nM\ln(M\mathbb{E}_\infty[T])/\epsilon^2)$, matching the known asymptotically optimal scaling with respect to false alarm run time $\mathbb{E}_\infty[T]$. Our approach enables quantum devices to autonomously monitor their own calibration status without requiring ancillary systems, entangling operations, or a trusted reference device, offering a practical solution for robust quantum computing with contemporary noisy devices.

quant-ph

Learning Hamiltonians in the Heisenberg limit with static single-qubit fields

Learning the Hamiltonian governing a quantum system is a central task in quantum metrology, sensing, and device characterization. Existing Heisenberg-limited Hamiltonian learning protocols either require multi-qubit operations that are prone to noise, or single-qubit operations whose frequency or strength increases with the desired precision. These two requirements limit the applicability of Hamiltonian learning on near-term quantum platforms. We present a protocol that learns a quantum Hamiltonian with the optimal Heisenberg-limited scaling using only single-qubit control in the form of static fields with strengths that are independent of the target precision. Our protocol is robust against the state preparation and measurement (SPAM) error. By overcoming these limitations, our protocol provides new tools for device characterization and quantum sensing. We demonstrate that our method achieves the Heisenberg-limited scaling through rigorous mathematical proof and numerical experiments. We also prove an information-theoretic lower bound showing that a non-vanishing static field strength is necessary for achieving the Heisenberg limit unless one employs an extensive number of discrete control operations.

quant-ph

Convergence of the Cumulant Expansion and Polynomial-Time Algorithm for Weakly Interacting Fermions

We propose a randomized algorithm to compute the log-partition function of weakly interacting fermions with polynomial runtime in both the system size and precision. Although weakly interacting fermionic systems are considered tractable for many computational methods such as the diagrammatic quantum Monte Carlo, a mathematically rigorous proof of polynomial runtime has been lacking. In this work we first extend the proof techniques developed in previous works for proving the convergence of the cumulant expansion in periodic systems to the non-periodic case. A key equation used to analyze the sum of connected Feynman diagrams, which we call the tree-determinant expansion, reveals an underlying tree structure in the summation. This enables us to design a new randomized algorithm to compute the log-partition function through importance sampling augmented by belief propagation. This approach differs from the traditional method based on Markov chain Monte Carlo, whose efficiency is hard to guarantee, and enables us to obtain a algorithm with provable polynomial runtime.

quant-ph

Improved Hamiltonian learning and sparsity testing through Bell sampling

We consider the problem of learning an $M$-sparse Hamiltonian and the related problem of Hamiltonian sparsity testing. Through a detailed analysis of Bell sampling, we reduce the total evolution time required by the state-of-the-art algorithm for $M$-sparse Hamiltonian learning to $\widetilde{\mathcal{O}}(M/\epsilon)$, where $\epsilon$ denotes the $\ell^{\infty}$ error, achieving an improvement by a factor of $M$ (ignoring the logarithmic factor) while only requiring access to forward time-evolution. We then establish a connection between Hamiltonian learning and Hamiltonian sparsity testing through Bell sampling, which enables us to propose a Hamiltonian sparsity testing with state-of-the-art total evolution time scaling.

quant-ph

Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate Prediction

In click-through rate prediction, click-through rate prediction is used to model users' interests. However, most of the existing CTR prediction methods are mainly based on the ID modality. As a result, they are unable to comprehensively model users' multi-modal preferences. Therefore, it is necessary to introduce multi-modal CTR prediction. Although it seems appealing to directly apply the existing multi-modal fusion methods to click-through rate prediction models, these methods (1) fail to effectively disentangle commonalities and specificities across different modalities; (2) fail to consider the synergistic effects between modalities and model the complex interactions between modalities. To address the above issues, this paper proposes the Diffusion-based Multi-modal Synergy Interest Network (Diff-MSIN) framework for click-through prediction. This framework introduces three innovative modules: the Multi-modal Feature Enhancement (MFE) Module Synergistic Relationship Capture (SRC) Module, and the Feature Dynamic Adaptive Fusion (FDAF) Module. The MFE Module and SRC Module extract synergistic, common, and special information among different modalities. They effectively enhances the representation of the modalities, improving the overall quality of the fusion. To encourage distinctiveness among different features, we design a Knowledge Decoupling method. Additionally, the FDAF Module focuses on capturing user preferences and reducing fusion noise. To validate the effectiveness of the Diff-MSIN framework, we conducted extensive experiments using the Rec-Tmall and three Amazon datasets. The results demonstrate that our approach yields a significant improvement of at least 1.67% compared to the baseline, highlighting its potential for enhancing multi-modal recommendation systems. Our code is available at the following link: https://github.com/Cxx-0/Diff-MSIN.

cs.IR

Multi-Modal Multi-Behavior Sequential Recommendation with Conditional Diffusion-Based Feature Denoising

The sequential recommendation system utilizes historical user interactions to predict preferences. Effectively integrating diverse user behavior patterns with rich multimodal information of items to enhance the accuracy of sequential recommendations is an emerging and challenging research direction. This paper focuses on the problem of multi-modal multi-behavior sequential recommendation, aiming to address the following challenges: (1) the lack of effective characterization of modal preferences across different behaviors, as user attention to different item modalities varies depending on the behavior; (2) the difficulty of effectively mitigating implicit noise in user behavior, such as unintended actions like accidental clicks; (3) the inability to handle modality noise in multi-modal representations, which further impacts the accurate modeling of user preferences. To tackle these issues, we propose a novel Multi-Modal Multi-Behavior Sequential Recommendation model (M$^3$BSR). This model first removes noise in multi-modal representations using a Conditional Diffusion Modality Denoising Layer. Subsequently, it utilizes deep behavioral information to guide the denoising of shallow behavioral data, thereby alleviating the impact of noise in implicit feedback through Conditional Diffusion Behavior Denoising. Finally, by introducing a Multi-Expert Interest Extraction Layer, M$^3$BSR explicitly models the common and specific interests across behaviors and modalities to enhance recommendation performance. Experimental results indicate that M$^3$BSR significantly outperforms existing state-of-the-art methods on benchmark datasets.

cs.IR

Qubit-Efficient Quantum Algorithm for Linear Differential Equations

As quantum hardware rapidly advances toward the early fault-tolerant era, a key challenge is to develop quantum algorithms that are not only theoretically sound but also hardware-friendly on near-term devices. In this work, we propose a quantum algorithm for solving linear ordinary differential equations (ODEs) with a provable runtime guarantee. Our algorithm uses only a single ancilla qubit, and is locality preserving, i.e., when the coefficient matrix of the ODE is $k$-local, the algorithm only needs to implement the time evolution of $(k+1)$-local Hamiltonians. We also discuss the connection between our proposed algorithm and Lindbladian simulation. By applying our algorithm to the interacting Hatano-Nelson model, a widely studied non-Hermitian model with rich phenomenology, and numerically simulating it under realistic noise models, we demonstrate its practical feasibility on near-term quantum devices.

quant-ph

Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models

Instruction-following capability has become a major ability to be evaluated for Large Language Models (LLMs). However, existing datasets, such as IFEval, are either predominantly monolingual and centered on English or simply machine translated to other languages, limiting their applicability in multilingual contexts. In this paper, we present an carefully-curated extension of IFEval to a localized multilingual version named Marco-Bench-MIF, covering 30 languages with varying levels of localization. Our benchmark addresses linguistic constraints (e.g., modifying capitalization requirements for Chinese) and cultural references (e.g., substituting region-specific company names in prompts) via a hybrid pipeline combining translation with verification. Through comprehensive evaluation of 20+ LLMs on our Marco-Bench-MIF, we found that: (1) 25-35% accuracy gap between high/low-resource languages, (2) model scales largely impact performance by 45-60% yet persists script-specific challenges, and (3) machine-translated data underestimates accuracy by7-22% versus localized data. Our analysis identifies challenges in multilingual instruction following, including keyword consistency preservation and compositional constraint adherence across languages. Our Marco-Bench-MIF is available at https://github.com/AIDC-AI/Marco-Bench-MIF.

cs.CL

Heisenberg-limited Hamiltonian learning continuous variable systems via engineered dissipation

Discrete and continuous variables oftentimes require different treatments in many learning tasks. Identifying the Hamiltonian governing the evolution of a quantum system is a fundamental task in quantum learning theory. While previous works mostly focused on quantum spin systems, where quantum states can be seen as superpositions of discrete bit-strings, relatively little is known about Hamiltonian learning for continuous-variable quantum systems. In this work we focus on learning the Hamiltonian of a bosonic quantum system, a common type of continuous-variable quantum system. This learning task involves an infinite-dimensional Hilbert space and unbounded operators, making mathematically rigorous treatments challenging. We introduce an analytic framework to study the effects of strong dissipation in such systems, enabling a rigorous analysis of cat qubit stabilization via engineered dissipation. This framework also supports the development of Heisenberg-limited algorithms for learning general bosonic Hamiltonians with higher-order terms of the creation and annihilation operators. Notably, our scheme requires a total Hamiltonian evolution time that scales only logarithmically with the number of modes and inversely with the precision of the reconstructed coefficients. On a theoretical level, we derive a new quantitative adiabatic approximation estimate for general Lindbladian evolutions with unbounded generators. Finally, we discuss possible experimental implementations.

quant-ph

Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs

Adversarial attacks aim to generate malicious inputs that mislead deep models, but beyond causing model failure, they cannot provide certain interpretable information such as ``\textit{What content in inputs make models more likely to fail?}'' However, this information is crucial for researchers to specifically improve model robustness. Recent research suggests that models may be particularly sensitive to certain semantics in visual inputs (such as ``wet,'' ``foggy''), making them prone to errors. Inspired by this, in this paper we conducted the first exploration on large vision-language models (LVLMs) and found that LVLMs indeed are susceptible to hallucinations and various errors when facing specific semantic concepts in images. To efficiently search for these sensitive concepts, we integrated large language models (LLMs) and text-to-image (T2I) models to propose a novel semantic evolution framework. Randomly initialized semantic concepts undergo LLM-based crossover and mutation operations to form image descriptions, which are then converted by T2I models into visual inputs for LVLMs. The task-specific performance of LVLMs on each input is quantified as fitness scores for the involved semantics and serves as reward signals to further guide LLMs in exploring concepts that induce LVLMs. Extensive experiments on seven mainstream LVLMs and two multimodal tasks demonstrate the effectiveness of our method. Additionally, we provide interesting findings about the sensitive semantics of LVLMs, aiming to inspire further in-depth research.

cs.CV

Training-Free Watermarking for Autoregressive Image Generation

Invisible image watermarking can protect image ownership and prevent malicious misuse of visual generative models. However, existing generative watermarking methods are mainly designed for diffusion models while watermarking for autoregressive image generation models remains largely underexplored. We propose IndexMark, a training-free watermarking framework for autoregressive image generation models. IndexMark is inspired by the redundancy property of the codebook: replacing autoregressively generated indices with similar indices produces negligible visual differences. The core component in IndexMark is a simple yet effective match-then-replace method, which carefully selects watermark tokens from the codebook based on token similarity, and promotes the use of watermark tokens through token replacement, thereby embedding the watermark without affecting the image quality. Watermark verification is achieved by calculating the proportion of watermark tokens in generated images, with precision further improved by an Index Encoder. Furthermore, we introduce an auxiliary validation scheme to enhance robustness against cropping attacks. Experiments demonstrate that IndexMark achieves state-of-the-art performance in terms of image quality and verification accuracy, and exhibits robustness against various perturbations, including cropping, noises, Gaussian blur, random erasing, color jittering, and JPEG compression.

cs.CV

State-space gradient descent and metastability in quantum systems

We propose a quantum algorithm, inspired by ADAPT-VQE, to variationally prepare the ground state of a quantum Hamiltonian, with the desirable property that if it fails to find the ground state, it still yields a physically meaningful local-minimum state that oftentimes corresponds to a metastable state of the quantum system. At each iteration, our algorithm reduces the energy using a set of local physical operations. The operations to perform are chosen using gradient and Hessian information that can be efficiently extracted from experiments. We show that our algorithm does not suffer from the barren plateau problem, which is a significant issue in many variational quantum algorithms. We use numerical simulation to demonstrate that our method reliably produces either the true ground state or a physically meaningful metastable state in typical physical systems with such states.

quant-ph