arXiv Science⌕ Search

arXiv subjects

Yu Meng

Publications and source records attributed to Yu Meng.

At least 37 records · Page 2Linked to original sources

Radiative decay of heavy-light mesons from lattice QCD

We present the first systematic study of the radiative decays of charmed mesons using $2+1$-flavor clover fermion gauge ensembles generated by the CLQCD collaboration. One of the ensembles is at the physical pion mass, and one has a fine lattice spacing $a\sim 0.05 ~\text{fm}$. We determine the coupling constants to be $g_{D^{\ast+} D^+ γ} = -0.204(22)$ GeV$^{-1}$, $g_{D^{\ast0} D^0 γ} = 1.73(37)$ GeV$^{-1}$, and $g_{D_s^{\ast+} D_s^+ γ} =-0.120(14)$ GeV$^{-1}$, respectively. Compared with previous studies, our results demonstrate significant improvements in precision. In particular, we carefully estimate the systematic uncertainty arising from matrix element fits, momentum transfer extrapolations, and chiral and continuum limit extrapolations, which are included in the reported total uncertainties. These couplings yield the following predictions of decay widths: $Γ_{D^{\ast+} \rightarrow D^+ γ} = 0.253(55)$ keV, $Γ_{D^{\ast0} \rightarrow D^0 γ} = 18.2(7.8)$ keV, and $Γ_{D_s^{\ast+}\rightarrow D_s^+ γ} = 0.094(22)$ keV. This work establishes first-principles results of the charmed meson radiative transitions and provides inputs for understanding the structure and properties of heavy-light mesons.

hep-lat↗

G-Zero: Self-Play for Open-Ended Generation from Zero Data

Self-evolving LLMs excel in verifiable domains but struggle in open-ended tasks, where reliance on proxy LLM judges introduces capability bottlenecks and reward hacking. To overcome this, we introduce G-Zero, a verifier-free, co-evolutionary framework for autonomous self-improvement. Our core innovation is Hint-$δ$, an intrinsic reward that quantifies the predictive shift between a Generator model's unassisted response and its response conditioned on a self-generated hint. Using this signal, a Proposer model is trained via GRPO to continuously target the Generator's blind spots by synthesizing challenging queries and informative hints. The Generator is concurrently optimized via DPO to internalize these hint-guided improvements. Theoretically, we prove a best-iterate suboptimality guarantee for an idealized standard-DPO version of G-Zero, provided that the Proposer induces sufficient exploration coverage and the data filteration keeps pseudo-label score noise low. By deriving supervision entirely from internal distributional dynamics, G-Zero bypasses the capability ceilings of external judges, providing a scalable, robust pathway for continuous LLM self-evolution across unverifiable domains.

cs.LG↗

CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning

Large Language Models (LLMs) have recently exhibited remarkable reasoning capabilities, largely enabled by supervised fine-tuning (SFT)- and reinforcement learning (RL)-based post-training on high-quality reasoning data. However, reproducing and extending these capabilities in open and scalable settings is hindered by three fundamental data-centric challenges: (1) the cold-start problem, arising from the lack of seed datasets with detailed, long Chain-of-Thought (CoT) trajectories needed to initialize reasoning policies; (2) limited domain coverage, as most existing open-source reasoning datasets are concentrated in mathematics, with limited coverage of broader scientific disciplines; and (3) the annotation bottleneck, where the difficulty of frontier-level reasoning tasks makes reliable human annotation prohibitively expensive or infeasible. To address these challenges, we introduce CHIMERA, a compact synthetic reasoning dataset comprising 9K samples for generalizable cross-domain reasoning. CHIMERA is constructed with three key properties: (1) it provides rich, long CoT reasoning trajectories synthesized by state-of-the-art reasoning models; (2) it has broad and structured coverage, spanning 8 major scientific disciplines and over 1K fine-grained topics organized via a model-generated hierarchical taxonomy; and (3) it employs a fully automated, scalable evaluation pipeline that uses strong reasoning models to cross-validate both problem validity and answer correctness. We use CHIMERA to post-train a 4B Qwen3 model. Despite the dataset's modest size, the resulting model achieves strong performance on a suite of challenging reasoning benchmarks, including GPQA-Diamond, AIME 24/25/26, HMMT 25, and Humanity's Last Exam, approaching or matching the reasoning performance of substantially larger models such as DeepSeek-R1 and Qwen3-235B.

cs.CL↗

Revisiting the Geminga halo at GeV energies with Fermi-LAT data

Nearby pulsars within $\sim1\,{\rm kpc}$ are considered to be possible sources of 10-500 GeV cosmic-ray positron excess measured by PAMELA and AMS-02. A TeV halo around Geminga is detected by HAWC, and the measurements of its surface brightness profile indicate a slow particle diffusion surrounding the source. This result challenges the pulsar interpretation of the positron excess. The observations at GeV energies provide direct information on the electron/positron density in the GeV nebula, which can offer more direct constraints on the origin of the positron excess. Two previous works have performed analyses on the GeV emission of the pulsar halo, but focused on the energy band above 8 GeV. In this work, we use a longer dataset from the Fermi Large Area Telescope (LAT) to re-analyze the GeV halo emission of Geminga, extending the analysis to cover the energy range of 1-1000 GeV. We find that the analysis in this wider energy range results in a low significance of the halo emission. This can be attributed to the Galactic interstellar emission model being unable to perfectly fit the background over this broader energy range, and due to the low measured halo flux at $<$ 10 GeV energies leading to a mismatch between the observation and model expectation. We also derive the spectral energy distribution of the tentative halo emission, which shows a very hard spectrum in the 1-10 GeV range.

astro-ph.HE↗

Form factors of the $D_s \to ϕ\ell ν_\ell$ semileptonic decay with (2+1)-flavor lattice QCD

We present a systematic lattice calculation of the vector and axial vector form factors $V$ and $A_i~(i=0,1,2)$ for the $D_s \to ϕ\ell ν_\ell$ semileptonic decay using (2+1)-flavor Wilson-clover fermion configurations generated by the CLQCD collaboration. Seven gauge ensembles with different lattice spacings, from $0.052~\text{fm}$ to $0.105~\text{fm}$, and different pion masses, from about $210~\text{MeV}$ to $320~\text{MeV}$ are utilized, enabling us to take both the continuum limit and physical pion mass extrapolation. The form factor ratios are obtained to be $r_V=1.614(19)$ and $r_2=0.741(31)$. Our results of form factors reach the precision of $1\%-4\%$, which greatly improves the previous lattice QCD results and obtains the most precise determination to date.

hep-lat↗

Do LLM Evaluators Prefer Themselves for a Reason?

Large language models (LLMs) are increasingly used as automatic evaluators in applications such as benchmarking, reward modeling, and self-refinement. Prior work highlights a potential self-preference bias where LLMs favor their own generated responses, a tendency often intensifying with model size and capability. This raises a critical question: Is self-preference harmful, or does it simply reflect the genuinely higher-quality outputs of stronger models? Answering this has been difficult as prior works mostly relied on subjective tasks that lack an objective ground truth, meaning that either preference can be reasonably justified. To address this ambiguity, we investigate self-preference using verifiable benchmarks (mathematical reasoning, factual knowledge, code generation) that allow objective ground-truth assessment. This enables us to distinguish harmful (favoring objectively worse responses) from legitimate (favoring genuinely superior ones) self-preference. Our large-scale experiments across 7 model families reveal three key findings: (1) While stronger models exhibit greater self-preference, much of this preference aligns with objectively superior performance, indicating stronger models prefer themselves mostly legitimately. (2) Harmful self-preference persists when evaluator models err as generators, and stronger models display more pronounced harmful self-preference bias when they do err. This suggests stronger models struggle more to recognize when they are wrong. (3) Inference-time scaling strategies, such as generating a long Chain-of-Thought before evaluation, effectively reduce the harmful self-preference. Additionally, we experiment with LMArena and show that our findings extend beyond verifiable benchmarks to real-world, subjective domains. These results provide a more nuanced understanding of LLM-based evaluation and practical insights for improving its reliability.

cs.CL↗

Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by integrating external knowledge sources to address their limitations in accessing up-to-date or specialized information. A natural strategy to increase the likelihood of retrieving relevant information is to expand the number of retrieved documents. However, involving more documents could introduce significant noise, as many documents may be irrelevant or misleading, thereby reducing the overall accuracy of the generated responses. To overcome the challenge associated with handling a larger number of documents, we propose WinnowRAG, a novel RAG framework designed to systematically filter out noisy documents while preserving valuable content -- a process we refer to as winnowing. WinnowRAG operates in two stages: In Stage I, we perform query-aware clustering to group similar documents and form distinct topic clusters. Each cluster is assigned to an LLM agent for generating a unique answer. In Stage II, we perform winnowing, wherein a critic LLM evaluates the outputs of multiple agents and iteratively separates useful documents from noisy ones. To retain useful documents when discarding agents, we propose two strategic merging techniques to ensure that only relevant knowledge is used for generating the final response. Crucially, WinnowRAG is model-agnostic and does not require any model fine-tuning, making it easily adaptable to various tasks. Extensive experiments on various realistic datasets demonstrate the effectiveness of WinnowRAG over state-of-the-art baselines.

cs.CL↗

DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment

Vision Language Foundation Models based on CLIP architecture for remote sensing primarily rely on short text captions, which often result in incomplete semantic representations. Although longer captions convey richer information, existing models struggle to process them effectively because of limited text-encoding capacity, and there remains a shortage of resources that align remote sensing images with both short text and long text captions. To address this gap, we introduce DGTRSD, a dual-granularity remote sensing image-text dataset, where each image is paired with both a short text caption and a long text description, providing a solid foundation for dual-granularity semantic modeling. Based on this, we further propose DGTRS-CLIP, a dual-granularity curriculum learning framework that combines short text and long text supervision to achieve dual-granularity semantic alignment. Extensive experiments on four typical zero-shot tasks: long text cross-modal retrieval, short text cross-modal retrieval, image classification, and semantic localization demonstrate that DGTRS-CLIP consistently outperforms existing methods across all tasks. The code has been open-sourced and is available at https://github.com/MitsuiChen14/DGTRS.

cs.CV↗

The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning

Reinforcement learning with verifiable rewards (RLVR) is a promising approach for training language models (LMs) on reasoning tasks that elicit emergent long chains of thought (CoTs). Unlike supervised learning, it updates the model using both correct and incorrect samples via policy gradients. To better understand its mechanism, we decompose the learning signal into reinforcing correct responses and penalizing incorrect ones, referred to as Positive and Negative Sample Reinforcement (PSR and NSR), respectively. We train Qwen2.5-Math-7B, Qwen3-4B and Llama-3.1-8B-Instruct on a mathematical reasoning dataset and uncover a surprising result: training with only negative samples -- without reinforcing correct responses -- can be highly effective: it consistently improves performance over the base model across the entire Pass@$k$ spectrum $k$ up to 256), often matching or surpassing PPO and GRPO. In contrast, reinforcing only correct responses improves Pass@1 but degrades performance at higher $k$, due to reduced diversity. These inference-scaling trends highlight that solely penalizing incorrect responses may contribute more to performance than previously recognized. Through gradient analysis, we show that NSR works by suppressing incorrect generations and redistributing probability mass toward other plausible candidates, guided by the model's prior beliefs. It refines the model's existing knowledge rather than introducing entirely new behaviors. Building on this insight, we propose a simple variant of the RL objective that upweights NSR, and show that it consistently improves overall Pass@$k$ performance on MATH, AIME 2025, and AMC23. Our code is available at https://github.com/TianHongZXY/RLVR-Decomposed.

cs.CL↗

Construction of general $N$-body lattice operators with arbitrary momenta

We present a systematic method for constructing lattice QCD operators for systems of an arbitrary number of particles with arbitrary momentum, spin, and internal quantum numbers. Explicit constructions are provided for one-, two-, three-, and four-hadron operators, covering all irreducible representations of the relevant lattice symmetry groups in both rest and moving frames. The construction procedure has been implemented in the open-source package \texttt{OpTion} (Operator construcTion), available at https://github.com/wittscien/OpTion. The paper and the package are designed to serve as a practical and extensible dictionary for future lattice QCD studies, as lattice calculations advance towards increasingly complex hadronic systems.

hep-lat↗

Study of the $D_s \to ϕ\ell ν_\ell$ semileptonic decay with (2+1)-flavor lattice QCD

We present a systematic lattice calculation of the $D_s \to ϕ\ell ν_\ell$ semileptonic decay using (2+1)-flavor Wilson-clover fermion configurations generated by the CLQCD collaboration. Seven gauge ensembles with different lattice spacings, from $0.052~\text{fm}$ to $0.105~\text{fm}$, and different pion masses, from about $210~\text{MeV}$ to $320~\text{MeV}$ are utilized, enabling us to take both the continuum limit and physical pion mass extrapolation. The ratios of form factors are obtained to be $r_V=1.614(19)$ and $r_2=0.741(31)$, with the precision improved by up to an order of magnitude compared to previous lattice studies. The branching fractions are given as $\mathcal{B}(D_s\toϕeν_e)=2.493(66)_{\text{stat}}(31)_{|V_{cs}|}\times 10^{-2}$ and $\mathcal{B}(D_s\toϕμν_μ)=2.351(60)_{\text{stat}}(29)_{|V_{cs}|}\times 10^{-2}$. The corresponding ratio of the branching fractions between the lepton $μ$ and $e$ is given by $\mathcal{R}_{μ/e}=0.9432(13)$, which provides essential theoretical support for future high-precision experimental tests of the lepton flavor universality. The CKM matrix element $|V_{cs}|$ is also extracted to be $0.952(12)_{\text{stat}}(23)_{\text{PDG}}$ and $0.945(12)_{\text{stat}}(24)_{\text{PDG}}$ for the $μ$ and $e$ channels, respectively.

hep-lat↗

Contextuality-based quantum key distribution with deterministic single-photon sources

Photons are central to quantum technologies, with photonic qubits offering a promising platform for quantum communication. Semiconductor quantum dots stand out for their ability to generate single photons on demand, a key capability for enabling long-distance quantum networks. In this work, we utilize high-purity single-photon sources based on self-assembled InAs(Ga)As quantum dots as quantum information carriers. We demonstrate that such on-demand single photons can generate quantum contextuality. This capability enables a novel protocol for semi-device-independent quantum key distribution over free-space channels. Crucially, our method does not require ideal or perfectly projective measurements, opening a new pathway for robust and practical quantum communication.

quant-ph↗

A quantum-coherent photon--emitter interface in the original telecom band

Quantum dots stand out as the most advanced and versatile light-matter interface available today. Their ability to deliver high-quality, high-rate, and pure photons has set benchmarks that far surpass other emitters. Yet, a critical frontier has remained elusive: achieving these exceptional capabilities at telecom wavelengths, bridging the gap to fiber-optic infrastructure and scalable silicon photonics. Overcoming this challenge demands high quality quantum materials and devices which, despite extensive efforts, have not been realized yet. Here, we demonstrate waveguide-integrated quantum dots and realize a fully quantum-coherent photon-emitter interface operating in the original telecommunication band. The quality is assessed by recording transform-limited linewidths only 8 % broader than the inverse lifetime and bright 41.7 MHz emission rate under 80 MHz $π$-pulse excitation, unlocking the full potential of quantum dots for scalable quantum networks.

physics.optics↗

Aligning Large Language Models via Fully Self-Synthetic Data

Traditional reinforcement learning from human feedback (RLHF) for large language models (LLMs) relies on expensive human-annotated datasets, while Reinforcement Learning from AI Feedback (RLAIF) also incurs significant costs, requiring the collection of diverse prompts and corresponding responses, often necessitating external reward models or proprietary models like GPT-4 to annotate preference pairs. In this work, we introduce Self-Alignment Optimization (SAO), a fully self-synthetic framework for LLM alignment, where all training data, including prompts (i.e., user queries), responses, and preferences, are generated by the model itself. Specifically, SAO first instructs the LLM to engage in persona role-play and generate diverse prompts and responses, which are then self-evaluated for preference optimization. Extensive experiments demonstrate that SAO effectively enhances the model's chat capabilities on standard benchmarks like AlpacaEval~2.0, while maintaining strong performance on downstream objective tasks (e.g., question-answering, math reasoning). Our work provides a practical solution for self-improvement in aligning LLMs, and the code for reproducing our results is available at: https://github.com/SJY8460/SAO.

cs.CL↗

Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents

Enabling large language models (LLMs) to utilize search tools offers a promising path to overcoming fundamental limitations such as knowledge cutoffs and hallucinations. Recent work has explored reinforcement learning (RL) for training search-augmented agents that interleave reasoning and retrieval before answering. These approaches usually rely on outcome-based rewards (e.g., exact match), implicitly assuming that optimizing for final answers will also yield effective intermediate search behaviors. Our analysis challenges this assumption: we uncover multiple systematic deficiencies in search that arise under outcome-only training and ultimately degrade final answer quality, including failure to invoke tools, invalid queries, and redundant searches. To address these shortcomings, we introduce DeSA (Decoupling Search-and-Answering), a simple two-stage training framework that explicitly separates search optimization from answer generation. In Stage 1, agents are trained to improve search effectiveness with retrieval recall-based rewards. In Stage 2, outcome rewards are employed to optimize final answer generation. Across seven QA benchmarks, DeSA-trained agents consistently improve search behaviors, delivering substantially higher search recall and answer accuracy than outcome-only baselines. Notably, DeSA outperforms single-stage training approaches that simultaneously optimize recall and outcome rewards, underscoring the necessity of explicitly decoupling the two objectives.

cs.AI↗

RS-OOD: A Vision-Language Augmented Framework for Out-of-Distribution Detection in Remote Sensing

Out-of-distribution (OOD) detection represents a critical challenge in remote sensing applications, where reliable identification of novel or anomalous patterns is essential for autonomous monitoring, disaster response, and environmental assessment. Despite remarkable progress in OOD detection for natural images, existing methods and benchmarks remain poorly suited to remote sensing imagery due to data scarcity, complex multi-scale scene structures, and pronounced distribution shifts. To this end, we propose RS-OOD, a novel framework that leverages remote sensing-specific vision-language modeling to enable robust few-shot OOD detection. Our approach introduces three key innovations: spatial feature enhancement that improved scene discrimination, a dual-prompt alignment mechanism that cross-verifies scene context against fine-grained semantics for spatial-semantic consistency, and a confidence-guided self-training loop that dynamically mines pseudo-labels to expand training data without manual annotation. RS-OOD consistently outperforms existing methods across multiple remote sensing benchmarks and enables efficient adaptation with minimal labeled data, demonstrating the critical value of spatial-semantic integration.

cs.CV↗

ProxyThinker: Test-Time Guidance through Small Visual Reasoners

Recent advancements in reinforcement learning with verifiable rewards have pushed the boundaries of the visual reasoning capabilities in large vision-language models (LVLMs). However, training LVLMs with reinforcement fine-tuning (RFT) is computationally expensive, posing a significant challenge to scaling model size. In this work, we propose ProxyThinker, an inference-time technique that enables large models to inherit the visual reasoning capabilities from small, slow-thinking visual reasoners without any training. By subtracting the output distributions of base models from those of RFT reasoners, ProxyThinker modifies the decoding dynamics and successfully elicits the slow-thinking reasoning demonstrated by the emerged sophisticated behaviors such as self-verification and self-correction. ProxyThinker consistently boosts performance on challenging visual benchmarks on spatial, mathematical, and multi-disciplinary reasoning, enabling untuned base models to compete with the performance of their full-scale RFT counterparts. Furthermore, our implementation efficiently coordinates multiple language models with parallelism techniques and achieves up to 38 $\times$ faster inference compared to previous decoding-time methods, paving the way for the practical deployment of ProxyThinker. Code is available at https://github.com/MrZilinXiao/ProxyThinker.

cs.CV↗

DragOSM: Extract Building Roofs and Footprints from Aerial Images by Aligning Historical Labels

Extracting polygonal roofs and footprints from remote sensing images is critical for large-scale urban analysis. Most existing methods rely on segmentation-based models that assume clear semantic boundaries of roofs, but these approaches struggle in off- nadir images, where the roof and footprint are significantly displaced, and facade pixels are fused with the roof boundary. With the increasing availability of open vector map annotations, e.g., OpenStreetMap, utilizing historical labels for off-nadir image annotation has become viable because remote sensing images are georeferenced once captured. However, these historical labels commonly suffer from significant positional discrepancies with new images and only have one annotation (roof or footprint), which fails to describe the correct structures of a building. To address these discrepancies, we first introduce a concept of an alignment token, which encodes the correction vector to guide the label correction. Based on this concept, we then propose Drag OpenStreetMap Labels (DragOSM), a novel model designed to align dislocated historical labels with roofs and footprints. Specifically, DragOSM formulates the label alignment as an interactive denoising process, modeling the positional discrepancy as a Gaussian distribution. During training, it learns to correct these errors by simulating misalignment with random Gaussian perturbations; during inference, it iteratively refines the positions of input labels. To validate our method, we further present a new dataset, Repairing Buildings in OSM (ReBO), comprising 179,265 buildings with both OpenStreetMap and manually corrected annotations across 5,473 images from 41 cities. Experimental results on ReBO demonstrate the effectiveness of DragOSM. Code, dataset, and trained models are publicly available at https://github.com/likaiucas/DragOSM.git.

cs.CV↗