arXiv ScienceSearch

arXiv subjects

Ying Zhu

Publications and source records attributed to Ying Zhu.

At least 19 recordsLinked to original sources

Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Group Relative Policy Optimization (GRPO). In contrast to Reinforcement Learning with Verifiable Rewards (RLVR) methods, our proposal does not involve a Kullback Leibler (KL) regularization or critic; the trainable policy and the reference anchor are coupled through the logit averaging structure to leverage the reasoning expertise of the trainable policy while maintaining the formatting advantage of SFT. Our method is evaluated on MATH, cn-k12, and MMLU, and the results show a higher accuracy or at least comparable accuracy relative to the canonical KL-regularized GRPO.

cs.LG

Phase calibration of quantum oscillations in the magnetostrictive coefficient using the topological antiferromagnet YbMnBi$_2$

The Berry phase accumulated along a cyclotron orbit encodes important information about electronic band topology and is commonly inferred from the phase of quantum oscillations. Measurements of the ac magnetostrictive coefficient have recently emerged as a sensitive thermodynamic probe of quantum oscillations, but the phase offset has not been experimentally calibrated. Here, using the topological antiferromagnet YbMnBi$_2$, we calibrate this offset by directly comparing quantum oscillations in magnetization with those in the ac magnetostrictive coefficient. Measurements of both responses on the same single crystal reveal a single fundamental frequency of approximately 160 T in fields up to 14 T, enabling a direct phase comparison free from ambiguities associated with multiple frequencies. We observe an approximately $π/2$ relative phase shift between the two oscillatory responses, consistent with the Maxwell relation linking the magnetostrictive coefficient to the stress derivative of magnetization. Our results establish the appropriate phase needed to extract cyclotron-orbit phase information from quantum oscillations in the ac magnetostrictive coefficient.

cond-mat.str-el

Approximating invariant functions with the sorting trick is theoretically justified

Many machine learning models leverage group invariance which is enjoyed with a wide-range of applications. For exploiting an invariance structure, one common approach is known as \emph{frame averaging}. One popular example of frame averaging is the \emph{group averaging}, where the entire group is used to symmetrize a function. Another example is the \emph{canonicalization}, where a frame at each point consists of a single group element which transforms the point to its orbit representative, for example, sorting. Compared to group averaging, canonicalization is more efficient computationally. However, it results in non-differentiability or discontinuity of the canonicalized function. As a result, the theoretical performance of canonicalization has not been given much attention. In this work, we establish an approximation theory for canonicalization. Specifically, we bound the point-wise and $L^2(\mathbb{P})$ approximation errors as well as the eigenvalue decay rates associated with a canonicalization trick applied to reproducing kernels. We discuss two key insights from our theoretical analyses and why they point to an interesting future research direction on how one can choose a design to fully leverage canonicalization in practice.

cs.LG

FourierSampler: Unlocking Non-Autoregressive Potential in Diffusion Language Models via Frequency-Guided Generation

Despite the non-autoregressive potential of diffusion language models (dLLMs), existing decoding strategies demonstrate positional bias, failing to fully unlock the potential of arbitrary generation. In this work, we delve into the inherent spectral characteristics of dLLMs and present the first frequency-domain analysis showing that low-frequency components in hidden states primarily encode global structural information and long-range dependencies, while high-frequency components are responsible for characterizing local details. Based on this observation, we propose FourierSampler, which leverages a frequency-domain sliding window mechanism to dynamically guide the model to achieve a "structure-to-detail" generation. FourierSampler outperforms other inference enhancement strategies on LLADA and SDAR, achieving relative improvements of 20.4% on LLaDA1.5-8B and 16.0% on LLaDA-8B-Instruct. It notably surpasses similarly sized autoregressive models like Llama3.1-8B-Instruct.

cs.CL

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1

cs.SD

Complex Temperature-dependent Thermal Conductivity in a Sawtooth Chain Magnet Fe$_\mathrm{2}$SiSe$_\mathrm{4}$

Geometrically frustrated magnets provide an ideal platform for exploring the interplay between lattice geometry and spin degrees of freedom. Here, we investigate the interactions between lattice and spin via thermal-transport measurements on the triangular sawtooth-lattice olivine magnet Fe$_\mathrm{2}$SiSe$_\mathrm{4}$, which exhibits successive magnetic transitions at $T_1 = 110$ K (antiferromagnetic) and $T_2 = 50$ K (ferrimagnetic). Although phonons dominate the thermal conductivity, its temperature dependence displays a pronounced double-peak structure arising from spin-phonon coupling. In the intermediate temperature range between $T_1$ and $T_2$ , resonant scattering of phonons by magnetic excitations around 5 meV produces a broad maximum around 60 K. Below $T_2$, the resonant spin-phonon scattering is strongly suppressed, leading to a rapid increase in thermal conductivity upon cooling and a pronounced low-temperature peak near 11 K, characteristic of heat transport governed by conventional phonon scattering mechanisms. Notably, this low-temperature peak is enhanced by a factor of $\sim 5$ compared to the broad maximum at higher temperatures. These results demonstrate the strong sensitivity of thermal transport to spin-lattice interactions and highlight spin-phonon scattering as an effective mechanism for tailoring thermal conductivity in geometrically frustrated magnets.

cond-mat.str-el

Turbulence and far-from-equilibrium equation of state of Bogoliubov waves in Bose-Einstein Condensates

Bogoliubov waves are fundamental excitations of Bose-Einstein Condensates (BECs). They emerge from a perturbed ground state and interact nonlinearly, triggering turbulent cascades. Here, we study turbulent BECs theoretically and numerically using the 3D Gross-Pitaevskii model and its associated wave-kinetic equations. We derive a new Kolmogorov-like stationary spectrum for short Bogoliubov waves and find a complete analytical expression for the spectrum in the long-wave acoustic regime. We then use our predictions to explain the BEC equation of state reported by Dogra et al. (Nature 620, 521, 2023), and to suggest new experimental settings.

cond-mat.quant-gas

Less is More: Lightweight Prompt Compression for Question Answering Applications on Edge Devices

In agent-driven question answering (QA) applications, retrieval-augmented generation (RAG) is commonly introduced to enhance the response accuracy of large language models (LLMs) by providing additional context. Due to the inherent noise in retrieval results and the coarse granularity of document-level retrieval, the retrieved context often contains substantial redundant information. In this setting, the agent prompt, consisting of the user query and the associated retrieved context, leads to unnecessary computational overhead during LLM inference. Existing prompt compression methods typically rely on auxiliary small language models (SLMs) to estimate context importance. However, such approaches introduce significant memory and computational overhead, which limits their deployment on resource-constrained edge devices. In this paper, we propose CORE, a two-stage sentence-level prompt compression method that eliminates the need for SLMs. In the first stage, CORE constructs an answer set via named entity recognition (NER) and a clue set via semantic matching. In the second stage, CORE refines the clue set using an orthogonal residual retrieval strategy and designs a spatial proximity-based metric to filter the answer set. The two sets are then combined to form the final compressed context. We implement CORE on an NVIDIA Jetson AGX Orin edge device and a Huawei Nova smartphone. Experimental results demonstrate that within a 2000-token budget, CORE improves accuracy by at least 30.19% compared to state-of-the-art baselines, while reducing memory usage by at least 50.47% and achieving at least 1.94 times speedup on the edge device. Moreover, compared to the state-of-the-art LLMLingua2 method, CORE achieves a substantial energy reduction of 95.74% on the smartphone, highlighting its practicality and generalizability for mobile deployments.

cs.CL

Collective Resonance of Superconducting/Normal Domain Walls in the Intermediate State of type-I superconductor

The dynamics of phase boundaries, such as superconducting/normal (S/N) interfaces in type-I superconductors, are typically obscured in conventional magnetic measurements, which are dominated by surface barriers and over-damped flux processes. Here, we employ ac magnetostriction as a sensitive probe to reveal the distinct bulk dynamics of these domain walls in the intermediate state of lead. In contrast to the Debye-type relaxation observed in magnetic susceptibility, we discover a pronounced quasiresonant response characterized by a sign reversal of the imaginary component and a non-monotonic evolution of the real part with frequency. We attribute this behavior to the collective oscillations of S/N interfaces driven by eddy currents generated within the normal domains. This work uncovers a fundamental dynamical channel in superconducting modulated phases and establishes ac magnetostrictive coefficient as a powerful tool for probing hidden interface physics.

cond-mat.supr-con

Strong and weak wave turbulence regimes in Bose-Einstein condensates

When a turbulent Bose-Einstein condensate is driven out-of-equilibrium at a scale much smaller than the system size, nonlinear wave interactions transfer particles towards large scales in an inverse cascade process. In this work, we study numerically wave turbulence in a three-dimensional Bose-Einstein condensate in forced and dissipated inverse cascade settings. We observe that when the forcing rate increases, thereby increasing the particle flux, the turbulence spectrum gradually transitions from the weak-wave Kolmogorov-Zakharov cascade to a critical balance state characterized by a range of scales with balanced linear and nonlinear dynamic timescales. Further forcing increases lead to a coherent condensate component superimposed with Bogoliubov-type acoustic turbulence. The role of vortices in such a strongly forced state is marginal, which makes this new state very different from the strongly turbulent state composed of a tangle of quantized vortex lines. We then use our predictions and numerical data to formulate a new out-of-equilibrium equation of state for the 3D inverse cascade.

cond-mat.quant-gas

Evaluating Prompting Strategies for Chart Question Answering with Large Language Models

Prompting strategies affect LLM reasoning performance, but their role in chart-based QA remains underexplored. We present a systematic evaluation of four widely used prompting paradigms (Zero-Shot, Few-Shot, Zero-Shot Chain-of-Thought, and Few-Shot Chain-of-Thought) across GPT-3.5, GPT-4, and GPT-4o on the ChartQA dataset. Our framework operates exclusively on structured chart data, isolating prompt structure as the only experimental variable, and evaluates performance using two metrics: Accuracy and Exact Match. Results from 1,200 diverse ChartQA samples show that Few-Shot Chain-of-Thought prompting consistently yields the highest accuracy (up to 78.2\%), particularly on reasoning-intensive questions, while Few-Shot prompting improves format adherence. Zero-Shot performs well only with high-capacity models on simpler tasks. These findings provide actionable guidance for selecting prompting strategies in structured data reasoning tasks, with implications for both efficiency and accuracy in real-world applications.

cs.CL

Universal Behavior on the Relaxation Dynamics of Far-From-Equilibrium Quantum Fluids

Investigating the initial conditions that lead many-body quantum systems to an out-of-equilibrium state is fundamental for understanding their thermalization dynamics. In this work we observe the relaxation for two regimes of excitation that can drive the turbulent Bose-Einstein condensate into two distinct final states, and are defined by the amount of energy injected into the system. The subcritical regime is characterized by a lower injection of energy, which can lead to an inverse particle cascade and, consequently, to the BEC mode repopulation during the relaxation process. The supercritical regime is marked by a higher energy injection, that may lead to the BEC dissolution and a final thermal state. In both cases we observe relaxation stages that exhibit the same key features: a direct cascade, a non-thermal fixed point with the same exponents, a prethermalization region and, finally, the thermalization of the system. In the final thermalization stage, universal scaling is observed for both regimes, even though their final states are completely different. By analyzing the coherence length of our turbulent cloud, we clearly visualize the recovery and the loss of the coherence for the subcritical and supercritical regimes after relaxation. These results indicate that the evolution of turbulence occurs independent of its initial conditions and of the final state achieved.

cond-mat.quant-gas

Predicting Tennis Serve directions with Machine Learning

Serves, especially first serves, are very important in professional tennis. Servers choose their serve directions strategically to maximize their winning chances while trying to be unpredictable. On the other hand, returners try to predict serve directions to make good returns. The mind game between servers and returners is an important part of decision-making in professional tennis matches. To help understand the players' serve decisions, we have developed a machine learning method for predicting professional tennis players' first serve directions. Through feature engineering, our method achieves an average prediction accuracy of around 49\% for male players and 44\% for female players. Our analysis provides some evidence that top professional players use a mixed-strategy model in serving decisions and that fatigue might be a factor in choosing serve directions. Our analysis also suggests that contextual information is perhaps more important for returners' anticipatory reactions than previously thought.

cs.LG

MOVA: Towards Scalable and Synchronized Video-Audio Generation

Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and degrade overall quality. While systems such as Veo 3 and Sora 2 emphasize the value of simultaneous generation, joint multimodal modeling introduces unique challenges in architecture, data, and training. Moreover, the closed-source nature of existing systems limits progress in the field. In this work, we introduce MOVA (MOSS Video and Audio), an open-source model capable of generating high-quality, synchronized audio-visual content, including realistic lip-synced speech, environment-aware sound effects, and content-aligned music. MOVA employs a Mixture-of-Experts (MoE) architecture, with a total of 32B parameters, of which 18B are active during inference. It supports IT2VA (Image-Text to Video-Audio) generation task. By releasing the model weights and code, we aim to advance research and foster a vibrant community of creators. The released codebase features comprehensive support for efficient inference, LoRA fine-tuning, and prompt enhancement.

cs.CV

Controllable Dance Generation with Style-Guided Motion Diffusion

Dance plays an important role as an artistic form and expression in human culture, yet automatically generating dance sequences is a significant yet challenging endeavor. Existing approaches often neglect the critical aspect of controllability in dance generation. Additionally, they inadequately model the nuanced impact of music styles, resulting in dances that lack alignment with the expressive characteristics inherent in the conditioned music. To address this gap, we propose Style-Guided Motion Diffusion (SGMD), which integrates the Transformer-based architecture with a Style Modulation module. By incorporating music features with user-provided style prompts, the SGMD ensures that the generated dances not only match the musical content but also reflect the desired stylistic characteristics. To enable flexible control over the generated dances, we introduce a spatial-temporal masking mechanism. As controllable dance generation has not been fully studied, we construct corresponding experimental setups and benchmarks for tasks such as trajectory-based dance generation, dance in-betweening, and dance inpainting. Extensive experiments demonstrate that our approach can generate realistic and stylistically consistent dances, while also empowering users to create dances tailored to diverse artistic and practical needs. Code is available on Github: https://github.com/mucunzhuzhu/DGSDP

cs.CV

Rotating fluorescent nanodiamond assemblies with focused Laguerre-Gaussian beams

Optical tweezers which utilize structured light fields enable the rotation of trapped nanoparticles through the transfer of orbital angular momentum (OAM) from holographically generated Laguerre-Gaussian (LG) modes. In this research we use OAM transfer to demonstrate controlled rotation of bright fluorescent nanodiamond clusters assembled in a focused higher-order LG beam. We find that the assemblies can be effectively rotated in a two-dimensional optical trap with orbital frequencies of up to 5 Hz. We use video tracking to explore the Brownian dynamics of such a trapping arrangement and look at the impact of orientation stability on measurements of optically detected magnetic resonance (ODMR) with an applied weak external magnetic field. By collecting ODMR spectra at multiple points along the orbit, we show that the constrained two-dimensional motion can provide additional insights for vector magnetic field reconstruction.

physics.optics

DiRL: An Efficient Post-Training Framework for Diffusion Language Models

Diffusion Language Models (dLLMs) have emerged as promising alternatives to Auto-Regressive (AR) models. While recent efforts have validated their pre-training potential and accelerated inference speeds, the post-training landscape for dLLMs remains underdeveloped. Existing methods suffer from computational inefficiency and objective mismatches between training and inference, severely limiting performance on complex reasoning tasks such as mathematics. To address this, we introduce DiRL, an efficient post-training framework that tightly integrates FlexAttention-accelerated blockwise training with LMDeploy-optimized inference. This architecture enables a streamlined online model update loop, facilitating efficient two-stage post-training (Supervised Fine-Tuning followed by Reinforcement Learning). Building on this framework, we propose DiPO, the first unbiased Group Relative Policy Optimization (GRPO) implementation tailored for dLLMs. We validate our approach by training DiRL-8B-Instruct on high-quality math data. Our model achieves state-of-the-art math performance among dLLMs and surpasses comparable models in the Qwen2.5 series on several benchmarks.

cs.LG

Redundancy as a Structural Information Principle for Learning and Generalization

We present a theoretical framework that extends classical information theory to finite and structured systems by redefining redundancy as a fundamental property of information organization rather than inefficiency. In this framework, redundancy is expressed as a general family of informational divergences that unifies multiple classical measures, such as mutual information, chi-squared dependence, and spectral redundancy, under a single geometric principle. This reveals that these traditional quantities are not isolated heuristics but projections of a shared redundancy geometry. The theory further predicts that redundancy is bounded both above and below, giving rise to an optimal equilibrium that balances over-compression (loss of structure) and over-coupling (collapse). While classical communication theory favors minimal redundancy for transmission efficiency, finite and structured systems, such as those underlying real-world learning, achieve maximal stability and generalization near this equilibrium. Experiments with masked autoencoders are used to illustrate and verify this principle: the model exhibits a stable redundancy level where generalization peaks. Together, these results establish redundancy as a measurable and tunable quantity that bridges the asymptotic world of communication and the finite world of learning.

cs.LG