arXiv Science⌕ Search

arXiv subjects

Xiaoqian Liu

Publications and source records attributed to Xiaoqian Liu.

At least 19 recordsLinked to original sources

When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning

Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-Language Models (LALMs), an unintuitive phenomenon exists: post-training models for structured reasoning trajectories results in marginal or even negative gains compared to post-training for direct answering. To investigate it, we introduce CAFE, an evaluation framework designed to precisely quantify audio reasoning errors. Evaluation results reveal LALMs struggle with perception during reasoning and encounter a critical bottleneck: reasoning performance suffers from audio perception decay as reasoning length extends. To address it, we propose MPAR$^2$, a paradigm that encourages dynamic perceptual reasoning and decomposes complex questions into perception-rich sub-problems. Leveraging reinforcement learning, MPAR$^2$ improves perception performance on CAFE from 31.74% to 63.51% and effectively mitigates perception decay, concurrently enhancing reasoning capabilities to achieve a significant 74.59% accuracy on the MMAU benchmark. Further analysis demonstrates that MPAR$^2$ reinforces LALMs to attend to audio input and dynamically adapts reasoning budget to match task complexity.

cs.SD↗

AudioKV: KV Cache Eviction in Efficient Large Audio Language Models

Large Audio-Language Models (LALMs) have set new benchmarks in speech processing, yet their deployment is hindered by the memory footprint of the Key-Value (KV) cache during long-context inference. While general KV cache compression techniques excel in LLMs, they often fail in the audio domain by overlooking the intrinsic temporal continuity of acoustic signals. To bridge this gap, we propose AudioKV, a novel framework that robustly prioritizes audio-critical attention heads through a hardware-friendly semantic-acoustic alignment mechanism. Specifically, we identify these modality-specialized heads by analyzing attention scores in ASR tasks and dynamically allocate KV cache budgets preferentially to them. Furthermore, we introduce Spectral Score Smoothing (SSS), an FFT-based global filtering strategy designed to suppress high-frequency noise and recover smooth global trends from importance scores, ensuring more balanced token selection with unprecedented precision. Extensive evaluations across multiple LALMs, including Qwen and Gemma series, demonstrate that AudioKV significantly outperforms baselines while enhancing computational efficiency. Notably, at a 40% compression ratio, AudioKV maintains near-full accuracy on Qwen3-Omni-30B with only a 0.45% drop, whereas traditional methods suffer from catastrophic performance degradation and repetition. Our code will be released after acceptance.

cs.SD↗

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.

cs.CV↗

BiHom-four-angle Hopf modules and BiHom-Yetter-Drinfel'd modules

In this paper, we introduce the notion of four-angle Hopf modules over a BiHom-Hopf algebra $H$. We show that the category ${}_{H}^{H}\mathfrak{M}_{H}^{H}$ of BiHom-four-angle Hopf modules admits a strict monoidal category structure with respect to either the BiHom-tensor product $\otimes_{H}$ or the BiHom-cotensor product $\square_{H}$ as its monoidal product. We prove that the category $\mathcal{YD}_{H}^{H}(m,n,p,q)$ of BiHom-$(m,n,p,q)$-Yetter-Drinfel'd modules with parameters $m,n,p,q\in\mathbb{Z}$ forms a strict braided monoidal category equipped with a new monoidal product. Furthermore, we establish monoidal equivalences between the monoidal categories $\mathcal{YD}_{H}^{H}(m,n,p,q)$ and ${}_{H}^{H}\mathfrak{M}_{H}^{H}$, where ${}_{H}^{H}\mathfrak{M}_{H}^{H}$ carries either $\otimes_{H}$ or $\square_{H}$ as its monoidal product. Finally, we construct braiding structures for the monoidal categories $({}_{H}^{H}\mathfrak{M}_{H}^{H},\otimes_{H})$ and $({}_{H}^{H}\mathfrak{M}_{H}^{H},\square_{H})$.

math.QA↗

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state. Using the integral relation between trajectories and velocity fields, we derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps. Under a multi-reference setup, single-state FlowCTS-OPD outperforms vanilla KL-based OPD with faster convergence. FlowCTS-OPD improves GenEval from 0.90 to 0.93, OCR from 0.90 to 0.92, and PickScore from 22.75 to 23.06, while outperforming a mixed-reward RL baseline across all target metrics. Further analysis reveals a clear temporal supervision mismatch in vanilla KL-based OPD arising from its auxiliary SDE transition kernels. Beyond on-policy setting,FlowCTS also consistently outperforms vanilla SFT , particularly on OCR, while increasing supervision steps exhibit a trade-off between richer trajectory information and greater optimization difficulty.

cs.LG↗

Towards Ultra-High Reliability in Wi-Fi 8: IEEE 802.11bn Core Mechanisms, mmWave Integration, and Performance Verification

As the demand for wireless connectivity expands from high-speed data transmission to high-reliability applications, such as the Industrial Internet of Things and immersive communications, traditional Wi-Fi technologies optimized primarily for peak throughput face new challenges in reliability and latency. Consequently, Wi-Fi 8 aims to achieve ultra-high reliability (UHR), improve communication performance in complex environments, and drive the transition from high-speed connectivity to highly reliable intelligent connectivity. This article provides a comprehensive review of the core mechanisms of Wi-Fi 8 and conducts system-level performance verification. We focus on the key enhancement mechanisms at the physical (PHY) and medium access control (MAC) layers in IEEE 802.11bn, elaborating on their theoretical principles and key application scenarios. Additionally, this paper explores the potential role of integrated millimeter-wave (IMMW) technology as a complementary solution for spectrum expansion in the Wi-Fi 8 era, analyzing its basic architecture and implementation. Finally, system-level simulations are performed to verify the effectiveness of the key technologies in IEEE 802.11bn in achieving their performance targets, while further validating the robust performance of the IMMW scheme under practical hardware impairments.

cs.NI↗

Electronic properties and topological aspects of graphene nanohelicoids

We introduce graphene nanohelicoids, geometric analogues of graphene nanoribbons, in which the honeycomb lattice is embedded on a helicoidal surface. Starting from the three-dimensional helical structure, we construct effective one-dimensional lattice models with band structures characterized by a momentum-shifted particle-hole relation $E_v(k)=-E_c(k+π)$ that reflects an anti-chiral symmetry arising from the nonsymmorphic symmetry. A systematic investigation of graphene nanohelicoids using the tight-binding approximation reveals a number of trends upon varying width and edge orientation, for instance, alternating transitions between semiconducting and metallic regimes. As the structure width varies, the band gap periodically closes and reopens, accompanied by an alternating Zak phase that switches between trivial and nontrivial. We derive an analytic tight-binding model and introduce a continuous deformation of the graphene nanohelicoids that explains the origin of width-dependent band inversion and alternating Zak phase.

cond-mat.mes-hall↗

DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast

Text-guided audio editing aims to modify the language-specified acoustic content while preserving edit-irrelevant source components. Existing training-free methods typically rely on inversion-based editing. While inversion-free editing is appealing as it decreases computational overhead and reconstruction errors, it remains largely unexplored for audio editing. The key challenge is to construct a source-to-target editing path through diffusion denoising dynamics. In this paper, we introduce DirectAudioEdit, the first attempt to develop a training-free and inversion-free method for audio editing. Experiments on music and event-level benchmarks across two backbones show that DirectAudioEdit reduces macro-averaged FAD and KL by 15.9% and 15.8% compared with DDPM inversion, while achieving up to 64.5% editing speedup.

cs.SD↗

An Interpretable and Scalable Framework for Evaluating Large Language Models

Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the inherent stochasticity of LLM outputs and the heterogeneity of benchmark items. Item Response Theory (IRT) offers a principled framework for modeling latent model abilities and item characteristics, but conventional methods are computationally expensive and numerically unstable, limiting large-scale implementations. To address these challenges, we propose an interpretable and scalable framework for LLM evaluation based on the majorization-minimization principle. Our approach reformulates the problem as a sequence of constrained matrix factorization subproblems, enabling stable and efficient parameter estimation with theoretical guarantees for identifiability and convergence. Experiments on synthetic and real-world datasets, including MATH-500 and six Open LLM Leaderboard benchmarks, demonstrate that our method achieves superior scalability and interpretability. It delivers orders-of-magnitude speedups over competing methods while maintaining comparable or even higher estimation accuracy. Our results align with established scaling laws and offer insights into item difficulty and discrimination, informing more principled benchmark design.

stat.ML↗

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models

Recent large audio language models (LALMs) demonstrate remarkable capabilities in processing extended multi-modal sequences, yet incur high inference costs. Token compression is an effective method that directly reduces redundant tokens in the sequence. Existing compression methods usually assume that all attention heads in LALMs contribute equally to various audio tasks and calculate token importance by averaging scores across all heads. However, our analysis demonstrates that attention heads exhibit distinct behaviors across diverse audio domains. We further reveal that only a sparse subset of attention heads actively responds to audio, with completely different performance when handling semantic and acoustic tasks. In light of this observation, we propose HeadRouter, a head-importance-aware token pruning method that perceives the varying importance of attention heads in different audio tasks to maximize the retention of crucial tokens. HeadRouter is training-free and can be applied to various LALMs. Extensive experiments on the AudioMarathon and MMAU-Pro benchmarks demonstrate that HeadRouter achieves state-of-the-art compression performance, exceeding the baseline model even when retaining 70% of the audio tokens and achieving 101.8% and 103.0% of the vanilla average on Qwen2.5-Omni-3B and Qwen2.5-Omni-7B, respectively.

cs.SD↗

Four-angle Hopf modules for Hom-Hopf algebras

In this paper, we introduce the notion of a four-angle Hopf module for a Hom-Hopf algebra $(H,β)$ and show that the category $\!^{H}_{H}\mathfrak{M}^{H}_{H}$ of four-angle Hopf modules is a monoidal category with either a Hom-tensor product $\otimes_{H}$ or a Hom-cotensor product $\Box_{H}$ as a monoidal product. We study the category $\mathcal{YD}^{H}_{H}$ of Yetter-Drinfel'd modules with bijective structure map can be organized as a braided monoidal category, in which we use a new monoidal structure. Finally, We prove an equivalence between the monoidal category $(~\!^{H}_{H}\mathfrak{M}^{H}_{H},\otimes_{H})$ or $(~\!^{H}_{H}\mathfrak{M}^{H}_{H},\Box_{H})$ of four-angle Hopf modules, and the monoidal category $\mathcal{YD}^{H}_{H}$ of Yetter-Drinfel'd modules, and furthermore, we give a braiding structure of the monoidal categorys $(~\!^{H}_{H}\mathfrak{M}^{H}_{H},\otimes_{H})$ (and $(~\!^{H}_{H}\mathfrak{M}^{H}_{H},\Box_{H})$).

math.RA↗

Transfer Learning for Robust Structured Regression with Bi-level Source Detection

High-dimensional data in modern applications, such as COVID-19 mortality, often span multiple domains. Leveraging auxiliary information from source domains to improve performance in a target domain motivates the use of transfer learning. However, a practical issue that has been overlooked is data contamination, which induces heterogeneity and can significantly degrade transfer learning performance. To address this challenge, we propose a novel approach that tackles transfer learning under data contamination within a structured regression setting. By employing the robust L2E criterion, we develop the TransL2E method that accounts for contamination in both target and source data while effectively transferring relevant information. Beyond robust estimation, TransL2E introduces a data-driven bi-level source detection mechanism, operating at both individual and cohort levels, which possesses multiple advantages over existing source detection approaches. Comprehensive simulation studies and a real data application demonstrate the superior performance of TransL2E in both robust estimation and structure recovery in the presence of data limitation and contamination.

stat.ME↗

On the Emotion Understanding of Synthesized Speech

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding results a plausible reward or evaluation metric for assessing emotional expressiveness in speech synthesis. In this work, we critically examine this assumption by systematically evaluating Speech Emotion Recognition (SER) on synthesized speech across datasets, discriminative and generative SER models, and diverse synthesis models. We find that current SER models can not generalize to synthesized speech, largely because speech token prediction during synthesis induces a representation mismatch between synthesized and human speech. Moreover, generative Speech Language Models (SLMs) tend to infer emotion from textual semantics while ignoring paralinguistic cues. Overall, our findings suggest that existing SER models often exploit non-robust shortcuts rather than capturing fundamental features, and paralinguistic understanding in SLMs remains challenging.

cs.CL↗

APR: Penalizing Structural Redundancy in Large Reasoning Models via Anchor-based Process Rewards

Test-Time Scaling (TTS) has significantly enhanced the capabilities of Large Reasoning Models (LRMs) but introduces a critical side-effect known as Overthinking. We conduct a preliminary study to rethink this phenomenon from a fine-grained perspective. We observe that LRMs frequently conduct repetitive self-verification without revision even after obtaining the final answer during the reasoning process. We formally define this specific position where the answer first stabilizes as the Reasoning Anchor. By analyzing pre- and post-anchor reasoning behaviors, we uncover the structural redundancy fixed in LRMs: the meaningless repetitive verification after deriving the first complete answer, which we term the Answer-Stable Tail (AST). Motivated by this observation, we propose Anchor-based Process Reward (APR), a structure-aware reward shaping method that localizes the reasoning anchor and penalizes exclusively the post-anchor AST. Leveraging the policy optimization algorithm suitable for length penalties, our APR models achieved the performance-efficiency Pareto frontier at 1.5B and 7B scales averaged across five mathematical reasoning datasets while requiring substantially fewer computational resources for RL training.

cs.CL↗

Electronic states at twist stacking faults in rhombohedral graphite

Flat bands in graphitic materials emerged as a platform for realizing tunable correlated physics. As a nodal-line semimetal, rhombohedral graphite features flat drumhead surface states in the vicinity of the Dirac points, which carry a nontrivial topological charge. We present a comprehensive study on rhombohedral graphite with twist stacking faults. Using both the continuum models and the realistic tight-binding models, we show that the twist angle between the graphene layers can tune the interface states at such stacking faults. The evolution of interface states originates from the interplay between the moiré periodicity and Zak phase topology, predicting the occurrence of nearly flat bands throughout the moiré Brillouin zone. We further investigate the disorder-induced layer polarization and tunable Chern number for flat band, and characterize the relationship between the disorder strength and Chern number in twisted rhombohedral graphite.

cond-mat.mes-hall↗

SageLM: A Multi-aspect and Explainable Large Language Model for Speech Judgement

Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose \texttt{SageLM}, an end-to-end, multi-aspect, and explainable speech LLM for comprehensive S2S LLMs evaluation. First, unlike cascaded approaches that disregard acoustic features, SageLM jointly assesses both semantic and acoustic dimensions. Second, it leverages rationale-based supervision to enhance explainability and guide model learning, achieving superior alignment with evaluation outcomes compared to rule-based reinforcement learning methods. Third, we introduce \textit{SpeechFeedback}, a synthetic preference dataset, and employ a two-stage training paradigm to mitigate the scarcity of speech preference data. Trained on both semantic and acoustic dimensions, SageLM achieves an 82.79\% agreement rate with human evaluators, outperforming cascaded and SLM-based baselines by at least 7.42\% and 26.20\%, respectively.

cs.CL↗

MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction

Current direct speech-to-speech translation methods predominantly employ speech tokens as intermediate representations. However, a single speech token is not dense in semantics, so we generally need multiple tokens to express a complete semantic unit. To address this limitation, we introduce multi-token prediction (MTP) loss into speech-to-unit translation (S2UT) models, enabling models to predict multiple subsequent tokens at each position, thereby capturing more complete semantics and enhancing information density per position. Initial MTP implementations apply the loss at the final layer, which improves output representation but initiates information enrichment too late. We hypothesize that advancing the information enrichment process to intermediate layers can achieve earlier and more effective enhancement of hidden representation. Consequently, we propose MTP-S2UT loss, applying MTP loss to hidden representation where CTC loss is computed. Experiments demonstrate that all MTP loss variants consistently improve the quality of S2UT translation, with MTP-S2UT achieving the best performance.

cs.CL↗

AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs

Processing long-form audio is a major challenge for Large Audio Language models (LALMs). These models struggle with the quadratic cost of attention ($O(N^2)$) and with modeling long-range temporal dependencies. Existing audio benchmarks are built mostly from short clips and do not evaluate models in realistic long context settings. To address this gap, we introduce AudioMarathon, a benchmark designed to evaluate both understanding and inference efficiency on long-form audio. AudioMarathon provides a diverse set of tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0 seconds, which correspond to encoded sequences of 2,250 to 7,500 audio tokens, respectively, full domain coverage across speech, sound, and music, and complex reasoning that requires multi-hop inference. We evaluate state-of-the-art LALMs and observe clear performance drops as audio length grows. We also study acceleration techniques and analyze the trade-offs of token pruning and KV cache eviction. The results show large gaps across current LALMs and highlight the need for better temporal reasoning and memory-efficient architectures. We believe AudioMarathon will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.

cs.SD↗