arXiv ScienceSearch

arXiv subjects

Kai Huang

Publications and source records attributed to Kai Huang.

At least 19 recordsLinked to original sources

Alignment-Path Distillation from Non-streaming ASR-LLMs for Streaming Speech Recognition

In this paper, we propose an alignment-path distillation framework for streaming automatic speech recognition (ASR) with large language models (LLMs). Interleaved streaming ASR-LLMs use forced alignments (FA) from alignment models, such as those trained with connectionist temporal classification (CTC), to construct speech-text training sequences. However, alignments obtained from a separate acoustic model may be inconsistent with those learned by LLM-based ASR. This motivates us to transfer alignment information from a non-streaming ASR-LLM to improve streaming recognition. Specifically, we extract monotonic alignment paths from a non-streaming teacher's soft text-audio attention and use them to construct interleaved training sequences. The framework also includes logit and hidden-state distillation to learn from the teacher's output distributions and internal representations. Experimental results show that, without logit or hidden-state distillation, training with teacher-derived alignment paths achieves a 5.2% relative error rate reduction compared with training using forced alignments. When both models use logit and hidden-state distillation, teacher-derived alignments yield a 3.9% relative error rate reduction, with similar mean emission latency but higher flicker. The complete framework achieves a 16.6% relative error rate reduction compared with training using forced alignments without logit or hidden-state distillation.

eess.AS

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

Autonomous driving involves coupled decision-making and scene evolution under multi-mode uncertainty. To capture this coupling and uncertainty, we introduce MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair. Each hypothesis is initialized from a structured action prior and an independent future scene source, which are then co-evolved through a modality-aware diffusion Transformer. To support efficient multi-mode rollout, MM-Future compresses multi-view video into planning-oriented representations, dubbed MM-Tokens. Finally, a future-conditioned proposal scorer ranks trajectory candidates by shared history context and their paired predicted future. On NAVSIM navtest, MM-Future achieves 94.0 PDMS and 91.5 EPDMS, while attaining a 32.3 HD-Score in zero-shot closed-loop evaluation on HUGSIM. Ablations show consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.

cs.CV

Purely Electric-Field Control of Topological Magnetism in Two-Dimensional Magnets

Electrical control of topological magnetism is central to realizing energy-efficient topological spintronics. Yet most electric-field approaches modify the competing magnetic interactions through a nonselective rearrangement of low-energy electronic states, yielding coarse magnetic phase control that often requires external magnetic fields to stabilize topological quasiparticles, while few schemes solely based on electric fields are restricted to specific conducting materials. Here, we establish a general approach for purely electric-field control of topological magnetism, in which an applied electric field electrostatically dopes a selected orbital-angular-momentum-polarized band edge of a two-dimensional (2D) van der Waals (vdW) magnetic semiconductor via proximity to an adjacent nonmagnetic vdW metal. We show that the resulting electrostatic doping predominantly tunes magnetic anisotropy of the 2D magnet, while leaving exchange interaction and Dzyaloshinskii-Moriya interaction nearly unchanged, thereby reversibly driving the system through ferromagnetic, skyrmion, spiral, and bimeron phases. We demonstrate this mechanism for CrBr3/graphene and Cr2Ge2Te6/TaS2 vdW heterostructures hosting electron and hole pockets of different orbital characters. These results establish a broadly applicable strategy for purely electric-field control in topological spintronics.

cond-mat.mes-hall

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

A unified audio model must recognize and understand linguistic, paralinguistic, and environmental information while supporting speech synthesis and editing. A key challenge is representation: understanding favors compact features suited to long-context modeling, whereas speech generation requires reconstructible features that preserve fine-grained acoustic detail. We introduce FireRedAudio, a general-purpose audio language model with a shared 9B-parameter LLM. To the best of our knowledge, it is the first publicly disclosed unified audio-language model to provide separate continuous input representations for understanding and generation within a single trainable autoregressive LLM. Audio to be recognized or analyzed is processed by a dedicated Audio Encoder, while speech inputs for generation use a RedAE-based pathway. The LLM directly generates text or conditions a flow-matching DiT to produce continuous acoustic latents. Through progressive multitask training, FireRedAudio supports ASR and audio understanding, with the latter extending to recordings of up to one hour, as well as zero-shot TTS, Instruct TTS, and semantic and acoustic speech editing. Its structured organization of long-form audio achieves second-level timestamp accuracy. Across comprehensive evaluations, FireRedAudio achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero-shot TTS, leading instruction following in Instruct TTS, and substantial improvements over Ming-UniAudio-Edit in both semantic and acoustic speech editing. These results demonstrate the viability of decoupled continuous input representations for unifying audio understanding and continuous-latent speech generation in a model of moderate scale. Our code is available at https://github.com/FireRedTeam/FireRedAudio.

cs.CL

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.

cs.AI

Unveiling competitions between carrier recombination pathways in semiconductors via mechanical damping

The total rate of carrier recombination in semiconductors has conventionally been expressed using an additive model, r_total = Σr_i , which rules out the interactions between carrier recombination pathways. Here we challenge this paradigm by demonstrating pathway competitions using our newly developed light-induced mechanical absorption spectroscopy (LIMAS), which allows us to probe genuine recombination dynamics in semiconductors via mechanical damping. We show that the total recombination rate in zinc sulfide (ZnS), a model semiconductor material, follows a multiplicative weighting model, r_total \propto Πr_i ^(w_i) with Σw_i=1. Under both steady-state and switch-on illuminations, the weighting factors w_i for each recombination pathway-direct, trap-assisted, and sublinear-are dictated by the carrier generation mechanism: (i) interband transition favors direct recombination; (ii) single-defect level-mediated generation promotes trap-assisted recombination; (iii) generation involving multiple saturated defect levels gives rise to sublinear recombination. Upon light switch-off, localized state changes drive a dynamic evolution of w_i, altering pathway competitions. These findings reshape our fundamental understanding of carrier dynamics and provide a new strategy to optimize next-generation optoelectronic devices.

cond-mat.mtrl-sci

Beyond Textual Repository Exploration: Dual-Modal Structural Reasoning for Agentic Issue Resolution

Recent advances in agentic program repair have significantly improved issue resolution by enabling iterative repository exploration. However, existing approaches predominantly rely on sequential, text-based code navigation, which fundamentally limits their ability to reason over large-scale long-horizon repositories with complex and long-range dependencies. As issue-resolution agents traverse repositories through fragmented textual observations, structural information such as module organization, call relationships, and dependency chains must be repeatedly reconstructed across interaction steps, often leading to exploration drift and incomplete localization. We present DUALVIEW, a dual-modal structural scaffolding framework that brings visual reasoning into repository exploration for issue-resolution agents. DUALVIEW represents repository structure through four complementary graph views: Module Coupling Graph (MCG), Function Call Graph (FCG), Class Hierarchy Graph (CHG), and Program Dependence Graph (PDG), and exposes them through a queryable interface with visual and textual responses. Rather than reconstructing repository structure from a sequence of textual observations, agents can directly reason over persistent visual representations of code dependencies, enabling more effective exploration and understanding of long-horizon codebases. We evaluate DUALVIEW on SWE-bench Pro and Verified. Results show that DUALVIEW consistently improves issue-resolution performance across different agent architectures and model families. Further ablation studies demonstrate that the gains arise not only from textual structural information but also from visual externalization of repository dependencies, which better supports long-horizon repository exploration.

cs.SE

Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation

Language-conditioned robot manipulation is an emerging field aimed at enabling seamless communication and cooperation between humans and robotic agents by teaching robots to comprehend and execute instructions conveyed in natural language. This interdisciplinary area integrates scene understanding, language processing, and policy learning to bridge the gap between human instructions and robot actions. In this comprehensive survey, we systematically explore recent advancements in language-conditioned robot manipulation. We categorize existing methods based on the primary ways language is integrated into the robot system, namely language for state evaluation, language as a policy condition, language for cognitive planning and reasoning, and language in unified vision-language-action models. Specifically, we further analyze state-of-the-art techniques from five axes of action granularity, data and supervision regimes, system cost and latency, environments and evaluations, and task specification. Additionally, we highlight the key debates in the field. Finally, we discuss open challenges and future research directions, focusing on potentially enhancing generalization capabilities and addressing safety issues in language-conditioned robot manipulators.

cs.RO

Local-global principle for triangularizability and diagonalizability of matrices

Given a number field $k$ with the ring of integers $\mathcal{O}_k$ and a matrix $M\in \mathrm{M}_{n}(\mathcal{O}_k)$. We prove that if $\mathcal{O}_k$ is a principal ideal domain, the local-global principle for triangularizability and diagonalizability of $M$ holds. To explain the possible failures of the local-global principle, we prove that the stratified Brauer--Manin obstruction is the only obstruction to the local-global principle for triangularizability and diagonalizability of $M$ in some special cases.

math.NT

A11YRepair: Bridging Web Accessibility Barriers via Knowledge-Enhanced Divide-and-Conquer Repair

Web accessibility (A11Y), which ensures web content is perceivable and usable for users with disabilities, is a critical requirement for modern web applications. Yet existing tooling overwhelmingly focuses on detecting A11Y violations rather than repairing them. Automated program repair (APR) techniques appear promising for this setting, but our study shows that state-of-the-art APR systems perform poorly when applied to real-world A11Y violations. Unlike conventional sparse-bug scenarios, web A11Y issues often manifest as multiple structurally related violations per page, requiring coordinated edits across multiple files. Existing repair systems fail to manage this multi-fault scale, as they handle each bug individually without considering their relationships or incorporating domain rules such as the Web Content Accessibility Guidelines (WCAG). We propose A11YRepair, an LLM-based framework for web A11Y repair. A11YRepair introduces a divide-and-conquer workflow that first clusters violations requiring coordinated edits to reduce redundant localization, and then decomposes each cluster by root cause so the LLM can generate focused and consistent patches. The framework further incorporates WCAG-driven knowledge to strengthen domain awareness during both fault localization and patch synthesis. To support systematic evaluation, we construct A11YBench, a benchmark of 60 real-world web projects collected from GitHub. Experimental results show that A11YRepair achieves higher repair effectiveness and lower cost than state-of-the-art baselines, and ablation studies confirm the importance of its divide-and-conquer design and selective domain knowledge integration. Specifically, patches generated by A11YRepair have been merged into open-source projects from Google, Microsoft, Facebook, IBM, K8s, Docker, and Alibaba, demonstrating its practical value.

cs.SE

Blow-Up Constructions and Applications to Segre Classes and Multidegree Formulas

By comparing simultaneous multigraded blow-ups with suitable iterated blow-up construction, we establish a birational correspondence between the associated exceptional divisors. We further investigate blow-ups arising from rational maps to multiprojective spaces and derive intersection-theoretic formulas for the Chern classes of the resulting exceptional divisors and pullback of tautological line bundles. These constructions have two main applications. First, they yield a proof of the product formula for Segre classes without any pure-dimensionality hypothesis. Second, they lead to a degree formula for arbitrary closed subschemes of multiprojective spaces, recovering and extending the classical formula of van der Waerden.

math.AG

Re-acceleration of Energetic Ions via Small-Scale Reconnection in Magnetic Fusion Plasmas

We report the first observation on the EXL-50U spherical torus that energetic particles injected by neutral beam injection (NBI) can be stably accelerated to significantly higher energies - reaching up to 2.5 times the injection energy, occurring without significant large-scale magnetohydrodynamic (MHD) bursts. Simulations based on EXL-50U parameters indicate that small-scale magnetic reconnection, mediated by multiple magnetic islands, fails to accelerate bulk thermal ions but efficiently energizes seed fast ions. Unlike global MHD events, such small-scale reconnection is ubiquitous in magnetic confinement devices and does not degrade core confinement. This mechanism offers a novel and potentially universal channel for auxiliary ion heating in future fusion reactors.

physics.plasm-ph

Spin Dynamics in the van der Waals Ferromagnet CrTe2 Engineered by Niobium Doping

Understanding and controlling spin dynamics in two-dimensional (2D) van der Waals (vdW) ferromagnets is essential for their application in magnonics and hybrid quantum platforms. Here, we investigate the spin dynamics of the vdW ferromagnet 1T-CrTe_{2} and demonstrate their systematic tunability via niobium (Nb) substitution in Cr_{1-x}Nb_{x}Te_{2}(x=0-0.2). Ferromagnetic resonance (FMR) spectroscopy reveals that Nb doping enables wide-band tuning of the resonance frequency from 40 GHz down to the few-GHz regime, accompanied by a moderate increase in the Gilbert damping constant from ~0.066 to ~0.14, while preserving robust room-temperature ferromagnetism. Complementary magnetometry shows a concurrent reduction of the Curie temperature and saturation magnetization with increasing Nb content. Density functional theory calculations attribute the observed spin-dynamic trends to Nb-induced modifications of magnetic anisotropy and magnetic exchange interactions. Furthermore, CrTe_{2} flakes (~80 nm thick) exhibit lower resonance frequencies and damping than bulk crystals, consistent with thickness, surface/interface, and shape-dependent magnetic anisotropy. These results establish Nb-doped CrTe_{2} as a tunable vdW ferromagnet with controllable spin dynamics, extending its functionality from spintronics to broadband magnonics and quantum magnonics.

cond-mat.mtrl-sci

VisualNeo: Bridging the Gap between Visual Query Interfaces and Graph Query Engines

Visual Graph Query Interfaces (VQIs) empower non-programmers to query graph data by constructing visual queries intuitively. Devising efficient technologies in Graph Query Engines (GQEs) for interactive search and exploration has also been studied for years. However, these two vibrant scientific fields are traditionally independent of each other, causing a vast barrier for users who wish to explore the full-stack operations of graph querying. In this demonstration, we propose a novel VQI system built upon Neo4j called VisualNeo that facilities an efficient subgraph query in large graph databases. VisualNeo inherits several advanced features from recent advanced VQIs, which include the data-driven gui design and canned pattern generation. Additionally, it embodies a database manager module in order that users can connect to generic Neo4j databases. It performs query processing through the Neo4j driver and provides an aesthetic query result exploration.

cs.DB

KaLDeX: Kalman Filter based Linear Deformable Cross Attention for Retina Vessel Segmentation

Background and Objective: In the realm of ophthalmic imaging, accurate vascular segmentation is paramount for diagnosing and managing various eye diseases. Contemporary deep learning-based vascular segmentation models rival human accuracy but still face substantial challenges in accurately segmenting minuscule blood vessels in neural network applications. Due to the necessity of multiple downsampling operations in the CNN models, fine details from high-resolution images are inevitably lost. The objective of this study is to design a structure to capture the delicate and small blood vessels. Methods: To address these issues, we propose a novel network (KaLDeX) for vascular segmentation leveraging a Kalman filter based linear deformable cross attention (LDCA) module, integrated within a UNet++ framework. Our approach is based on two key components: Kalman filter (KF) based linear deformable convolution (LD) and cross-attention (CA) modules. The LD module is designed to adaptively adjust the focus on thin vessels that might be overlooked in standard convolution. The CA module improves the global understanding of vascular structures by aggregating the detailed features from the LD module with the high level features from the UNet++ architecture. Finally, we adopt a topological loss function based on persistent homology to constrain the topological continuity of the segmentation. Results: The proposed method is evaluated on retinal fundus image datasets (DRIVE, CHASE_BD1, and STARE) as well as the 3mm and 6mm of the OCTA-500 dataset, achieving an average accuracy (ACC) of 97.25%, 97.77%, 97.85%, 98.89%, and 98.21%, respectively. Conclusions: Empirical evidence shows that our method outperforms the current best models on different vessel segmentation datasets. Our source code is available at: https://github.com/AIEyeSystem/KalDeX.

eess.IV

EviRCOD: Evidence-Guided Probabilistic Decoding for Referring Camouflaged Object Detection

Referring Camouflaged Object Detection (Ref-COD) focuses on segmenting specific camouflaged targets in a query image using category-aligned references. Despite recent advances, existing methods struggle with reference-target semantic alignment, explicit uncertainty modeling, and robust boundary preservation. To address these issues, we propose EviRCOD, an integrated framework consisting of three core components: (1) a Reference-Guided Deformable Encoder (RGDE) that employs hierarchical reference-driven modulation and multi-scale deformable aggregation to inject semantic priors and align cross-scale representations; (2) an Uncertainty-Aware Evidential Decoder (UAED) that incorporates Dirichlet evidence estimation into hierarchical decoding to model uncertainty and propagate confidence across scales; and (3) a Boundary-Aware Refinement Module (BARM) that selectively enhances ambiguous boundaries by exploiting low-level edge cues and prediction confidence. Experiments on the Ref-COD benchmark demonstrate that EviRCOD achieves state-of-the-art detection performance while providing well-calibrated uncertainty estimates. Code is available at: https://github.com/blueecoffee/EviRCOD.

cs.CV

Remarks on Brauer-Manin obstruction for Weil restrictions

Given a finite extension $K/k$ of number fields and a smooth quasi-projective variety $X$ over $K$. If the abelianized fundamental group of $X$ is trivial, we prove that there is a natural identification between Brauer-Manin sets of $X$ and its Weil restriction $R_{K/k}X$. If $X$ is projective and $Pic(X\times_{K}\overline{k})$ is a torsion-free abelian group, we prove that there is a natural identification between algebraic Brauer-Manin sets of $X$ and $R_{K/k}X$.

math.NT