arXiv ScienceSearch

arXiv subjects

Zhipeng Zhang

Publications and source records attributed to Zhipeng Zhang.

At least 19 recordsLinked to original sources

Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool

Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost operators, but the dominant attention cost scales with the sum of squared sequence lengths. Thus, equally sized packed sequences drawn from a long-tailed corpus can carry substantially different attention workloads, creating data-parallel stragglers and pipeline bubbles. Existing approaches either balance at the granularity of sequences or microbatches, where an outlier can dominate an assignment, or disaggregate attention over a global worker pool whose communication domain grows with the data-parallel (DP) degree. We present Libra, which operationalizes the law of large numbers (LLN) as a scaling principle for load balancing: the attention-balancing pool need not grow with the DP degree. Libra groups packed sequences and their CP groups into fixed-size sequence pools. As DP scales out, Libra adds pools rather than enlarging each one, bounding every attention exchange. Variance-Reduced Sequence Placement makes this effective for finite, long-tailed workloads by co-locating sequences with complementary attention workloads to reduce residual inter-pool skew. Within each pool, Tiled Attention Pooling dispatches sequence-head SH-Tiles across GPUs, while a pipelined runtime overlaps tile exchange with attention. Libra exposes a drop-in context-parallel attention operator and a pluggable data sampler, requiring no changes to model layers, optimizers, or pipeline schedules. On three production Qwen3 models (8B, 30B, 235B) and 256K- and 1M-token production workloads, Libra improves end-to-end training throughput over the strongest evaluated baseline (WLB-LLM) by 44% on average and up to 68% at 256 GPUs. Libra has run for hundreds of thousands of GPU-hours in production on jobs spanning 32K to 1M tokens.

cs.DC

Low Mach number limit of the compressible Euler--Vlasov--Fokker--Planck system in the whole space

Although there are many important contributions on compressible and incompressible fluid-particle interaction models respectively, how to connect the two-type fluid-particle models via the low Mach number limit remains a challenging open problem. In this paper, we resolve it for the compressible isentropic fluid-particle model (Euler--Vlasov--Fokker--Planck (Euler--VFP) system) in the whole space $\mathbb{R}^3$. First, we establish the global-in-time {\it a priori estimates} of strong solutions that are uniform with respect to the Mach number $\varepsilon$ near the global Maxwellian. The proof relies on a refined energy method that combines the relaxation structure $b^\varepsilon-u^\varepsilon$ induced by the fluid-particle interaction and the symmetrized acoustic structure of the compressible Euler part in the model. Under the assumption of well-prepared initial data, we derive a {\it global-in-time} uniform error estimate in the $H^2$ framework between the solution of the compressible Euler--VFP system and that of the limiting incompressible Euler--VFP system. A key point is to introduce the corrected acoustic variable $ q^\varepsilon-\varepsilon [P'(1)]^{-1}π$, which captures the pressure corrector in the low Mach number limit. This also allows us to exploit the exact cancellation of the singular acoustic terms and to close the {\it global-in-time} error estimate. The damping term $b^\varepsilon-u^\varepsilon$, which is absent in the pure Euler equations, plays an essential role in recovering the relative velocity dissipation and in controlling the coupled fluid-particle dynamics. As a consequence, we prove the low Mach number limit of the compressible Euler--VFP system with the convergence rate $\mathcal O(\varepsilon)$ in the time-continuous $H^2$ topology.

math.AP

Low Mach number limit for the Navier--Stokes--Korteweg equations with a stationary force

In this paper, we investigate the low Mach number limit for the three-dimensional compressible Navier--Stokes--Korteweg equations in the whole space under a small stationary external force. We first construct a family of small stationary solutions uniformly with respect to the Mach number $ε$ and prove that both the stationary density fluctuation and the compressible component of the stationary velocity are of order $ε^2$. For ill-prepared non-stationary perturbations around these stationary solutions, we establish the existence and uniqueness of global strong solution by combining uniform high-order energy estimates with a low-frequency Besov estimate and a Kawashima-type compensating functional. The main difficulty is that Korteweg tensor not only changes the elliptic structure of the stationary problem, but also modifies the dispersive mechanism of the acoustic modes. In Korteweg-symmetric variables, the associated spectral projections are uniformly bounded zero-order Fourier multipliers, while the acoustic-capillary phase is wave-like at low frequencies and Schrödinger-like at high frequencies. Since the source terms generated by the stationary coefficients are generally not integrable in time, we decompose the Duhamel source according to its time-integrability and frequency behavior. Dyadic dispersive estimates, high-frequency damping estimates, and maximal regularity for the heat equation yield the global-in-time convergence rate $ε^{\min\{1/r,\,1/2-1/p\}}$ in the mixed Besov norms $L^r(0,\infty;\dot B^s_{p,1})$. As a consequence, Besov embeddings also yield quantitative convergence in the mixed Lebesgue norms $L^r(0,\infty;L^p)$.

math.AP

Stable but Wrong: When Learning Stabilizes Away from the Truth

Stable training is often treated as evidence that learning is succeeding, but stability characterizes optimization behavior rather than correctness relative to an external objective. We study what happens when the signal being optimized remains persistently biased. We define Stable but Wrong (SBW) as a learning state in which the learning process remains stable under a task-appropriate operational criterion while the learned outcome remains systematically displaced from an independently defined objective. A minimal strongly convex model shows that a persistent bias in the update direction can shift the unique convergence point away from the true optimum. Controlled experiments in reinforcement learning, supervised learning, and continual fine-tuning of a large language model reveal a recurring separation between apparently normal optimization and correctness under static and feedback-coupled biases. Recovery-stage clean-data access and exploration interventions further show that subsequent trajectories can remain modifiable, although the interventions differ in protocol and do not imply a shared mechanism. The results expose a basic limit of optimization stability as a reliability signal in persistent and feedback-coupled learning systems.

cs.LG

Compressible Navier--Stokes equations with a potential force: global well-posedness and optimal time-decay rates for arbitrarily large $L^2$ initial data

We study the Cauchy problem for the three-dimensional barotropic compressible Navier--Stokes equations with a time-independent potential force near a spatially nonconstant stationary state. The potential is controlled in unweighted homogeneous Besov spaces; in particular, no polynomial spatial-weight condition involving $(1+|x|)^j\nabla^jϕ$ is imposed. For initial data relative to the stationary state that are sufficiently small in $\dot H^{\frac12-δ}\cap\dot H^3$, we establish the existence and uniqueness of a global strong solution in $H^3$, while allowing the initial $L^2$ norm to be arbitrarily large. If the initial data are bounded in $\dot B^s_{2,\infty}$ for $s\in[-\frac32,-1)$, then the solution and its first spatial derivative decay at the optimal rates $(1+t)^{-\frac{k-s}{2}}$ with $k=0$ and $1$, respectively. The analysis relies on refined homogeneous energy estimates and a frequency-localized description for the dissipative and asymptotic structures of the system.

math.AP

Qwen-CUA: Native Computer Use for (almost) Everything

Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. It observes only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, or task-specific APIs. Its scaffold maintains up to 20 active screenshots and folds older visual history in fixed-size blocks to retain recent evidence while preserving reusable prompt prefixes. For training, we build a cloud rollout fleet with access to nearly 100,000 vCPUs and tens of thousands of concurrent environments, construct approximately 40,000 verifiable tasks, and collect personalized long-horizon workflows across everyday and professional software. We optimize complete trajectories with verifiable rewards and trajectory slicing, while iterative training runs refresh supervised data and recalibrate reinforcement-learning tasks. Across eight benchmarks, Qwen-CUA outperforms Qwen3.7 and remains competitive with leading proprietary systems, reaching 86.2 on OSWorld-Verified and 18.5/48.4 binary/partial completion on OSWorld 2.0. Scaling the same recipe to a model with over one trillion parameters yields Qwen-CUA-Max, improving these scores to 87.6 and 21.2/53.3. Qwen-CUA also reduces RedTeamCUA attack success from 36.6 to 16.4 relative to Qwen3.7. Efficiency analyses, a browser deployment, and Bash-augmented experiments further characterize practical behavior. These results establish native computer use as a broadly capable agent foundation and highlight scalable verifiable interaction and hybrid tool use as key directions.

cs.LG

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

Vision-Language-Action (VLA) models excel in robotic manipulation but suffer catastrophic performance drops when canonical instructions are simply paraphrased. Although this brittleness is typically addressed through costly data scaling, our probing reveals that the root cause is architectural rather than a lack of semantic understanding. Specifically, we demonstrate that current VLAs successfully retain the correct task identity internally. The failure actually stems from the joint encoding of dynamic visual observations and text, which introduces systematic feature shifts. Because the downstream action policy is highly vulnerable to these variations, it fails to translate the preserved semantics into correct control commands. To resolve this structural bottleneck, we propose Grounded Semantic Re-binding (GSR), an elegant intervention that bypasses unstable joint routing by explicitly fusing independently extracted task semantics with native visual features to train a completely re-initialized action expert from scratch. This targeted intervention dramatically restores paraphrastic invariance using only canonical demonstrations. On the LIBERO-Para benchmark, GSR improves success rates by up to 44.6 percent. It enables lightweight models to rival massively scaled baselines and pushes state-of-the-art models to a new record PRIDE score of 70.4, outperforming the recently introduced large-scale pretrained model Xiaomi-Robotics-0 in instruction generation capabilities. Building on these insights, we also introduce ParaVLA, a natively decoupled 0.33B-parameter model exhibiting near-perfect robustness to instruction rewording. Ultimately, our work proves that robust semantic grounding can be achieved through elegant structural design, bypassing the inefficient brute-force data scaling paradigm.

cs.RO

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. Recent systems increasingly generate multiple dissemination formats, but a practical workflow must also keep the outputs editable in native tools and bound into one navigable deliverable for revision and reuse. We present ResearchStudio-Reel, a native-editable dissemination workspace that binds its three artifacts into one interactive deliverable at the experience level, implemented as five skills executable in Claude Code and Codex: one shared extractor, three editable artifact generators, and one interactive convergence layer. A shared asset bundle feeds a PowerPoint poster and video deck, plus a bilingual Word blog; rather than re-rendering the paper into a fourth format, Paper2Reel converges these already-produced artifacts at the experience level, binding poster regions, video segments, and blog passages into one interactive viewer. Artifact-specific release checks make this delivery contract testable, and Paper2Poster additionally uses a measured-fill loop. On the Paper2Poster benchmark, our Claude Code configuration achieves the best scores among automated systems on all three aesthetic sub-criteria and the best or tied-best scores on two of three information sub-criteria. Under two VLMjudges, it exceeds the authors' posters in average aesthetics (3.56 vs. 3.03) and wins on overall quality on 74 and 95 of the 100 papers under the two judges. The full pipeline additionally packages the native-editable source artifacts and their aligned viewer. Project is available at https://aka.ms/ResearchStudio

cs.CV

Quantitative estimates of propagation of chaos for multi-species cross-diffusion equations

In this paper, we prove the quantitative propagation of chaos results that allow us to derive multi-species cross-diffusion equations from moderately interacting stochastic particle system. The quantitative propagation of chaos result in $L^1$-norm is obtained by the relative entropy method, and the proof is carried out in two steps. In the first step, we quantify the relative entropy between the joint distribution of the particle system and the tensorised solution of the PDE at the intermediate level. In the second step, we establish a rigorous convergence rate to the multi-species cross-diffusion equations by analyzing the $L^2$-distance between the solution of the intermediate-level PDE and that of the limiting PDE. Furthermore, combining the strong $L^1$-convergence for the propagation of chaos with the $L^p$-estimates $(2\le p<\infty)$ for the marginal distribution of multi-species particle system, we derive the corresponding $L^q$-result $(1<q<\infty)$ via interpolation.

math.AP

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers

Transforming static research papers into dynamic media such as posters, slides, and videos is essential for effective dissemination but remains a labor-intensive challenge. Existing automated approaches often treat these formats in isolation and consequently fail to maintain semantic consistency across the entire presentation suite. We address this fragmentation by formalizing the task of unified presentation suite generation and proposing $\textbf{OmniPresent}$ to orchestrate the creation of coherent deliverables. Our framework adopts a renderable HTML representation to enable centralized content planning and a self-correcting verify-and-repair loop that actively resolves conflicts across modalities. We further facilitate scalable research in this domain by releasing $\textbf{OmniPreBench}$, a comprehensive dataset comprising over one thousand papers with paired artifacts, and establishing a rigorous VLM-based evaluation protocol. Empirical results confirm that our method generates high-quality and faithful presentation suites that significantly surpass strong baselines in both accuracy and visual appeal.

cs.SE

An automated method of identifying incorrectly labelled images based on the sequences of loss functions of deep learning networks

Deep learning is widely applied in medical image analysis, but up to 10% of manually labelled images may be incorrect, degrading model performance. This paper proposes an automated method to identify incorrectly labelled medical images by analyzing sequences of loss functions from deep learning classification networks over multiple training epochs. Identified images can be reviewed and relabelled by experts, improving dataset quality and model performance. Two experiments validate the method on a fundus image dataset for referable diabetic retinopathy screening. In the first, 6% (648) of 10,788 gold-standard labels were intentionally flipped. The method identified 75.31% (488) of the flipped samples, with only 4.85% (492) false positives among correctly labelled samples. In the second, reviewing and correcting the 980 identified samples (9.1% of the dataset) and retraining the model improved best accuracy on an independent test set from 95.93% (with 6% label noise) to 96.50% (with 1.5% noise), approaching the ideal 96.57% (with 0% noise). The results demonstrate the method's effectiveness in improving model performance through automated label quality control.

cs.CV

Domain Knowledge Based Temporal-Spatial Graph Convolution Network for ECG Recognition

In light of strides in Arti cial Intelligence (AI) and its wide spread application, challenges persist in the interpretability of AI models, particularly within specialized domains like healthcare, such as electro cardiograph (ECG) recognition. Rather than relying solely on end-to-end convolutional neural networks, this paper introduces a novel approach using a domain knowledge-based graph convolution network for ECG recognition. Key landmarks points of PRQST, vital to ECG interpreta tion, are incorporated as domain knowledge. The double-stream directed graph is employed to model both intra and inter ECG cycles. Speci cally, spatial directed graphs capture the positional relationships among key points, while temporal directed graphs delineate temporal dependencies between adjacent cycles in extended ECG sequences. Experimental re sults on the First Chinese ECG Intelligent Competition dataset, which speci cally classify ECG into nine categories, prove the e cacy of the proposed model. The overall average F1 score is 88.1%, the average F1 score of rare categories is 76.3%, both outperform the state-of-the-art models. The introduction of domain knowledge did enhance the detec tion performance, especially for rare categories.

cs.LG

Human-Agent Collaborative Paper-to-Page Crafting

In the quest for scientific progress, communicating research is as vital as the discovery itself. Yet, researchers are often sidetracked by the manual, repetitive chore of building project webpages to make their dense papers accessible. While automation has tackled static slides and posters, the dynamic, interactive nature of webpages has remained an unaddressed challenge. To bridge this gap, we reframe the problem, arguing that the solution lies not in a single command, but in a collaborative, hierarchical process. We introduce $\textbf{AutoPage}$, a novel multi-agent system that embodies this philosophy. AutoPage deconstructs paper-to-page creation into a coarse-to-fine pipeline from narrative planning to multimodal content generation and interactive rendering. To combat AI hallucination, dedicated "Checker" agents verify each step against the source paper, while optional human checkpoints ensure the final product aligns perfectly with the author's vision, transforming the system from a mere tool into a powerful collaborative assistant. To rigorously validate our approach, we also construct $\textbf{PageBench}$, the first benchmark for this new task. Experiments show AutoPage not only generates high-quality, visually appealing pages but does so with remarkable efficiency in under 15 minutes for less than \$0.1. Code and dataset will be released at $\href{https://mqleet.github.io/AutoPage_ProjectPage/}{Webpage}$.

cs.SE

Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance

Writing effective rebuttals is a high-stakes task that demands more than linguistic fluency, as it requires precise alignment between reviewer intent and manuscript details. Current solutions typically treat this as a direct-to-text generation problem, suffering from hallucination, overlooked critiques, and a lack of verifiable grounding. To address these limitations, we introduce $\textbf{RebuttalAgent}$, the first multi-agents framework that reframes rebuttal generation as an evidence-centric planning task. Our system decomposes complex feedback into atomic concerns and dynamically constructs hybrid contexts by synthesizing compressed summaries with high-fidelity text while integrating an autonomous and on-demand external search module to resolve concerns requiring outside literature. By generating an inspectable response plan before drafting, $\textbf{RebuttalAgent}$ ensures that every argument is explicitly anchored in internal or external evidence. We validate our approach on the proposed $\textbf{RebuttalBench}$ and demonstrate that our pipeline outperforms strong baselines in coverage, faithfulness, and strategic coherence, offering a transparent and controllable assistant for the peer review process.

cs.AI

CasaMaestro: Multi-View Panoramas for House-Scale 3D Reconstruction

The rise of home-deployed embodied AI systems is driving a growing need for fast, metric 3D reconstruction of residential spaces to support navigation, interaction, and long-horizon task execution. However, the commonly used pinhole-camera 3D reconstruction pipelines struggle to model large indoor residences efficiently due to their limited field of view, to which achieving full coverage across multiple rooms often requires thousands of images and incurs drift from long chains of incremental alignment. In this work, we present CasaMaestro (Spanish words meaning ``house'' and ``master''), a feedforward model that can take only twenty to fifty sparse multi-view indoor panoramas as input and directly predicts metric depth along with camera poses, allowing fast point-cloud reconstruction of the entire house with full coverage. CasaMaestro is the first model that supports house-scale reconstruction with multi-view panoramas. Experiments show that CasaMaestro can robustly provide high quality results in both real-world and synthetic scenes, which can serve as a strong foundation for acquiring house-scale 3D indoor assets to be applied in close-loop simulation.

cs.CV

Towards Long-Form Spatio-Temporal Video Grounding

In real scenarios, videos can span several minutes or even hours. However, existing research on spatio-temporal video grounding (STVG), given a textual query, mainly focuses on localizing targets in short videos of tens of seconds, typically less than one minute, which limits real-world applications. In this paper, we explore Long-Form STVG (LF-STVG), which aims to locate targets in long-term videos. Compared with short videos, long-term videos contain much longer temporal spans and more irrelevant information, making it difficult for existing STVG methods that process all frames at once. To address this challenge, we propose an AutoRegressive Transformer architecture for LF-STVG, termed ART-STVG. Unlike conventional STVG methods that require the entire video sequence to make predictions at once, ART-STVG treats the video as streaming input and processes frames sequentially, enabling efficient handling of long videos. To model spatio-temporal context, we design spatial and temporal memory banks and apply them to the decoders. Since memories from different moments are not always relevant to the current frame, we introduce simple yet effective memory selection strategies to provide more relevant information to the decoders, significantly improving performance. Furthermore, instead of parallel spatial and temporal localization, we propose a cascaded spatio-temporal design that connects the spatial decoder to the temporal decoder, allowing fine-grained spatial cues to assist complex temporal localization in long videos. Experiments on newly extended LF-STVG datasets show that ART-STVG significantly outperforms state-of-the-art methods, while achieving competitive performance on conventional short-form STVG. Our code is at: https://github.com/HengLan/ART-STVG.

cs.CV

Projective systems and bounds on the length of codes of non-zero defect

We derive bounds on the lengths of linear codes with fixed Singleton defect $s$, working within the framework of projective systems as advocated by Tsfasman and Vlǎduţ. This geometric perspective allows us to unify and extend a range of existing results. We introduce the parameter $m^s(k,q)$, denoting the maximum length of a non-degenerate $[n,k,d]_q$ A$^s$MDS code, and more generally $m^s_t(k,q)$, where the dual code is additionally required to be A$^t$MDS. We also study $κ(s,q)$, the maximum dimension $k$ for which a length-maximal A$^s$MDS code exists. Among our main results, we provide sufficient conditions on $n$ and $k$ under which the dual of an A$^s$MDS code is necessarily A$^s$MDS, addressing a gap in the existing literature. We show that codes of sufficient length must be projective, meet the Griesmer bound, and be dual to an AMDS code. Our bounds subsume or improve several results in the literature. Two conjectures on the non-existence of length-maximal codes of dimension $k\ge 5$ are proposed, supported by computational evidence.

math.CO

GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments

Training embodied agents in the real world requires skilled operators and expensive hardware. Simulation environments offer a compelling alternative by enabling large-scale, cost-effective data augmentation. Consequently, rapidly constructing high-fidelity simulation scenes with a minimal sim-to-real gap has become a critical objective in robot learning. While reconstruction-based methods provide superior visual quality, current workflows are hindered by inefficient data acquisition and subpar foreground object extraction. We thus propose GASE, a highly automated system for simulation scene construction. GASE leverages multi-view video streams from panoramic camera arrays to enable rapid environment scanning. To ensure high-quality asset generation, our pipeline introduces a camera-pose-based strategy that robustly extracts objects across frames in the 2D domain, followed by high-fidelity scene inpainting. Foreground objects and the static background are then reconstructed independently and seamlessly imported into physics simulators for policy training. Extensive experiments demonstrate that GASE outperforms existing 3D Gaussian-based methods in segmentation accuracy by over 10\% while achieving state-of-the-art inpainting quality. Furthermore, real-robot deployments across manipulation and navigation tasks maintains a performance gap of less than 10\% compared to policies trained purely on real-world data. These results confirm that GASE provides an efficient and highly effective solution for bridging the sim-to-real gap. Code will be released.

cs.RO