arXiv ScienceSearch

arXiv subjects

Lei Yu

Publications and source records attributed to Lei Yu.

At least 19 recordsLinked to original sources

A Sharp Small-Coefficient Variant of Khintchine's Inequality and the Sharp $\pi/2$ Theorem

We prove a refined quadratic normal approximation for the first absolute moment of normalized weighted Rademacher sums with bounded maximal coefficients. For any weight vector $w\in\mathbb{R}^{n}$ satisfying $\|w\|_{2}=1$ and $\|w\|_{\infty}\leq\beta$ with sufficiently small $\beta>0$, we establish the uniform error bound $|\mathbb{E}|\sum_{i=1}^{n}w_{i}X_{i}|-\sqrt{2/\pi}|=O(\beta^{2})$ over all admissible weight configurations. Our proof combines zero-bias Stein's method and refined small-ball probability estimates to exploit symmetry cancellation and control the non-smooth residual of the absolute-value test function. An explicit extremal construction further verifies the optimality of this quadratic convergence rate. As an application, we establish an asymptotically sharp refinement of the Friedgut--Kalai--Naor (FKN) theorem for Boolean functions, also known as the sharp $\pi/2$ theorem, characterizing the level-1 Fourier energy for functions deviating far from dictatorships.

math.PR

Exact Common Information and Exact Channel Synthesis for Correlated Gaussian Sources

In this paper, we resolve two conjectures posed by Yu and Tan in 2020 (in two separate papers published in the IEEE Trans. Inf. Theory). Specifically, we establish that: 1) the exact common information for a pair of $\rho$-correlated Gaussian sources is given by the conjectured expression $\frac{1}{2}\log\frac{1+\rho}{1-\rho}+\frac{\rho}{1+\rho}$; and 2) the admissible region for the shared randomness rate and the communication rate in exact channel synthesis is exactly the conjectured one. These results yield two important consequences. First, for any $\rho>0$, the exact common information of a correlated Gaussian pair strictly exceeds Wyner's common information. Second, for $\rho>0$, the exact channel synthesis of such a pair requires strictly higher rates than the total-variation version. The proof combines an exact optimal-transport representation of the worst-case Gaussian cross-entropy, Fathi's Gaussian transport inequality, and a determinant inequality arising from the covariance structure of the conditional means.

cs.IT

Cusp-singularity-enhanced Coriolis effect for ultrasensitive chip-scale gyroscopes

Gyroscopes, as fundamental inertial sensors, are crucial for rotation measurements in consumer electronics, automotive, and aerospace industries, with the most widely used kind relying on the Coriolis effect. The chip-scale Coriolis vibratory gyroscopes (CVGs) show reduced size, weight, and cost, but remain far lower performance than traditional macroscale CVGs, as the weak intrinsic Coriolis factor sets a fundamental limit on scaling the sensitivity against the inherently louder Brownian noise in microchips compared to the macroscale ones. Here, to overcome this physical limit, for the first time, we propose and experimentally demonstrate the use of third-order singularities lying within cusp catastrophes in the phase-tracked oscillations of an on-chip CVG to facilitate a cubic-root scaling of the Coriolis-effect-induced frequency modulation. Employing this effect, we achieve a three-order-of-magnitude enhancement in the Coriolis factor, yielding a 253-fold improvement in signal-to-noise ratio and a 297-fold increase in precision. Moreover, the cusp singularity enables a previously unattainable ultrasensitive phase-modulated sublinear measurement, achieving a world-record signal-to-noise ratio performance for silicon-chip gyroscopes. These findings not only provide revolutionary advancements in gyroscope technologies, by filling the gap in observing and controlling the singularity-enhanced Coriolis effect, but also shed new light on other ultrasensitive sensing applications.

physics.ins-det

On The Most Discriminative Boolean Functions for Correlated Sources

Motivated by a conjecture of Amari and Kobayashi, we study the problem of identifying pairs of Boolean functions that maximize the Kullback-Leibler divergence between two distributions obtained by separately compressing two correlated sources. When the reference distribution corresponds to independent sources, this problem reduces to the problem of maximizing mutual information, for which the optimality of dictator functions has been proved by Pichler, Piantanida, and Matz. For the problem of maximizing Fisher information, which can be viewed as a local version of the problem studied in this paper, Amari and Kobayashi conjectured that parity functions are optimal. For unbiased pairs of Boolean functions, and for identical pairs in the nonnegative correlation regime, we prove that both the divergence and the Fisher information are maximized by level-$k$ functions, namely, functions whose Fourier coefficients are supported only on level $k$. Since level-$k$ functions include parity functions, this gives a partial resolution of the conjecture of Amari and Kobayashi. Furthermore, in the framework of Bayesian distributed one-bit hypothesis testing, we prove that level-$k$ functions are optimal among all pairs of functions. Finally, we also discuss the one function version of the problem studied in this paper, which can be regarded as the divergence analogue of the Courtade and Kumar conjecture.

cs.IT

From Stacking Disorder to Cubic Order: Ice Crystallization from Deeply Supercooled Water

Crystallization far from equilibrium can generate morphologies that defy classical crystal habits, yet the microscopic mechanisms linking atomic-scale disorder to emergent macroscopic order remain elusive. Here we use in situ cryogenic transmission electron microscopy with a membrane-encapsulated microdroplet platform to directly visualize the freezing of deeply supercooled water at molecular resolution. We show that homogeneous nucleation produces stacking-disordered ice composed of mixed hexagonal and cubic sequences, in which cubic ice initially exists only as isolated monolayers. The gradual thickening of these cubic layers constitutes the key kinetic mechanism that governs the entire crystallization pathway. As thickening proceeds, nanoscale, defect-free cubic ice germs nucleate on the basal planes of the disordered lattice. These faceted cubic germs act as facet-registered kinetic seeds that enforce cubic twinning and sequentially multiply growth branches. This kinetic pathway reproducibly generates robust eight-branched dendrites with global cubic (octahedral) symmetry, even though each branch remains highly stacking-disordered. At later stages, latent heat release drives a crossover to the thermodynamically favored hexagonal phase; remarkably, the pre-established global cubic symmetry is retained. These results reveal how strong kinetic driving forces convert microscopic disorder into emergent macroscopic symmetry, providing a general framework for understanding and controlling rapid crystallization far from equilibrium.

cond-mat.mtrl-sci

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts

Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation. However, VAR still suffers from inherent deficiencies in multi-scale representation learning. Specifically, lower scales primarily capture global semantics, while higher scales focus on fine-grained details. Employing a shared architecture across scales induces optimization conflicts. Moreover, due to the causal autoregressive process, inaccurate semantics at early scales can propagate and significantly degrade the final output. To address these issues, we introduce a scale-aware token-routed Mixture of Experts (MoE) architecture, allowing scale-adaptive expert selection, thereby facilitating decoupled representation learning across scales. In addition, we enhance semantic modeling at early scales by incorporating external self-supervised features. Unlike naive alignment, we analyse and design a residual feature aggregation scheme tailored to the VAR paradigm. Extensive experiments show that our method significantly improves both training efficiency and generation quality. On the ImageNet 256*256 benchmark, our model achieves a superior FID compared to the dense baseline while requiring only half of the default training epochs and a smaller parameter budget, with a merely marginal increase in training cost. Moreover, the performance gap further widens with larger training epochs.

cs.CV

Bash-Commenter: Leveraging Syntax-Aware Preference Optimization to Reinforce Large Language Model for Bash Code Comment Generation

Bash script comprehension is challenging due to Bash's syntactic freedom and complex command structures. Despite its critical role in system administration, Bash scripts often lack adequate comments, hindering readability and maintainability. Existing automated comment generation approaches face two main challenges: (1) limited training datasets that inadequately represent real-world Bash usage patterns; and (2) insufficient understanding of Bash-specific concepts by Large Language Models (LLMs). To address these, we propose Bash-Commenter, an advanced comment generation method based on LLaMA-3.1-8B. First, we construct a comprehensive dataset of complex, multi-line Bash scripts with high-quality comments. Second, we conduct Continual Pre-training (CPT) on large-scale Bash data, followed by Supervised Fine-tuning (SFT), strengthening the model's foundational knowledge of Bash syntax and semantics. Finally, we introduce Syntax-Aware Preference Optimization (SAPO), which constructs preference pairs by applying atomic operations to a script's Abstract Syntax Tree (AST), creating minimal pairs of correct and subtly incorrect scripts for fine-grained semantics learning. Our method outperforms state-of-the-art baselines, achieving 33.40% BLEU-4, 58.26% METEOR, and 57.03% ROUGE-L for 1,064 single-line commands, and 22.15% BLEU-4, 43.89% METEOR, and 32.80% ROUGE-L for 1,046 multi-line scripts. Human and LLM evaluations further confirm superior comment quality in correctness, completeness, and naturalness.

cs.SE

BashCoder-R1: Towards Robust and Explainable Bash Code Generation with Robustness-Aware Group Relative Policy Optimization

Bash scripts are critical for system administration, DevOps, and CI/CD, where code quality affects stability and security. However, LLM-generated scripts often lack reasoning and contain robustness flaws such as mishandled edge cases and unchecked failures. This is particularly critical in production environments where even minor errors can lead to service disruptions. We propose BashCoder-R1, a framework that jointly addresses both issues by treating explainability as a design goal. The pipeline has three stages. Continual Pre-training adapts to Bash syntax. Long Chain-of-Thought Supervised Fine-Tuning on expert-validated samples teaches risk-averse reasoning before code generation. Robustness-Aware Group Relative Policy Optimization optimizes a weighted reward for syntax correctness, robustness (verified by shellcheck), and format adherence. This staged design ensures that the model progressively acquires syntax knowledge, reasoning capability, and robust decision-making. On our BashBench benchmark (952 real-world tasks, 773 single-line and 179 multi-line), BashCoder-R1 achieves SyntaxPass of 100.00/94.97, RobustWarnRate of 4.01/16.47, RobustPass of 95.99/79.33, FuncRate of 93.01/93.85, and FullRate of 90.04/73.18 for single-line and multi-line tasks, respectively. These are relative FullRate improvements of 37.82 and 20.18 percent over the strongest baseline, DeepSeek-V3.2 (Reasoning). Human evaluation confirms its reasoning chains are highest in quality.

cs.SE

Learning with a Single Rollout via Monte Carlo Pass@k Critic

Estimating token-level advantages in reinforcement learning (RL) for language models remains challenging because scaling up episodic experience collection is expensive. The difficulty intensifies for baseline advantage estimation methods, where repeated sampling causes trajectories to diverge into substantially different reasoning prefixes. In this context, RL algorithms such as GRPO prove limited: an outcome reward is too sparse to be attributed to specific actions like intermediate steps, and comparisons across sampled traces are non-trivial because they are heterogeneous. To mitigate both the computational cost of repeated sampling and the difficulty of credit assignment, we study single-rollout proximal policy optimization (SR-PPO) featuring token-level credit assignment in RL for language models. Instead of estimating advantages by normalizing episodic returns within the candidate group, we train a calibrated token-level credit critic using Monte Carlo outcomes from one rollout per prompt. Specifically, we use the critic to predict the Pass@k success probability at the prompt prefix, which is derived from a Pass@1 attempt. This choice yields a more selective learning signal than Pass@1: it discounts easily solved prefixes while prioritizing hard ones whose success probability remains marginal. We show that as $k$ increases, Pass@k converges to a reachability indicator, reflecting whether a prefix can lead to at least one successful continuation. In an explicit state graph, the limit ($k \rightarrow \infty$) can be computed in $O(|V|+|E|)$ time, offering a promising surrogate for direct credit assignment without the need to sample contrastive traces. As an initial validation, SR-PPO exhibits stable learning dynamics, along with consistent gains in Pass@128 success rates on mathematical reasoning benchmarks such as HMMT26 and AIME24.

cs.LG

A Systematic Survey on Event Camera Representation Learning

Event cameras offer distinctive advantages, including microsecond-level latency and high dynamic range, rendering them promising for challenging perception tasks. Inspired by biological vision, they output asynchronous and sparse event streams rather than dense image frames, creating a fundamental mismatch with mainstream neural networks. This survey reviews recent advances in event camera representation learning from the perspective of converting raw event streams into learnable representations. We first organize existing methods according to whether they rely on a single principal event representation or jointly exploit multiple complementary representations. Single-representation methods are further categorized into dense-based representations, which regularize events into structured grid-like forms, and sparse-based representations, which preserve event-native discrete spatio-temporal structures. Multi-representation methods are organized into dense-dense and dense-sparse hybrid formulations that exploit complementary representation properties. This representation-centric taxonomy clarifies how different paradigms balance structural regularity, temporal fidelity, sparsity preservation, and architectural compatibility. For each paradigm, we examine the underlying design choices, modeling principles, and task-level implications. We further summarize standard benchmarks and evaluation settings across representative high-level perception and low-level vision tasks. Finally, we discuss open problems and outline future directions from fixed representation design toward adaptive representation optimization, improved fidelity-efficiency trade-offs, and more scalable event-based perception systems.

eess.IV

ATLAS: Agentic Taxonomy of Large-Scale Software Ecosystems

The open-source ecosystem on GitHub lacks a systematic hierarchical taxonomy of software repositories. GitHub Topics, the dominant organizational mechanism, is flat, inconsistent, and covers only 67% of projects. We present ATLAS, the first framework that automatically constructs a hierarchical taxonomy for software repositories and classifies projects into it end-to-end. By combining LLM global knowledge with real repository distributions, ATLAS proposes meaningful splitting dimensions and iteratively corrects those that fail to accommodate real projects. A Designer Agent proposes splitting dimensions while a Classifier Agent assigns repositories; a self-corrective refinement loop uses classification failures to drive dimension revision through escalating strategies. We evaluate ATLAS on 54,387 GitHub repositories against six baselines spanning four paradigms, two downstream tasks, and three model families. On a stratified 2,001-repository benchmark, ATLAS achieves a Taxonomy Quality F-score (TQF) of 83.13%, outperforming the best baseline by 15 percentage points (on the full 54k corpus the approximate TQF is 73.0%, a gap driven by Path Granularity's all-or-nothing scoring on longer paths rather than lower classification accuracy). It is the only method to simultaneously achieve high structural quality and high practical applicability. On downstream tasks, ATLAS enables alternative discovery with P@1 = 85.71%, surpassing even human-curated lists (62.34%), and achieves the highest P@1 for repository retrieval. The taxonomy further reveals structural ecosystem trends that are difficult to obtain from flat tags or similarity methods: the shift from libraries to AI/ML applications (now 61% of newly community-adopted projects) becomes visible only through hierarchical, type-based categorization. An interactive taxonomy explorer is available at https://atlas-taxonomy.netlify.app/

cs.SE

See Better, Foresee Better, Act Wiser: Physically Grounded Proactive Modeling and Decision Making

Reliable proactive agents must choose an action and judge whether current evidence is sufficient to act. We study retail service from sparse third-person video: before an explicit customer request, an agent must use limited human-object interaction evidence to intervene or remain silent. Physical grounding here means converting observations into task-relevant retail state, not modeling low-level dynamics. We introduce the Proactive Intent World Model (PIWM): See constructs the perceptual basis, Foresee models counterfactual consequences, and Act selects an action. Performance is poor when the agent must extract information from raw video and decide directly, but improves substantially with structured inputs extracted and annotated from a professional retail perspective. AIDA-stage constraints and BDI-state ablations further support role- and goal-directed selection and organization of decision-relevant cues. Counterfactual prediction performs well in standalone evaluation, yet planning methods that query these forecasts at inference time degrade sharply: locally useful consequence prediction does not reliably improve action selection. This gap may reflect incomplete process understanding, uncertainty in fine-grained single-step outcomes, and insufficient joint modeling of scenes and temporal evolution. Hold remains the hardest action in structured-state evaluation, exposing a related challenge in temporal awareness. PIWM advances static intent recognition toward intent world modeling by organizing observations under task knowledge, anticipating candidate interventions, and treating intervention and non-intervention jointly. Future work will introduce long-horizon interaction trajectories and temporal consequence supervision to improve sustained reasoning and intervention timing.

cs.CL

Advancing Mathematics Research with AI-Driven Formal Proof Search

Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research. A mitigation is using LLMs to generate formal proofs in languages like Lean. We perform the first large-scale evaluation of this method's ability to solve open problems. Our most capable agent autonomously resolved 9 of 353 open Erd\H{o}s problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures, and is being deployed in combinatorics, optimization, graph theory, algebraic geometry, and quantum optics research. A basic agent alternating LLM-based generation with Lean-based verification replicated the Erd\H{o}s successes but proved costlier on the hardest problems. These findings demonstrate the power of AI-aided formal proof search and shed light on the agent designs that enable it.

cs.AI

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices

The Diffusion Transformer (DiT) architecture is the state-of-the-art paradigm for high-fidelity image generation, underpinning models like Stable Diffusion-3 and FLUX.1. However, deploying these models on resource-constrained mobile devices entails prohibitive computational and memory overhead. While efficiency-driven approaches like Linear-DiT and static pruning alleviate bottlenecks, they often incur quality degradation. Unlike cloud environments, mobile constraints require a single-model paradigm that dynamically balances fidelity and latency. We introduce ElasticDiT, which achieves this dynamic trade-off by adjusting spatial compression ratios and DiT block depths. By integrating Shift Sparse Block Attention (SSBA) and a Tiny DWT-Distilled VAE (T-DVAE), ElasticDiT reduces inference latency and memory footprint while maintaining image quality. Experiments confirm that ElasticDiT effectively covers a wide range of fidelity-latency trade-offs within a single set of parameters. By jointly adjusting compression and depth, a single ElasticDiT model can be reconfigured on-the-fly to outperform task-specific baselines. Specifically, our flex lite variant achieves an HPS of 32.87, surpassing the Flux model, while maintaining competitive quality at 84.16 percent average sparsity through SSBA. Furthermore, the plug-and-play T-DVAE provides SD3-level reconstruction with only 1/8x the computational cost of standard VAEs, and Flow-GRPO boosts semantic alignment (GenEval: 66.93 to 73.62). These results demonstrate that ElasticDiT offers a versatile, hardware-adaptive solution that eliminates the need for multiple specialized models, providing a promising path for future high-resolution image generation on mobile devices.

cs.CV

OxyGent: Making Multi-Agent Systems Modular, Observable, and Evolvable via Oxy Abstraction

Deploying production-ready multi-agent systems (MAS) in complex industrial environments remains challenging due to limitations in scalability, observability, and autonomous evolution. We present OxyGent, an open-source framework driven by two core novelties: a unified Oxy abstraction and the OxyBank evolution engine. The unified abstraction encapsulates agents, tools, LLMs, and reasoning flows as pluggable atomic components, enabling Lego-like scalable system composition and non-intrusive monitoring. To enhance observability, OxyGent introduces permission-driven dynamic planning that replaces rigid workflows with execution graphs generated at runtime, providing adaptive visualizations. Furthermore, to support continuous evolution, OxyBank serves as an AI asset management platform that drives automated data backflow, annotation, and joint evolution. Empirical evaluations and real-world case studies show that OxyGent provides a robust and scalable foundation for MAS. OxyGent is fully open-sourced under the Apache License 2.0 at https://github.com/jd-opensource/OxyGent.

cs.AI

MTServe: Efficient Serving for Generative Recommendation Models with Hierarchical Caches

Generative recommendation (GR) offers superior modeling capabilities but suffers from prohibitive inference costs due to the repeated encoding of long user histories. While cross-request Key-Value (KV) cache reuse presents a significant optimization opportunity, the massive scale of individual user states creates a storage explosion that far exceeds physical GPU limits. We propose MTServe, a hierarchical cache management system that virtualizes GPU memory by leveraging host RAM as a scalable backup store. To bridge the I/O gap between tiers, MTServe introduces a suite of system-level optimizations, including a hybrid storage layout, an asynchronous data transfer pipeline, and a locality-driven replacement policy. On both public and production datasets, MTServe delivers up to 3.1* speedup while maintaining near-perfect hit ratios (>98.5%).

cs.LG

Detoxification for LLM: From Dataset Itself

Existing detoxification methods for large language models mainly focus on post-training stage or inference time, while few tackle the source of toxicity, namely, the dataset itself. Such training-based or controllable decoding approaches cannot completely suppress the model's inherent toxicity, whereas detoxifying the pretraining dataset can fundamentally reduce the toxicity that the model learns during training. Hence, we attempt to detoxify directly on raw corpora with SoCD (Soft Contrastive Decoding), which guides an LLM to localize and rewrite toxic spans in raw data while preserving semantics, in our proposed HSPD (Hierarchical Semantic-Preserving Detoxification) pipeline, yielding a detoxified corpus that can drop-in replace the original for fine-tuning or other training. On GPT2-XL, HSPD attains state-of-the-art detoxification, reducing Toxicity Probability (TP) from 0.42 to 0.18 and Expected Maximum Toxicity (EMT) from 0.43 to 0.20. We further validate consistent best-in-class results on LLaMA2-7B, OPT-6.7B, and Falcon-7B. These findings show that semantics-preserving, corpus-level rewriting with HSPD effectively suppresses downstream toxicity while retaining data utility and allowing seamless source-level mitigation, thereby reducing the cost of later model behavior adjustment. (Code is available at: https://github.com/ntsw2001/data_detox_for_llm)

cs.CL

scpFormer: A Foundation Model for Unified Representation and Integration of the Single-Cell Proteomics

The integration of single-cell proteomic data is often hindered by the fragmented nature of targeted antibody panels. To address this limitation, we introduce scpFormer, a transformer-based foundation model designed for single-cell proteomics. Pre-trained on over 390 million cells, scpFormer replaces standard index-based tokenization with a continuous, sequence-anchored approach. By combining Evolutionary Scale Modeling (ESM) with value-aware expression embeddings, it dynamically maps variable panels into a shared semantic space without artificial discretization. We demonstrate that scpFormer generates global cell representations that perform competitively in large-scale batch integration and unsupervised clustering. Moreover, its open-vocabulary architecture facilitates in silico panel expansion, assisting in the reconstruction of biological manifolds in sparse clinical datasets. Finally, this learned protein co-expression logic is transferable to bulk-omics tasks, supporting applications like cancer drug response prediction. scpFormer provides a versatile, panel-agnostic framework to facilitate scalable biomarker discovery and precision oncology.

q-bio.QM