arXiv Science⌕ Search

arXiv subjects

Rui Liu

Publications and source records attributed to Rui Liu.

At least 37 records · Page 2Linked to original sources

How Should Diffusion Language Models Edit Code?

Code editing requires a model to decide where to make changes, generate the new content, and preserve everything else. We study how masked diffusion language models divide these responsibilities across four editing interfaces: whole-file rewriting, search-and-replace, locate-then-infill, and token-level editing. Experiments on CanItEdit reveal a composition gap: diffusion models can generate coordinated changes when the correct edit locations are supplied, but much of this capability is lost when those locations must be predicted. Access to the intact original code helps the model fill multiple edit regions, yet does not resolve the difficulty of selecting those regions. By varying the editable regions while holding the generation model and decoding procedure fixed, we identify two distinct requirements for successful editing: covering every required change and placing precise boundaries around it. Missing a required region prevents the corresponding change, while widening regions to ensure coverage can sharply reduce success by requiring unchanged code to be regenerated. A sentence-level Wiki editing probe shows the same qualitative gap between supplied and predicted locations beyond code. These findings show why strong infilling capability alone does not ensure reliable editing: the interface must expose all required changes while limiting regeneration of unchanged code.

cs.SE↗

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pretrained main network, has shown remarkable success in visual-condition controllable generation. Despite its widespread adoption, the role of the side branch and its training efficiency remain underexplored. In this paper, we first revisit this mainstream paradigm through the lens of score-based generative modeling: 1) The main network preserves visual perceptual quality by providing a prior unconditional score. 2) The side network steers conditional control by implicitly contributing a likelihood score. Guided by this perspective, we propose LIkelihood Score Alignment (LISA), an effective regularization method that explicitly aligns the intermediate feature of the side network with an approximated likelihood score. Specifically, we first hook features from a designated layer of the side network and project them into the score latent space by a lightweight decoder. Then, we construct an approximated likelihood score target and calculate the distance between the decoder's output and this target as an additional regularization loss. Finally, we jointly optimize the side network and decoder with both standard diffusion loss and our regularization loss. Experiments across various image/video tasks, architectures, and diffusion/flow models demonstrated that LISA can not only consistently accelerate the training convergence and improve final synthetic results, but also encourage the side network's features to be more disentangled for conditional modeling with negligible additional training cost and zero extra inference cost.

cs.CV↗

SAGE: A Statistical Acceptance Gate for Self-Evolving Agents

Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer's Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above two limitations, we propose a statistical acceptance gate for self-evolving agents (SAGE). Compared with previous work, SAGE has two contributions. First, SAGE proposes a per-item paired comparison that evaluates the current skill and the edited skill on identical validation items, which exposes regressions that an aggregate score hides and penalizes them asymmetrically. Second, SAGE also employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise. SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting. It commits only a subset of the baseline's edits, filtering out those whose gains are unreliable or purchased by breaking already-solved items. Across five benchmarks and four backbone LLMs under an equal-budget protocol, SAGE lowers the regression rate in 19 of 20 settings and matches the baseline in the remaining one, for example from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attains the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.

cs.AI↗

ManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold Perspective

Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form knowledge, suffering from severe edit forgetting and degradation of general capabilities. To address these challenges, we reframe knowledge editing from a manifold perspective, viewing it as a localized displacement of an edit sub-manifold within the global knowledge manifold. Under this formulation, the problem can be decomposed into two key questions: (i) how to identify representative edit points that effectively anchor the edit sub-manifold, and (ii) how to preserve the remaining manifold structure during the sub-manifold displacement process. Based on this perspective, we propose ManiEdit, a novel manifold-aware autoregressive editing framework consisting of two core components. Pivot Localization addresses the mediocre-point dilemma by identifying high-leverage pivots to anchor the edit sub-manifold. Manifold-Aware Preservation preserves different knowledge types through an energy-weighted penalty combined with recursive null-space alignment. Experiments on two base LLMs and four unstructured editing benchmarks demonstrate that ManiEdit achieves state-of-the-art performance, outperforming the strongest baseline by up to +27.81 BERTScore and +8.50 ROUGE-L, while maintaining near-original general capabilities across six representative downstream tasks. Our code is available at: https://github.com/Areyliu/ManiEdit

cs.CL↗

Inverse-Limit Detection Envelopes as Calkin Algebras

For a unital Banach algebra $A$ with $\|1_A\|=1$ and a decreasing sequence $(I_r)$ of proper closed two-sided ideals of finite codimension, we introduce the bounded detection envelope $$\widehat A:=\operatorname{Env}^{b}_{(I_r)}(A)=\varprojlim^{b}(A/I_r,π_r^s),\qquad \|(a_r)_r\|=\sup_r\|a_r\|.$$ Combining vector-valued Bourgain--Delbaen extensions, oracle Turing-machine recursion, and Argyros--Haydon operator analysis, we construct a separable Banach space $X$ such that, isometrically as unital Banach algebras, $$\operatorname{Cal}(X)=\mathcal{B}(X)/\mathcal{K}(X)\simeq\widehat A. $$ We also characterize these envelopes as the unital Banach algebras with norm-one units admitting an isometric predual and countably many jointly faithful weak-star-to-norm continuous finite-dimensional unital representations. Applications include arbitrary finite-dimensional unital Banach algebras with norm-one units and their countable $\ell_\infty$-sums, as well as unitized coordinatewise algebras over normalized $1$-unconditional boundedly complete Schauder bases. These include incidence algebras, $\ell_\infty$, and the unitization of Tsirelson's original space. Further applications concern analytic operator algebras, Fourier--Stieltjes and Fourier multiplier algebras of countable discrete groups, and classical and quantum convolution algebras. In particular, we realize $H^\infty(\mathbb D)$, $\operatorname{Mult}(H_2^2)$, $\mathcal L_2$, $M_{\mathrm{cb}}A(\mathbb F_2)$, and $C^u(O_3^+)^*$ as Calkin algebras, with their canonical norms and products.

math.FA↗

Duals of separable Banach Spaces as Calkin Algebras and Universal Ideal Quotients

For $\mathbb{K}=\mathbb{R}$ or $\mathbb{C}$ and every separable Banach space $V$, we use an oracle-relative two-sorted finite-extension construction, whose oracle records the local rational structure of $V$, to construct a separable Banach space $X_V$ such that $$ \operatorname{Cal}(X_V)=\mathcal{B}(X_V)/\mathcal{K}(X_V)\simeq \begin{pmatrix} \mathbb{K}&0\\ V^*&\mathbb{K} \end{pmatrix} $$ as Banach algebras. The identification of $V^*$ with the Jacobson radical is isometric. Consequently, every nonzero dual Banach space with a separable predual admits, after an equivalent renorming, a unital Banach-algebra structure isomorphic to the Calkin algebra of a separable Banach space. Moreover, there is a separable Banach space $X$ such that every separable Banach space is isometric to $\mathcal{J}/\mathcal{K}(X)$ for a closed two-sided ideal $\mathcal{J}$ of $\mathcal{B}(X)$, and the subspace--ideal correspondence preserves the canonical order.

math.FA↗

Benchmarking EEG Foundation Models at Scale: Lessons from 20,000 Evaluations

Electroencephalography (EEG) foundation models (FMs) promise transferable neural representations, yet their advantages over strong supervised baselines and their prospects for further scaling remain unclear. To address these questions, we introduce EEG-Arena, an open-source benchmark covering 30 EEG FMs and 25 supervised baselines evaluated on 57 downstream tasks from 23 public datasets. Through more than 20,000 evaluations across five experimental protocols, we assess downstream performance, pretraining benefits, model size scaling, pretraining data scaling, and robustness to channel configuration. We find that (1) EEG FMs outperform strong task-specific supervised baselines on most evaluated tasks, particularly under non-bipolar settings; (2) compared with architecture-matched supervised training from scratch, pretraining improves both early optimization and final downstream performance, with larger and more consistent gains as more labeled downstream data become available; (3) existing EEG FMs do not exhibit a consistent positive relationship between parameter count and downstream performance; (4) under a fixed architecture, increasing the pretraining data scale yields sustained downstream gains; and (5) channel-flexible FMs achieve higher absolute performance than channel-constrained models across most evaluated channel configurations. Together, these findings demonstrate the downstream value of EEG FMs and identify pretraining data expansion as a promising direction for further progress. To support continued research, we release EEG-Arena as an open-source evaluation framework that provides shared infrastructure for reproducible benchmarking, model comparison, and community-driven development.

cs.LG↗

Recursive Self-Improvement via On-Policy Distillation for Reasoning

On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns to mimic the privileged teacher model, while the teacher remains frozen throughout training. Previous work showed that freezing the teacher is useful for training stability, but we argue that this can prevent the teacher from incorporating the improvements learned by the student during training. Our primary contribution is to address this limitation with a recursive framework built around two complementary components. First, we let the privileged teacher co-evolve with the student so that revision learned in one round can guide the next, a process we refer to as Dynamic Co-Evolution (DCE). Second, because stronger revision can also make responses too verbose and self-critical, we additionally train on shorter, verified rewrites of the model's own on-policy responses. We call this complementary objective Self-Refined Concise Learning (SRCL). Overall, our comprehensive evaluations show that DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks. Specifically, on Qwen3-8B, DCE+SRCL reaches 65.97% Average@12, outperforming OPSD by 35.62 percentage points while reducing mean output length by 7.80% relative to DCE alone.

cs.CL↗

Parametric Study of the Torus Instability Threshold

The torus instability of an arched current channel has been suggested to initiate and drive major solar and stellar eruptions. Its threshold, given by the critical decay index of the equilibrium external poloidal field (the so-called strapping field) at the position of the current channel, is insufficiently known. Here, we carry out a parametric numerical study of the threshold, employing the force-free Titov-Démoulin (TD) equilibrium of a line-tied partial toroidal current channel and flux rope. This addresses the scatter of the threshold about its canonical value, $n_\mathrm{cr}=3/2$. Values scattering in the range $n_\mathrm{cr}\approx$\,1--2 are typically found in numerical and observational studies of flux rope eruptions on the Sun. For zero external toroidal (guide, or shear) field and approximately semicircular geometry (corresponding to minimal photospheric line-tying), we find the threshold to lie in the theoretically expected range of $\approx$\,1--1.5. An external toroidal field introduces a strong stabilizing effect on the instability, raising the threshold up to $\sim$\,2.5, which can explain observational and numerical results above the canonical value. Line-tying is found to act as a stabilizer as well. We also consider the approximate threshold based on the potential field and find a very good agreement with the exact numerical value, provided the horizontal component perpendicular to the flux rope axis is used to approximate the external poloidal field.

astro-ph.SR↗

OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling

Automated operations research (OR) modeling requires LLMs to translate natural-language decision problems into correct mathematical programs. Existing methods can improve individual formulations, but they often solve problems in isolation, retaining little reusable experience and repeating similar formulation errors. Prior memory-based approaches store examples, thoughts, or insights as references, while OR modeling requires reusable formulation skills that transfer across problem narratives and guide concrete modeling decisions. We propose OptiSkill, a skill-augmented framework that builds a hierarchical and evolving SkillBank for LLM-based OR modeling. SkillBank stores solver-verified experience as reusable skills, with Global Strategies for problem-level formulation skeletons and Step Experiences for local error-prevention rules. It is further refined through stable batch-level test-time evolution, where candidate skills are incorporated only after validation. Experiments on eight OR modeling benchmarks show that OptiSkill improves formulation accuracy across LLM backbones, outperforms strong agentic baselines, and gains further by expanding SkillBank coverage and reliability. Code and data are available at https://github.com/rachhhhing/OptiSkill

cs.AI↗

When More Is Not Better: Component Anti-Synergy in a P300 Speller

P300 brain-computer interface (BCI) spellers can provide hands-free communication for people with severe motor impairments. Modern pipelines combine multiple individually promising components, often assuming that 'more-is-better'. We tested this assumption using a four-component full-factorial experiment varying the inclusion of Euclidean Alignment (EA), xDAWN spatial filtering, subject calibration, and language model priors on a public P300 dataset. Performance was evaluated using accuracy, repetitions, and information transfer rate (ITR) with mixed-effects models. Results show that the value of components is conditional rather than additive. Calibration was the strongest singular contributor, while EA compensated for its absence in zero-calibration settings. Adding independently useful components could also reduce performance, revealing component anti-synergy. Contrary to conventional wisdom, LM support was not universally beneficial: its effect depends strongly on the strength of the underlying EEG pipeline, while results from a larger LM showed a similar pattern. Together, these findings challenge maximal 'all-on' pipeline design and highlight the value of selecting spatial and language-support components according to the quality of available EEG evidence.

cs.LG↗

On-chip squeezed light in the audio frequency band

Squeezed light in the audio-frequency band is a key resource for quantum metrology and quantum sensing. However, realizing stable audio-frequency squeezed light on integrated photonic platforms remains challenging due to technical noise and the difficulty of scalable phase referencing. Here, we demonstrate on-chip generation of audio-band two-mode squeezed states down to 60 Hz in a silica microcavity. To enable phase-stable operation without directly locking fragile quantum modes, we develop a coherent-comb control method in which a weak electro-optic reference comb co-propagates with the vacuum at the quantum frequency modes in an orthogonal polarization. This scheme provides quadrature measurement and long-timescale phase stability, thereby enabling covariance-matrix reconstruction. We verify the entanglement with the positive partial transposition criterion, which confirms inseparability via a minimum symplectic eigenvalue of 0.395 ($<0.5$). Our results establish an experimentally accessible route toward on-chip phase-stable audio-band squeezing and support the scalable framework for continuous-variable quantum information processing with integrated photonics.

quant-ph↗

City Editing: Hierarchical Agentic Execution for Dependency-Aware Urban Geospatial Modification

Urban renewal requires incremental modifications to existing geospatial plans, yet manually updating complex layouts under spatial constraints is labor-intensive and error-prone. To tackle this, we propose CEAE, a hierarchical agentic framework that formulates urban renewal as machine-executable GeoJSON editing from natural-language instructions. CEAE decomposes instructions into hierarchical geometric intents, executing edits from coarse to fine while preserving spatial consistency through a self-reflective execution-validation loop. Experimental results show that CEAE outperforms baselines in execution validity, robustness, and geometric accuracy.

cs.MA↗

Streaming P300 Acquisition and Statistical Signal Validation Across Five EEG Platforms: A Hardware-Agnostic BrainFlow/LSL Pipeline

P300 spellers offer people with severe motor impairment, such as ALS, an effective communication channel and remain one of the most established surgery-free alternatives to intracortical interfaces. Advanced language models have made spellers faster and more robust, yet the hardware beneath them is under-studied. We present a hardware-agnostic, real-time P300 acquisition pipeline built on BrainFlow and Lab Streaming Layer (LSL) that runs unchanged across consumer- and research-grade EEG headsets, with permutation tests of signal separability. Using a standard 6 x 6 row/column paradigm, we piloted five configurations: a custom dry system, a custom wet/gel system, Emotiv Flex, Emotiv EPOC X, and Muse 2. The custom systems and EPOC X showed weak or inconsistent signal separability, Muse 2 had the highest acquisition reliability despite limited centro-parietal coverage, and Flex showed the most promising signal. In 20 further Flex sessions varying subject, timing, and phrase length (131 target characters), a peak-amplitude permutation test and a cross-validated xDAWN decoder both detected a significant target response under two channel-exclusion policies, with decoder AUC reaching about 0.72 after 15 repetitions. Character accuracy depended heavily on evaluation methodology: in-sample majority voting reached 94.7%, whereas character-held-out accuracy was 31.3% with evidence accumulated across repetitions, about three times that of held-out majority voting. These analyses indicate that Flex captured a detectable, if still weak, P300 under the studied conditions, while broader participant-level validation and improved decoding remain necessary.

cs.HC↗

Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search

Feature transformation improves predictive performance on tabular data by constructing informative abstractions from raw features. Recent generative approaches encode transformation knowledge into continuous embedding spaces for efficient exploration of candidate strategies, but face three key limitations: (1) overlooking hierarchical relationships between low-level features, operations, and high-level abstractions; (2) enforcing order-sensitive embeddings on inherently permutation-invariant transformation sequences, thereby introducing systematic bias; and (3) relying on gradient-based search, which is ill-suited to non-convex transformation spaces. We propose a framework with two complementary components. First, a permutation-invariant hierarchical module captures interactions across features, operations, and abstraction levels, with a self-attention pooling mechanism that maps semantically equivalent structures to consistent embeddings aligned with downstream performance. Second, a policy-guided multi-objective reinforcement learning strategy initializes the search from empirically strong seeds and jointly optimizes predictive accuracy and transformation efficiency. Extensive experiments on diverse tabular benchmarks demonstrate the effectiveness and robustness of our framework against strong baselines. Our code and data are publicly available at: https://github.com/RayLiu1103/PHER.

cs.LG↗

SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use

High-quality multi-turn tool-use data is essential for training agentic models, yet existing data synthesis methods often underrepresent the argument-level dependencies that are critical to long-horizon tool use. As a result, even when a model selects the correct tool, task execution may still fail because the model fills tool arguments with fabricated, stale, or weakly grounded values. To address this problem, we propose \textbf{State-Guided Data Synthesis with Argument Provenance (SAP)}. SAP combines state guidance, tool-argument provenance constraints, and turn-level validation to efficiently construct tool-use trajectories with long-range dependencies and high accuracy. Using data generated by SAP, we build SAP-4B, which is highly competitive even when compared with much larger models across multiple benchmarks. Source code, synthesized data, and trained weights are available at https://github.com/Zichen1024/SAP.

cs.AI↗

Visual Representation and History Modeling for Navigation World Models

Navigation World Models (NWMs) predict action-conditioned visual futures for planning. Two practical challenges are central to their design: selecting a suitable visual representation and efficiently modeling observation history for repeated candidate queries. Standard Global-Softmax attention provides flexible interactions but repeatedly processes the same history, leading to increasing computation and memory costs for long contexts and multi-query planning. We study both problems within a unified conditional flow-transformer framework. We first compare five frozen visual representations under the same dynamics model and evaluation. To reduce redundant history computation, we design Cached-Linear, a hybrid architecture that combines local and shifted-window attention for target mixing with linear attention for reusable history access. We further develop Balanced Gated Delta Network (GDN), which augments this design with frame-wise recurrent memory for temporal history modeling. Experiments on RECON, SACSoN, and SCAND show that representation choice depends on the prediction objective: PAE-L performs best for reconstruction, RAE-B for direct prediction, and V-JEPA for long-horizon rollout. Under shared-history workloads, Cached-Linear substantially reduces computation and memory compared with Global-Softmax, while Balanced GDN improves selected direct-prediction endpoints with efficient context reuse. Overall, we systematically study visual representation and history modeling for NWMs and develop hybrid reusable-history architectures for efficient long-context and multi-query prediction.

cs.CV↗

From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editing. We therefore formulate depth and surface-normal prediction as image-form denoising targets, using these dense tasks as structured visual supervision within the same creation interface. Our framework decouples semantic interpretation from spatially aligned visual injection while sharing one multimodal diffusion transformer (MMDiT) backbone across all tasks. Mutual Context Attention (MCA), a paired-video data-construction procedure, and a progressive training curriculum then connect the learned structural cues to temporally localized editing and reference-conditioned creation. A single checkpoint obtains the highest overall score in the reported comparison of unified systems (4.15); adding dense supervision improves OpenVE Overall from 3.98 to 4.06 and Local Add from 3.92 to 4.18. These results support a deliberately bounded conclusion: perception-oriented dense supervision transfers useful structural knowledge to downstream creation, especially editing locality and preservation; we do not claim superiority as a standalone dense predictor.

cs.CV↗