arXiv ScienceSearch

arXiv subjects

Shuai Lu

Publications and source records attributed to Shuai Lu.

At least 19 recordsLinked to original sources

Adaptive Schauder Stochastic Mirror Descent in Banach Spaces

In this paper, we introduce an adaptive regularization strategy for stochastic mirror descent (SMD) to solve a class of risk functional minimization problems in infinite-dimensional Banach spaces. This regularization strategy centers on using a Schauder basis to construct a nested family of finite-dimensional subspaces, with the dimension chosen adaptively according to the sample size $n$. We then restrict each SMD subproblem to the corresponding subspace and project the stochastic gradient onto its dual space. This yields closed-form solutions to the SMD subproblems and coordinate-wise updates of the basis coefficients, enabling an implementation with low computational and storage complexity. The subspace dimension also serves as a regularization parameter that balances approximation and optimization errors. For risk functional minimization in $\mathcal{L}^p$ spaces with $1<p<\infty$, we construct Bregman distances adapted to the geometry of the underlying Banach spaces using $\max\{2,p\}$-convex functionals induced by their uniform convexity. At the non-uniformly convex $\mathcal{L}^1$ endpoint, we instead construct a locally strongly convex functional based on the entropy function. By developing a new analytical framework, we establish a convergence rate of $\mathcal O\left(n^{-\min\{\frac12,\frac1p\}}\right)$, up to logarithmic factors. In the misspecified setting, where the minimizer satisfies only weaker regularity conditions, we prove that the risk functional still converges to its minimum value. Finally, we apply the method to statistical inverse problems and illustrate its empirical performance through numerical experiments in both settings.

math.OC

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

Two structural insights have been overlooked in automated residential floor plan generation. First, design is inherently progressive. Architects begin with rough strokes and refine them over time, whereas existing methods typically require their conditioning representation to be fully specified before generation, a fundamental mismatch with how design actually works. Second, the 2D floor plan is not an optional intermediate but an irreplaceable spatial contract. Once room boundaries, doors, and windows are fixed, furnishing reduces from open-ended spatial reasoning to bounded constraint satisfaction. Bypassing this contract, as existing 3D systems do by delegating layout to language models, yields overlapping rooms and implausible proportions; directly calling general-purpose language models likewise produces geometrically invalid layouts. Guided by these insights, we present PlanCraft. SketchPlan supplies the missing training signal by replaying the architect's drawing process on 80K real floor plans, producing partial sketches at every completeness level. PlanCraft-Diff progressively sharpens an incomplete sketch into a geometrically precise, vectorizable floor plan through a coarse-to-fine strategy. With the spatial contract established, PlanCraft-Agent then furnishes the scene within well-defined room boundaries. Experiments show that PlanCraft achieves a 61.1\% lower FID than the best existing 2D method and surpasses existing 3D systems by 15 points in expert-rated spatial rationality, with a sketch at only 25\% completion already outperforming all fully specified baselines.

cs.CV

Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic Programming

Constant optimization refines the numerical coefficients of candidate expressions in tree-based genetic programming for symbolic regression. But its per-generation cost has led modern GPU-accelerated frameworks to omit it or restrict it to lightweight forms. We present a GPU-resident, batched Levenberg--Marquardt solver that optimizes constants across a structurally heterogeneous population of expression trees using a fixed number of population-wide CUDA launches per iteration. Reverse-mode automatic differentiation assembles the per-tree Jacobian in one backward sweep, making the dominant per-iteration cost independent of the number of constants per tree, and a double-precision delivery guard guarantees that returned constants are never worse than their initial values. On early-generation populations, the solver sustains up to $5.1{\times}10^{5}$ trees per second on an NVIDIA A100; at a GPU-saturated benchmark configuration it delivers roughly $9.9{\times}$ the throughput of Operon running on a 64-core EPYC 7763, while matching fp64-reference quality. Integrated in-process into EvoGP, the solver enables end-to-end search to recover governing equations on $10$ of $18$ constructed problems versus 0 for stock EvoGP. Our code is at https://github.com/TensorConv/CuSR.

cs.NE

Heat equations in spectral Barron spaces

Spectral Barron spaces, characterized by an \(L^1\)-based Fourier-Lebesgue norm, have earned significant attention in approximation theory due to their remarkable capacity to represent functions via shallow neural networks with controlled complexity. Meanwhile, recent theoretical advances have firmly established an intrinsic and profound connection between these function spaces and the regularity theory of elliptic partial differential equations. Building upon this foundational interplay, the present work undertakes a systematic and comprehensive investigation into the well-posedness of heat equations formulated within the spectral Barron spaces framework. Specifically, we rigorously establish the core aspects of well-posedness, including the existence, uniqueness, and stability of solutions, under suitable assumptions on the source terms and conductivity coefficients. We also investigate a typical parabolic inverse problem, namely the backward heat equation, for which we derive a logarithmic conditional stability estimate. To the best of our knowledge, this constitutes the first stability estimate for inverse problems within the spectral Barron space setting. Moreover, we extend our analytical results to address the more intricate setting of time-fractional heat equations, which govern anomalous diffusion phenomena and introduce nonlocal temporal memory effects. In this extended context, we provide a characterization of the corresponding heat kernels, deriving decay estimates, regularity properties, thereby enriching the theoretical landscape of evolutionary PDEs within the spectral Barron spaces setting.

math.AP

Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores

Convolution is a principal computational bottleneck in deep neural networks, and its efficiency depends on tight integration between algorithms and GPU hardware. Existing GPU convolution methods suffer from large memory overhead, poor cache utilization, limited effectiveness across kernel sizes, or numerical instability. This work extends the im2win paradigm -- a universal, memory-efficient convolution method with contiguous memory access for all kernel sizes -- to run efficiently in full precision on CUDA cores and half precision on tensor cores. By introducing new kernel designs and optimizations such as zig-zag memory access and asynchronous data movement, im2win efficiently exploits hardware-accelerated half-precision matrix multiply-accumulate operations. Across twelve CNN benchmarks, im2win achieves up to 2.8x higher TFLOPS than its CUDA core implementation, 1.4x higher than cuDNN, and 6.4x higher than GEMM-based convolution with cuBLAS, while using as little as 53% and 35% of their memory, respectively. These results establish im2win as a unified, high-performance convolution framework for modern GPU architectures.

cs.DC

Gal3D: Superellipsoid Modeling of Radial 3D Galaxy Structure in IllustrisTNG and EAGLE Simulations

Galaxy morphology and structure are key tracers of galaxy formation and evolution, making accurate measurements of intrinsic three-dimensional (3D) shape essential for linking morphology to galaxy assembly and for comparing numerical simulations. We present Gal3D, a framework that reconstructs smoothed density fields from particle data and quantifies the radial 3D structure of simulated galaxies by fitting superellipsoids to iso-density surfaces. The method recovers axis ratios, orientations, center offsets, and superellipsoid indices ($S_a$, $S_b$, $S_c$), enabling a flexible characterization of diverse galactic structures such as disks, classical bulges, box/peanut bulges, and triaxial components. Applying Gal3D to galaxies in the IllustrisTNG and EAGLE simulations, we find that the radial extent of flattened disk regions increases with stellar mass up to $M_{*,30}\sim10^{11}\,M_\odot$ and then declines sharply, with EAGLE galaxies showing a saturation at $M_{*,30}\sim10^{10.5}\,M_\odot$. The bar-related $ \varepsilon_{ab}\equiv 1-b/a$ strengthens above $M_{*,30}\sim10^{10.5}\,M_\odot$ in both simulations, but remains systematically weaker in EAGLE. In TNG, outer bar regions are commonly associated with elevated $S_a$ and $S_c$, indicating enhanced boxiness and more prominent box/peanut-shaped bulges, whereas such higher-order signatures are weak or absent in EAGLE. At the highest stellar masses, flattened disks become less prominent, while inner prolate or triaxial structures remain common and massive EAGLE galaxies have more prolate or triaxial outer stellar bodies than their TNG counterparts. These results demonstrate that Gal3D provides a practical framework for quantifying intrinsic radial 3D structure and comparing morphology across cosmological simulations.

astro-ph.GA

On Generalized Barron Spaces for Shallow Neural Networks

Classical Barron spaces are function spaces specifically designed for shallow neural networks mostly with ReLU, $\mathrm{ReLU}^k$ (RePU) or Lipschitz continuous activation functions. In the present work, we introduce a generalized Barron space \( B_σ^φ \) for shallow neural networks with a generic activation function possessing certain smoothness properties. The subscript \( σ\) denotes the activation function, while the superscript \( φ\) controls the smoothness of the generalized Barron spaces, defined via a \( φ\)-weighted integral norm imposed on neural network parameter measures. Under certain assumptions on \( φ\) and \( σ\), we show that \( B_σ^φ \) can be continuously embedded into Sobolev spaces. We also explore the relationships among various function spaces for shallow neural networks, demonstrating that our definition encompasses most conventional ones. As applications of the proposed generalized Barron spaces, we derive approximation rates within these spaces and establish error bounds for numerical differentiation with regularization penalized by the newly introduced generalized Barron norm. Numerical examples confirm that the proposed spaces allow the construction of neural networks with varying degrees of smoothness while using the same activation function.

math.NA

Inverse Geometric Diffraction by a Cone

Consider the inverse problem of recovering a strictly convex conical obstacle in $\mathbb{R}^3$ from the diffraction coefficients along with arrival directions (lens data) or arrival times of diffracted waves. The incident wave is a spherical pulse emanating from a point, and the measurements of diffracted waves are taken at an arbitrarily sized receiver placed within the reflection shadow. Specifically, the lens data or arrival times determine the location of the tip, whereas the diffraction coefficients reconstruct the shape of the cone. Since diffraction coefficients are described by half waves over the complement of the cone base in $\mathbb{S}^2$, we reduce inverse diffraction by a cone in $\mathbb{R}^3$ to identifying the reflected wavefront in $\mathbb{S}^2$ and recovering the obstacle using reflected rays on the sphere. The former is accomplished by constructing the Hadamard parametrix for half waves near the wavefront, whereas the latter relies on the topological properties of broken geodesics on $\mathbb{S}^2$. The framework developed in this paper exploits the analytic and geometric structures of diffracted wave fields characterized in the Geometrical Theory of Diffraction, and establishes, for the first time, a rigorous inverse theory corresponding to GTD.

math.AP

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.

cs.CV

CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings

Architectural drawings violate the usual assumption behind multi-view reasoning: plans and sections are cuts, while elevations are facade projections, so corresponding components change appearance in ways camera motion cannot explain. We introduce CrossProjection, an anchor-grounded diagnostic of whether vision-language models preserve component identity and externalize geometry across heterogeneous architectural views. It evaluates Matching, Registration, and Geometric Grounding through categorical judgments, candidate selection, and free point, line, and region localization. Across 23 real drawing sets and 1,954 categorical conditions per model, GPT-5.5 scores 82.4%, Qwen3-VL-32B-Instruct 62.2%, and GLM-4.5V 57.2%. A matched 200-target study crosses natural and vector-text-suppressed drawings with closed-candidate and free-geometry outputs. Candidate-supported performance is often higher, but free localization remains fragile: on natural drawings, point/region PCK@.05 is 54-76% for GPT, 8-10% for Qwen, and 14-36% for GLM; line endpoint PCK@.05 is 22%, 4%, and 0%. A coordinate grid recovers some GPT point/region precision but not lines. Three architecture-trained participants reach 87.3-93.3% categorical accuracy and 76-92% GT-region hit, supporting task feasibility rather than a population-level human ceiling. Because the categorical families do not form a same-item Matching-Registration contrast and interface controls alter multiple burdens, we avoid mechanistic claims. The supported conclusion is narrower: closed-choice or marked-element success does not entail reliable explicit geometric grounding. For drawing-guided CAD/BIM systems, categorical correctness should not be treated as evidence of candidate-free spatial reliability. Reusable on-sheet anchors, fixed-denominator scoring, and hash-locked artifacts establish an audit trail for this gap.

cs.CV

Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge

Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis. However, their clinical deployment is severely hindered by lacking domain-specific knowledge. In this work, we identify two structural deficiencies hindering reliable medical reasoning: 1) the Perception Gap, where general-purpose visual encoders fail to resolve fine-grained pathological cues (e.g., microaneurysms); and 2) the Reasoning Gap, where sparse visual evidence is progressively overridden by massive language priors in deeper transformer layers, leading to ungrounded hallucinations. To bridge these gaps, we propose EyExIn, a data-efficient framework designed to anchor retinal VLMs with expert knowledge via a Deep Expert Injection mechanism. Our architecture employs an Expert-Aware Dual-Stream encoding strategy that decouples visual representation into a general stream for anatomical context and a specialized expert stream for pathological semantics. To ensure high-fidelity integration, we design a Semantic-Adaptive Gated Fusion module, which dynamically amplifies subtle lesion signals while filtering irrelevant background noise. Furthermore, we introduce Adaptive Deep Expert Injection to embed persistent "Vision Anchors" by integrating fused visual features as residual biases directly into intermediate LLM layers. This mechanism creates a visual shortcut that forces the reasoning stack to remain strictly grounded in visual evidence. Extensive experiments across four benchmarks demonstrate that our model consistently outperforms massive proprietary systems. EyExIn significantly enhances domain-specific knowledge embedding and achieves state-of-the-art precision in ophthalmic visual question answering, advancing the development of trustworthy ophthalmic AI.

cs.CV

A mixed residual method for biharmonic equations in spectral Barron spaces

We propose a mixed residual method (MIM) for numerically solving the biharmonic equation with nonhomogeneous clamped boundary conditions. By establishing the well-posedness of the biharmonic equation in spectral Barron spaces, we derive an error bound for MIM that relates shallow neural network approximations to the exact solution and overcomes the curse of dimensionality. This error bound consists of two components: the first corresponds to the approximation error of the neural network, while the second represents the generalization error arising from randomly sampled training data. Several numerical experiments are presented to demonstrate the effectiveness of the proposed method.

math.NA

Inverse Scattering by Diffracted Waves

In addition to reflection and refraction, another form of wave deviation is defined as diffraction. Notably, when incident waves strike a corner, diffracted waves emanate from the corner tip and propagate omnidirectionally. This paper proposes a novel framework for detecting rigid cornered obstacles using measured diffracted wave data. The framework first transforms the underlying initial-boundary value problems into initial value problems on conic manifolds via the method of images. Subsequently, the retrieval of obstacle information is achieved through Cheeger--Taylor functional calculus and microlocal analysis on conic manifolds. Specifically, we prove that for a given pulse, measurements of the resulting diffracted waves captured by a curve receiver uniquely determine both the location and shape of the visible portion of a polygonal obstacle. The proof is constructive, explicitly formulating the corresponding recovery scheme. This methodology offers two key advantages: first, the size and placement of the receiver can be arbitrary; second, the inversion only requires measurements of diffracted waves and obviates the need to solve wave equations within the cornered domain as in conventional methods.

math.AP

An inverse coefficient problem for a semilinear wave equation by the first order linearization

This paper investigates recovery of an unknown coefficient in a semilinear wave equation defined on a bounded, open, and strictly convex domain in \(\mathbb{R}^{1+n}\) with \(n \ge 2\). We demonstrate that the unknown coefficient \(q\) appearing in the semilinear wave equation \(\square u + q u^m = 0\) with Neumann boundary conditions can be reconstructed with Hölder stability from the linearized Neumann-to-Dirichlet (NtD) map. Our approach combines first-order linearization with the Principle of Inclusion-Exclusion (PIE) identity, and employs geometric optics solutions for wave equations in two distinct regimes: the case \(m=2\) with \(q = q(x)\), and the case \(m \ge 3\) with \(q = q(t,x)\). Furthermore, numerical examples illustrate that the unknown coefficient can also be effectively reconstructed using a neural network-based inversion algorithm within the framework of the least squares method.

math.AP

Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions

While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving efficient, scalable, real-time high-resolution video generation a fundamental open challenge. To bridge this gap, we present Ultra Flash, a cascaded streaming framework capable of real-time high-resolution video generation. Ultra Flash achieves ~30 FPS at 1K resolution and ~18 FPS at 2K resolution on a single GPU through three key contributions: (1) an architecture-preserving T2V-to-TV2V super-resolution training paradigm coupled with an AIGC-oriented data degradation pipeline that effectively preserves the generative capability of the base model, enabling enhanced high-resolution detail when cascaded after mainstream low-resolution generative models; (2) a causal streaming latent upsampler paired with a high-resolution decoder, which enhances spatiotemporal coherence while enabling efficient latent spatial scaling and precise high-resolution decoding with negligible computational overhead; and (3) a cascade high-resolution streaming video generation optimization scheme that first performs hybrid-reward-enhanced sparse causalization and single-step distillation of the super-resolution model, then introduces cascaded streaming self-forcing preference optimization with dynamic cache management, jointly enhancing overall coherence, improving quality, and enabling real-time high-resolution streaming video generation. Extensive experiments demonstrate that Ultra Flash reliably produces ultra-high-resolution streaming video while maintaining state-of-the-art visual quality and superior efficiency. Project Page: https://xin1u.github.io/UltraFlash/

cs.CV

Pull Requests as a Training Signal for Repo-Level Code Editing

Repository-level code editing requires models to understand complex dependencies and execute precise multi-file modifications across a large codebase. While recent gains on SWE-bench rely heavily on complex agent scaffolding, it remains unclear how much of this capability can be internalised via high-quality training signals. To address this, we propose Clean Pull Request (Clean-PR), a mid-training paradigm that leverages real-world GitHub pull requests as a training signal for repository-level editing. We introduce a scalable pipeline that converts noisy pull request diffs into Search/Replace edit blocks through reconstruction and validation, resulting in the largest publicly available corpus of 2 million pull requests spanning 12 programming languages. Using this training signal, we perform a mid-training stage followed by an agentless-aligned supervised fine-tuning process with error-driven data augmentation. On SWE-bench, our model significantly outperforms the instruction-tuned baseline, achieving absolute improvements of 13.6% on SWE-bench Lite and 12.3% on SWE-bench Verified. These results demonstrate that repository-level code understanding and editing capabilities can be effectively internalised into model weights under a simplified, agentless protocol, without relying on heavy inference-time scaffolding.

cs.SE

From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents

Supervised fine-tuning (SFT) on long teacher trajectories is the dominant way to instill investigation and reasoning in open software-engineering (SWE) agents. Since every retained response becomes an imitation target, the student inherits the final outcome and intermediate flaws, including ungrounded leaps and redundant loops. High-quality training data must be effective(each step is grounded and narrows the agent's epistemic gap to the correct fix) and efficient(each step is information-bearing rather than redundant or looping). Existing recipes filter or relabel teacher rollouts using only a binary terminal verifier, which does not directly target these axes and provides no supervision on instances where the teacher fails. Most real issue includes a developer-authored reference patch, $p^\star$, revealing the file paths, runtime behaviors, and coding conventions presupposed by the correct fix, yet standard pipelines discard it. We propose Patches-to-Trajectories (P2T), which uses $p^\star$ as privileged information during curation and formulates trajectory construction as bi-objective optimization over per-step effectiveness and trajectory length. A reverse phase distills $p^\star$ into a latent process graph, $G^\star$, of contextual facts and solution milestones. A forward phase curates trajectories from blinded teacher continuations by scoring per-step progress against $G^\star$ under a leakage-blocking groundedness check and retaining the shortest effective segments. Using only 1.8k curated SWE-Gym instances, P2T improves effectiveness and efficiency over outcome-filtered SFT and its tool-error-masking variant. On SWE-bench Verified, it raises Pass@1 by up to 10.8 points while reducing per-instance inference cost by ~15%, with consistent gains on SWE-bench Lite. Size-matched ablations and qualitative analysis further isolate trajectory quality from data scale.

cs.SE