arXiv ScienceSearch

arXiv subjects

Yuwen Li

Publications and source records attributed to Yuwen Li.

At least 19 recordsLinked to original sources

Shallow neural network approximation in mixed Sobolev spaces

We investigate the best $L_2$ approximation of mixed Sobolev spaces by shallow neural networks with $n$ neurons and general activation functions. We first establish an activation-independent Fourier-block principle: if an activation has univariate approximation order $\rho$ in the sense of the Fourier-block property, then the global approximation rate has algebraic order $\min\{\alpha,\rho\}$ for target functions of mixed smoothness $\alpha$, up to explicit logarithmic factors. To verify this property for concrete activations, we introduce a structured univariate approximation condition that implies the Fourier-block property with explicit parameters. For $\mathrm{ReLU}^k$, a matching algebraic lower bound identifies $\min\{\alpha,k+1\}$ as the optimal algebraic approximation exponent in any dimension, up to logarithmic factors in the upper bound. The framework also yields the exponent $\min\{\alpha,k+1\}$ for cardinal B-splines and soft-$\mathrm{ReLU}^k$, and the full mixed-smoothness exponent $\alpha$ for ELU and cosine activations, again up to logarithmic~factors.

math.NA

Closure complexity of longest-edge bisection for triangular meshes

On triangular meshes, we analyze local mesh refinement based on longest-edge bisection equipped with the serial longest-edge propagation-path closure. Ties are resolved by terminal priority: if the incoming shared edge is a longest edge of the neighboring triangle, the pair is declared terminal and that edge is bisected. For every adaptive mesh sequence $\mathcal{T}_0, \mathcal{T}_1, \ldots, \mathcal{T}_L$ with a sequence of marked sets $\mathcal{M}_0, \mathcal{M}_1, \ldots, \mathcal{M}_{L-1}$, we prove the cumulative closure estimate $$\#\mathcal{T}_L-\#\mathcal{T}_0 \lesssim\sum_{\ell=0}^{L-1}\#\mathcal{M}_\ell.$$ The proof derives single-mark locality from the uniform multiplicative gap between possible descendant diameters implied by finite similarity classes, and converts this locality into the cumulative estimate through a Binev--Dahmen--DeVore-type charging argument.

math.NA

Sharp embeddings between quasi-Banach Besov spaces and shallow ReLU variation spaces

Let $\mathcal D$ be the normalized ridge dictionary generated by $\operatorname{ReLU}^k$ on a bounded Lipschitz domain $\Omega\subset\mathbb R^d$. We establish sharp embeddings between isotropic Besov spaces and the associated variation space $\mathcal L_1(\mathcal D)$ in the quasi-Banach range $0 k+d/p$ for $1<q\le\infty$. A rescaled-bump construction shows that this smoothness threshold is sharp. Conversely, for $0<p<1$, \[ \mathcal L_1(\mathcal D)\hookrightarrow B^{k+1}_{p,2}(\Omega), \] and both the smoothness $k+1$ and the fine index $2$ are optimal. The forward embedding converts known Besov regularity, in particular for solutions of partial differential equations, into controlled approximation error bounds and convergence of greedy algorithms based on shallow $\text{ReLU}^k$ neural networks. The proofs combine Littlewood--Paley localization, Fourier--Radon representations, measure-valued derivatives, and vector-valued singular-integral estimates.

math.FA

Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reasoning Across Diseases and Populations

Cardiologists interpret electrocardiograms by localizing waveform components, measuring rhythm and interval patterns, and translating these structured observations into diagnostic evidence. Whether this expert reading process can serve as an effective prior for ECG agents remains unclear. To address this question, we introduce LuminaECG, a clinically structured ECG reasoning framework that reformulates ECG interpretation as measurement-grounded visual reading. ECG signals are rendered on standard electrocardiographic grid paper to preserve the spatial and scale cues used in clinical reading. P-wave, QRS-complex, and T-wave boundaries are explicitly delineated, and color-coded segmentation decomposes the waveform into discrete visual measurement primitives. A general 2B vision-language backbone is then trained with low-rank supervised fine-tuning to associate these primitives with diagnostic reasoning, without architectural modification. Across open, proprietary, and ECG-specialist zero-shot baselines, LuminaECG improves both waveform measurement and diagnostic recovery. It reaches a clinically meaningful reader tier on the CODE-test benchmark, transfers across geographically diverse ECG datasets without retraining, and generates reports whose structure contains an emergent prognostic signal. These findings suggest that effective ECG agents require not only larger models, but supervision that preserves the alignment between measurable waveform evidence and clinical knowledge.

eess.SP

Nearly optimal Kolmogorov widths under holomorphic mappings

This paper establishes essentially optimal asymptotic bounds for Kolmogorov widths under holomorphic mappings between complex Banach spaces. Given a compact parameter set whose Kolmogorov widths decay algebraically with rate s, we prove that the widths of its image under a holomorphic mapping decay algebraically with every rate t<s, thereby answering an open question raised by Cohen and DeVore. As an application, we obtain a sharp characterization of the approximability of solution manifolds associated with inf-sup stable parametrized PDEs. We also construct an explicit example showing that the arbitrarily small loss in the algebraic decay exponent is unavoidable. Finally, we provide similar characterization of asymptotic bounds for Kolmogorov widths in the exponentially decaying regime. Our analysis involves multilinear Taylor expansion in Banach spaces and a novel block dyadic expansion-truncation technique.

math.FA

Localized pointwise a posteriori error estimates for nonconforming finite element methods

This paper establishes the first localized pointwise a posteriori error estimates for nonconforming finite element discretizations of the Poisson and biharmonic equations. For the Poisson problem, we derive localized estimates for the function-value and broken gradient errors of the Crouzeix--Raviart method. For the Morley discretization of the biharmonic equation, we derive a localized a posteriori estimate that controls the local Hessian error. This provides the first pointwise a posteriori error analysis for the biharmonic equation, for either conforming or nonconforming finite element methods. The key ingredient, absent from pointwise analysis of second order PDEs, is the design of two novel weight functions that facilitate sharp estimates of the $L^1$ norms of derivatives of a regularized Green's function for the biharmonic operator.

math.NA

Evaluation Protocols and Cross-Subject Generalization in EEG Emotion Recognition

Reported accuracy in electroencephalography (EEG) emotion recognition depends on the complete evaluation procedure, not only the classifier. We separate the target quantity, development procedure, and reporting rule, then use one archived dynamical graph convolutional neural network (DGCNN) pathway on SEED and SEED-IV as an illustrative case. In a protocol-matched subject-dependent check, the SEED result was within 1.47 percentage points of the public reference value; the 3.40-point SEED-IV difference remained unresolved. Across 30 matched SEED subject-session trajectories, checkpoint selection based on repeated test-set evaluation increased mean window accuracy from 0.7855 at epoch 80 to 0.8892. Under five-fold subject-disjoint evaluation, validation-selected checkpoints achieved training-participant trial accuracies of 0.9990 on SEED and 0.9920 on SEED-IV. Accuracy for entirely held-out participants was 0.5348 (95% conditional subject-level bias-corrected and accelerated [BCa] interval [0.4667, 0.5985]) on SEED. The SEED-IV estimate was 0.3954 ([0.3343, 0.4648]) and is reported only as secondary sensitivity evidence because its protocol-matched compatibility check remained unresolved. The observed train-to-held-out-subject gaps are inconsistent with simple optimization underfitting, but they do not isolate subject identity from implementation, preprocessing, representation, or distributional factors. Supporting analyses further showed that participant rankings depended on representation and time scale, while a development-selected tail-risk ensemble did not establish a positive gain in a separate final evaluation. Subject-dependent, subject-disjoint, and cross-session results should therefore be reported as answers to different questions.

cs.LG

Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment

Adult and pediatric electrocardiogram (ECG) interpretation relies on age-sensitive criteria, and models pretrained mainly on adult ECGs often transfer poorly to pediatric populations when pediatric labels are scarce. Existing multimodal ECG--text methods typically align waveforms and text at the global sample level, entangling evidence from co-occurring diagnoses and limiting transfer under this gap. We propose Pediatric-Adult ECG Alignment via Cross-modal Enhancement (PEACE), a knowledge-guided framework pretrained on the largely adult MIMIC-IV ECG corpus. PEACE describes each diagnosis along rhythm, morphology, and ST--T axes and, per recording, composes only positive-label descriptors into three axis tokens and a fused embedding. A label query network (LQN) uses diagnostic labels as queries to cross-attend over ECG tokens and axis tokens, while label set aware bidirectional contrastive learning (LSBC) aligns pooled ECG features with the fused embedding when recordings share diagnoses. Curriculum adaptive fusion (CAF) gates alignment strength according to smoothed classification loss and training progress, limiting disruption during early optimization. The knowledge branch is used only for training supervision; inference uses ECG signals alone. On ZZU-pECG, PEACE reaches macro average AUCs of 59.39%, 81.74%, and 91.56% under zero-shot, 50-shot, and full fine-tuning, with the clearest gains over foundation and knowledge-pretraining baselines under limited supervision; versus domain adaptation initializations, zero-shot improves substantially while 50-shot AUC is comparable to DANN. After fine-tuning on PTB-XL, PEACE reaches 96.90% macro average AUC over nine harmonized labels. Ablations confirm that label-conditioned knowledge alignment, rather than global text fusion, is the key driver of pediatric transfer gains.

cs.LG

Superconvergence in finite element method by smoothing

This paper develops a smoothing-based postprocessing method for superconvergence in finite element methods. The method applies a few smoothing iterations, such as damped Jacobi, Gauss-Seidel, or conjugate gradient, with initial guess being the current finite element solution embedded in an enriched finite element space. The resulting procedure is algebraic, easy to implement, and applicable to high-order and three-dimensional discretizations. For symmetric and positive-definite problems, we prove superconvergence of the smoothed solutions under additive and multiplicative smoothers. Effectiveness of the proposed method is demonstrated by numerical experiments for the Poisson, Maxwell, biharmonic and Helmholtz equations.

math.NA

An Adaptive Finite Element Method Based on Generalized Barycentric Coordinates

This work derives a posteriori error estimate of polygonal finite element methods based on Wachspress barycentric coordinates. In particular, we prove that the classical residual-based a posteriori error estimator is both an upper and lower bounds for the discretization error. The analysis relies a Scott-Zhang type interpolation and homogeneity arguments for rational functions on polygonal elements. Numerical experiments on square and L-shaped domains demonstrate the effectiveness of the adaptive algorithm.

math.NA

Label-Conditioned Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Curriculum-Gated Contrastive Alignment

Automated pediatric electrocardiogram (ECG) interpretation remains challenging because developmental differences in heart rate, intervals, and waveforms limit the transferability of models trained mainly on adult data, while expert-labeled pediatric ECG cohorts are scarce. We propose PEACE (Pediatric-Adult ECG Alignment via Cross-modal Enhancement), an adult-to-pediatric ECG transfer framework pretrained on MIMIC-IV ECGs and adapted to pediatric targets. PEACE integrates label-specific bidirectional contrastive learning (LSBC) to align ECG representations with diagnostic semantics and curriculum adaptive fusion (CAF) to stabilize optimization under limited pediatric supervision. Label-conditioned short text descriptors provide auxiliary semantic supervision during training, whereas inference requires ECG signals only. On ZZU-pECG, PEACE achieves macro-average AUCs of 59.39%, 81.74%, and 91.56% under zero-shot, 50-shot, and full fine-tuning settings, respectively, outperforming ECG-only, multimodal, and generic domain adaptation baselines including DANN and MMD. On PTB-XL, it reaches 96.90% macro-average AUC after full fine-tuning over nine harmonized labels with nonzero mapped incidence. Gradient-based attention maps show increased saliency around QRS voltage and morphology regions for chamber-related RVH and around QRS-to-T/repolarization intervals for LQTS, broadly consistent with ECG regions commonly inspected during routine interpretation. These results suggest that adult-scale ECG pretraining coupled with rhythm, morphology, and ST-T repolarization semantic descriptors improves transferable pediatric diagnosis under label scarcity while preserving clinically interpretable waveform focus.

cs.LG

Higher Order Approximation Rates for ReLU CNNs in Korobov Spaces

This paper investigates the $L_p$ approximation error for higher order Korobov functions using deep convolutional neural networks (CNNs) with ReLU activation. For target functions having a mixed derivative of order m+1 in each direction, we improve classical approximation rate of second order to (m+1)-th order (modulo a logarithmic factor) in terms of the depth of CNNs. The key ingredient in our analysis is approximate representation of high-order sparse grid basis functions by CNNs. The results suggest that higher order expressivity of CNNs does not severely suffer from the curse of dimensionality.

cs.LG

IQuest-Coder-V1 Technical Report

In this report, we introduce the IQuest-Coder-V1 series-(7B/14B/40B/40B-Loop), a new family of code large language models (LLMs). Moving beyond static code representations, we propose the code-flow multi-stage training paradigm, which captures the dynamic evolution of software logic through different phases of the pipeline. Our models are developed through the evolutionary pipeline, starting with the initial pre-training consisting of code facts, repository, and completion data. Following that, we implement a specialized mid-training stage that integrates reasoning and agentic trajectories in 32k-context and repository-scale in 128k-context to forge deep logical foundations. The models are then finalized with post-training of specialized coding capabilities, which is bifurcated into two specialized paths: the thinking path (utilizing reasoning-driven RL) and the instruct path (optimized for general assistance). IQuest-Coder-V1 achieves state-of-the-art performance among competitive models across critical dimensions of code intelligence: agentic software engineering, competitive programming, and complex tool use. To address deployment constraints, the IQuest-Coder-V1-Loop variant introduces a recurrent mechanism designed to optimize the trade-off between model capacity and deployment footprint, offering an architecturally enhanced path for efficacy-efficiency trade-off. We believe the release of the IQuest-Coder-V1 series, including the complete white-box chain of checkpoints from pre-training bases to the final thinking and instruction models, will advance research in autonomous code intelligence and real-world agentic systems.

cs.AI

Smoother-type a posteriori error estimates for finite element methods

This work develops user-friendly a posteriori error estimates of finite element methods, based on smoothers of linear iterative solvers. The proposed method employs simple smoothers, such as Jacobi or Gauss-Seidel iteration, on an auxiliary finer mesh to process the finite element residual for a posteriori error control. The implementation has linear complexity and requires only a coarse-to-fine prolongation operator. For symmetric problems, we prove the reliability and efficiency of smoother-type error estimators under a saturation assumption. Numerical experiments for various PDEs demonstrate that the proposed smoother-type error estimators outperform residual-type estimators in accuracy and exhibit robustness with respect to parameters and polynomial degrees.

math.NA

Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing

Enabling Large Language Models (LLMs) to reliably invoke external tools remains a critical bottleneck for autonomous agents. Existing approaches suffer from three fundamental challenges: expensive human annotation for high-quality trajectories, poor generalization to unseen tools, and quality ceilings inherent in single-model synthesis that perpetuate biases and coverage gaps. We introduce InfTool, a fully autonomous framework that breaks these barriers through self-evolving multi-agent synthesis. Given only raw API specifications, InfTool orchestrates three collaborative agents (User Simulator, Tool-Calling Assistant, and MCP Server) to generate diverse, verified trajectories spanning single-turn calls to complex multi-step workflows. The framework establishes a closed loop: synthesized data trains the model via Group Relative Policy Optimization (GRPO) with gated rewards, the improved model generates higher-quality data targeting capability gaps, and this cycle iterates without human intervention. Experiments on the Berkeley Function-Calling Leaderboard (BFCL) demonstrate that InfTool transforms a base 32B model from 19.8% to 70.9% accuracy (+258%), surpassing models 10x larger and rivaling Claude-Opus, and entirely from synthetic data without human annotation.

cs.CL

Some p-robust a posteriori error estimates based on auxiliary spaces

This work develops polynomial-degree-robust (p-robust) equilibrated a posteriori error estimates for $H(\rm curl)$, $H(\rm div)$ and $H(\rm divdiv)$ problems, based on $H^1$ auxiliary space decomposition. The proposed framework employs auxiliary space preconditioning and regular decompositions to decompose the finite element residual into $H^{-1}$ residuals that are further controlled by classical p-robust equilibrated a posteriori error analysis. As a result, we obtain novel and simple p-robust a posteriori error estimates of $H(\rm curl)$/$H(\rm div)$ conforming methods and mixed methods for the biharmonic equation. In addition, we prove guaranteed a posteriori upper error bounds under convex domains or certain boundary conditions. Numerical experiments demonstrate the effectiveness and p-robustness of the proposed error estimators for the Nédélec edge element methods and the Hellan--Herrmann--Johnson methods.

math.NA

Masked Autoencoders that Feel the Heart: Unveiling Simplicity Bias for ECG Analyses

The diagnostic value of electrocardiogram (ECG) lies in its dynamic characteristics, ranging from rhythm fluctuations to subtle waveform deformations that evolve across time and frequency domains. However, supervised ECG models tend to overfit dominant and repetitive patterns, overlooking fine-grained but clinically critical cues, a phenomenon known as Simplicity Bias (SB), where models favor easily learnable signals over subtle but informative ones. In this work, we first empirically demonstrate the presence of SB in ECG analyses and its negative impact on diagnostic performance, while simultaneously discovering that self-supervised learning (SSL) can alleviate it, providing a promising direction for tackling the bias. Following the SSL paradigm, we propose a novel method comprising two key components: 1) Temporal-Frequency aware Filters to capture temporal-frequency features reflecting the dynamic characteristics of ECG signals, and 2) building on this, Multi-Grained Prototype Reconstruction for coarse and fine representation learning across dual domains, further mitigating SB. To advance SSL in ECG analyses, we curate a large-scale multi-site ECG dataset with 1.53 million recordings from over 300 clinical centers. Experiments on three downstream tasks across six ECG datasets demonstrate that our method effectively reduces SB and achieves state-of-the-art performance.

eess.SP

Some Super-approximation Rates of ReLU Neural Networks for Korobov Functions

This paper examines the $L_p$ and $W^1_p$ norm approximation errors of ReLU neural networks for Korobov functions. In terms of network width and depth, we derive nearly optimal super-approximation error bounds of order $2m$ in the $L_p$ norm and order $2m-2$ in the $W^1_p$ norm, for target functions with $L_p$ mixed derivative of order $m$ in each direction. The analysis leverages sparse grid finite elements and the bit extraction technique. Our results improve upon classical lowest order $L_\infty$ and $H^1$ norm error bounds and demonstrate that the expressivity of neural networks is largely unaffected by the curse of dimensionality.

cs.LG