arXiv ScienceSearch

arXiv subjects

Defeng Sun

Publications and source records attributed to Defeng Sun.

At least 19 recordsLinked to original sources

Diffusion Based Unpaired Data Learning for Inverse Problems

Data is important in many deep learning-based inverse problem solvers. However, obtaining sufficient paired data in many scenarios remains highly challenging, while unpaired data is cheap. To maximize data utilization, this paper proposes LUD-DIF, a diffusion-based approach for solving inverse problems with unpaired data. Starting from the evidence lower bound (ELBO) of the joint distribution, we decouple it into two independent diffusion processes under the weak-coupling assumption. The method provides theoretical support from a variational inference perspective, derives the loss function, quantitatively analyzes the error bound introduced by the assumption, and offers a theorem-motivated heuristic for hyperparameter selection. Experimental results demonstrate that LUD-DIF achieves outstanding performance on multiple image inverse problems, validating its effectiveness and generalization capability in unpaired inverse problem settings.

cs.CV

ADMM Fails to Achieve an $O(K^{-1})$ Ergodic KKT Residual Bound

The Karush--Kuhn--Tucker (KKT) residual is a fundamental measure of first-order optimality and, under an error bound condition, is comparable to the distance to the KKT solution set up to constant factors. Despite the $O(K^{-1})$ ergodic rates known for objective error and feasibility violations, we show that the KKT residual of classical ADMM cannot, in general, satisfy a uniform $O(K^{-1})$ bound. Specifically, we construct a fixed-dimensional, horizon-dependent family of two-block convex optimization problems for which the KKT residual is $\Omega(K^{-1/2})$ at both the last iterate and the equal-weight ergodic average at the prescribed horizon $K$. Consequently, a uniform $O(K^{-1})$ KKT residual bound is impossible for either output.

math.OC

Robust mean field control: an application to optimal execution under composite uncertainty

We provide a framework for robust mean field control problems that describe multi-dimensional optimal liquidation problems under uncertainty from both the underlying stochastic process and the deterministic model parameters. The verification results are established with Hamilton-Jacobi-Bellman-Isaacs (HJBI) equations where the variables are probability measures and the Hamiltonian nonlinearly involves the joint distribution of position and momentum. Using novel a priori estimates, we establish the well-posedness of the HJBI equations featuring general or quadratic Hamiltonians that are neither displacement convex nor concave in their momentum. The a priori estimates and well-posedness results are extended during their application to optimal liquidation problems, where we allow the Hamiltonian to have derivatives of linear growth and solve the constrained multi-dimensional linear quadratic optimal liquidation problem under composite uncertainty.

math.OC

Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width

In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width parameter $N$ and the depth parameter $L$, thereby granting greater architectural flexibility. Existing works using the $(N,L)$-characterization focus on function classes with finite smoothness $s$, establishing a typical approximation rate of $\mathcal{O}\left(N^{-2s/d}L^{-2s/d}\right)$ with $d$ denoting the input dimension, which indicates that network depth and width play symmetric roles for these classes. In contrast, this paper establishes upper bounds for the approximation of analytic functions, which possess infinite smoothness, via ReLU networks under the $(N,L)$-characterization. Specifically, we derive approximation rates of $\mathcal{O}\left(N^{-C L^{\tau}}\right)$, where $C>0$ is some constant and $\tau>0$ is a parameter influenced by the relation between $L$ and $N$. In particular, $\tau=1$ if $N$ scales roughly as $L^d$. Our findings reveal that depth plays a more critical role than width in the context of analytic function approximation. The main technical difficulty of obtaining such upper bounds lies in the trade-off between the smoothness parameters and the approximation accuracy. To overcome this difficulty, we employ refined constructions of several ReLU networks to approximate power functions, multivariate multiplication, and polynomials, which may be of independent interest.

stat.ML

Maximal monotonicity of piecewise polyhedral mappings

Maximal monotone mappings that are piecewise polyhedral arise from subdifferentials of convex functions and saddle functions that are piecewise linear-quadratic and enter into algorithmic constructions of importance in linear-quadratic optimization and associated splitting methods. The question of whether those constructions preserve maximal monotonicity is then crucial. The usual answers to that invoke constraint qualifications involving the nonemptiness of intersections of certain relative interiors, but it is shown here that the piecewise polyhedral structure allows the relative interiors to be bypassed.

math.OC

Convergence Analysis of the Restarted Moving-Anchored Extra-Gradient Method in the Absence of Local Lipschitz Continuity

In this paper, we introduce the moving-anchored extra-gradient (MAEG) method for solving monotone inclusion problems involving the sum of a continuous monotone operator and a maximal monotone operator. Notably, the distance from the anchor point to the solution set is designed to be monotonically non-increasing. Under Lipschitz continuity of the forward operator, MAEG attains an $\mathcal{O}(1/k)$ non-asymptotic iteration complexity, and when a positive anchor-update parameter is used, it further achieves an $o(1/k)$ asymptotic rate. Furthermore, leveraging the specific behavior of the anchor point, we propose a tailored restart strategy. We demonstrate that this strategy ensures convergence even in the absence of local Lipschitz continuity, while preserving the original iteration complexity guarantees whenever the Lipschitz condition holds.

math.OC

A Semismooth Newton Augmented Lagrangian Method for Sparse Spectral Risk Optimization

Empirical risk minimization is a standard and effective paradigm for learning predictive models by minimizing average loss. In high-stakes decision-making, however, an average-loss criterion may underrepresent rare but severe losses. Spectral risk measures (SRMs) provide a principled framework by incorporating weighted order statistics of losses, but the induced nonsmoothness and nonseparability from sorting make the resulting optimization problems challenging. We propose a relative inexact proximal augmented Lagrangian method with a semismooth Newton subproblem solver for solving SRM-based optimization problems. Exploiting a dual reformulation and properties of the Moreau envelope, we reduce the subproblems to structured dual-variable formulations, significantly simplifying computation. We provide explicit generalized Jacobian characterizations and tailor the pool adjacent violators algorithm for their efficient evaluation. Numerical results on synthetic and real-data instances show that the proposed method attains lower running times than the tested ADMM baseline while producing comparable stationarity residuals and sparse solutions.

math.OC

Factorized low-rank matrix recovery problem, Schatten-$q$ quasi-norm, Error bound for critical point, Kurdyka-\L ojasiewicz property, Inexact proximal alternating linearized minimization

The Schatten-$q$ quasi-norm is a widely used nonconvex rank surrogate and matrix factorization is an effective approach to reduce computational cost. In this paper, we consider the equivalent group-sparse factorized reformulation of Schatten-$q$ norm regularized low-rank matrix recovery problem. Though this factorized model exhibits favorable performance, two issues remain: (i) the error bound of critical points is unexplored; (ii) the proximal operator of $\|\cdot\|_2^q$ lacks a closed-form solution for general $q$, limiting algorithms to adopt fixed $q$ like $1/2$ or $2/3$. This paper addresses both issues. We investigate the properties of critical points for the factorized problem and show that, compared to nuclear norm, the Schatten-$q$ norm implicitly endows critical points with column orthogonality. From this insight, we introduce the notion of S-critical points under mild conditions that ensure column orthogonality with easily operable criterion for identifying. We show that global optimal points must be S-critical points and we derive an error bound between S-critical points and the true matrix. We further present an inexact proximal alternating linearized minimization method for the factorized problem, along with practically computable inexact proximal operator for $\|\cdot\|_2^q$ and criteria to find solutions satisfying inexactness conditions, and we establish the whole sequence convergence and a convergence rate guarantee under Kurdyka--\L ojasiewicz condition. Moreover, we prove that the factorized model with least-squares loss has KL exponent $1/2$ at S-critical points, then the iteration converges linearly under suitable condition. Extensive numerical experiments validate the effectiveness of our algorithm and confirm the theoretical properties of the factorized model.

math.OC

Low-Rank Tensor Completion using Tensor Train Decomposition via Riemannian Optimization on the Quotient Geometry

Owing to the effectiveness of Tensor Train (TT) decomposition in managing high-order tensors, low-rank tensor completion within the TT-format has emerged as a prominent research focus. In this paper, we leverage the left-orthogonal property of the TT-decomposition to construct a novel quotient manifold and introduce a family of admissible Riemannian metrics. Within this geometric framework, we propose a new approach to constructing retractions compatible with the quotient structure, realized via two novel retractions based on recursive polar and QR decompositions that respect the recursive orthogonalization structure of the TT format. We then derive Riemannian gradient descent and conjugate gradient methods to solve the tensor completion problem. Theoretically, our approach streamlines the horizontal projection by reducing the number of unknowns per block from a quadratic dependence on the TT-ranks to a near-half scaling, thereby enhancing computational efficiency over conventional quotient-based methods. Numerical experiments demonstrate that the proposed algorithms achieve reconstruction accuracy comparable to state-of-the-art TT-based geometric methods.

math.NA

A symmetric Gauss-Seidel based alternating proximal ALM for generalized Nash Equilibrium problems in Banach spaces

In this paper, we study a class of monotone generalized Nash equilibrium problems (GNEPs) with jointly linear constraints. The players' strategy spaces are real Hilbert spaces, while the joint constraint is formulated in a Banach space. To solve such problems, we propose a novel symmetric Gauss-Seidel (sGS) based alternating proximal augmented Lagrangian method (sGS-APALM) which incorporates newly designed quadratic surrogates. In contrast to existing regularization and ALM-type methods, the proposed method avoids solving coupled Nash equilibrium subproblems at each iteration and instead updates the players' strategies alternately by solving a sequence of unconstrained quadratic programs. Moreover, unlike many existing splitting-based methods, our global convergence analysis and convergence rate estimation require only monotonicity and Lipschitz continuity of the pseudo-gradient mapping, without imposing stronger assumptions such as strong monotonicity or co-coercivity. Finally, we apply the method to a class of risk-neutral PDE-constrained GNEPs with joint state constraints, and preliminary numerical results demonstrate its efficiency and effectiveness.

math.OC

A Level Set Method with Secant Iterations for the Least-Squares Constrained Nuclear Norm Minimization

We present an efficient algorithm for least-squares constrained nuclear norm minimization, a computationally challenging problem with broad applications. Our approach combines a level set method with secant iterations and a proximal generation method. As a key theoretical contribution, we establish the nonsingularity of the Clarke generalized Jacobian for a general class of projection norm functions over closed convex sets. This property and the (strong) semismoothness of our value function yield fast local convergence of the secant method. For the resulting nuclear norm regularized subproblems, we develop a proximal generation method that exploits low-rank structures without compromising convergence. Extensive numerical experiments demonstrate the superior performance of our approach compared to state-of-the-art methods.

math.OC

DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrieval

Large Language Model (LLM) agents can automate data-science workflows, but many rigorous statistical methods implemented in R remain underused because LLMs struggle with statistical knowledge and tool retrieval. Existing retrieval-augmented approaches focus on function-level semantics and ignore data distribution, producing suboptimal matches. We propose DARE (Distribution-Aware Retrieval Embedding), a lightweight, plug-and-play retrieval model that incorporates data distribution information into function representations for R package retrieval. Our main contributions are: (i) RPKB, a curated R Package Knowledge Base derived from 8,191 high-quality CRAN packages; (ii) DARE, an embedding model that fuses distributional features with function metadata to improve retrieval relevance; and (iii) RCodingAgent, an R-oriented LLM agent for reliable R code generation and a suite of statistical analysis tasks for systematically evaluating LLM agents in realistic analytical scenarios. Empirically, DARE achieves an NDCG at 10 of 93.47%, outperforming state-of-the-art open-source embedding models by up to 17% on package retrieval while using substantially fewer parameters. Integrating DARE into RCodingAgent yields significant gains on downstream analysis tasks. This work helps narrow the gap between LLM automation and the mature R statistical ecosystem.

cs.IR

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets

The tremendous success of Transformer models in fields such as large language models and computer vision necessitates a rigorous theoretical investigation. To the best of our knowledge, this paper is the first work proving that standard Transformers can approximate H\"older functions $ C^{s,\lambda}\left([0,1]^{d\times n}\right) $$ (s\in\mathbb{N}_{\geq0},0<\lambda\leq1) $ under the $L^t$ distance ($t \in [1, \infty]$) with arbitrary precision. Building upon this approximation result, we demonstrate that standard Transformers achieve the minimax optimal rate in nonparametric regression for H\"older target functions. It is worth mentioning that, by introducing two metrics: the size tuple and the dimension vector, we provide a fine-grained characterization of Transformer structures, which facilitates future research on the generalization and optimization errors of Transformers with different structures. As intermediate results, we also derive the upper bounds for the Lipschitz constant of standard Transformers and their memorization capacity, which may be of independent interest. These findings provide theoretical justification for the powerful capabilities of Transformer models.

stat.ML

Beyond the Prompt in Large Language Models: Comprehension, In-Context Learning, and Chain-of-Thought

Large Language Models (LLMs) have demonstrated remarkable proficiency across diverse tasks, exhibiting emergent properties such as semantic prompt comprehension, In-Context Learning (ICL), and Chain-of-Thought (CoT) reasoning. Despite their empirical success, the theoretical mechanisms driving these phenomena remain poorly understood. This study dives into the foundations of these observations by addressing three critical questions: (1) How do LLMs accurately decode prompt semantics despite being trained solely on a next-token prediction objective? (2) Through what mechanism does ICL facilitate performance gains without explicit parameter updates? and (3) Why do intermediate reasoning steps in CoT prompting effectively unlock capabilities for complex, multi-step problems? Our results demonstrate that, through the autoregressive process, LLMs are capable of exactly inferring the transition probabilities between tokens across distinct tasks using provided prompts. We show that ICL enhances performance by reducing prompt ambiguity and facilitating posterior concentration on the intended task. Furthermore, we find that CoT prompting activates the model's capacity for task decomposition, breaking complex problems into a sequence of simpler sub-tasks that the model has mastered during the pretraining phase. By comparing their individual error bounds, we provide novel theoretical insights into the statistical superiority of advanced prompt engineering techniques.

cs.CL

DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems

Recent LLM-based data agents aim to automate data science tasks ranging from data analysis to deep learning. However, the open-ended nature of real-world data science problems, which often span multiple taxonomies and lack standard answers, poses a significant challenge for evaluation. To address this, we introduce DSAEval, a benchmark comprising 641 real-world data science problems grounded in 285 diverse datasets, covering both structured and unstructured data (e.g., image and text). DSAEval incorporates three distinctive features: (1) Multimodal Environment Perception, which enables agents to interpret observations from multiple modalities, including text and vision; (2) Multi-Query Interactions, which mirror the iterative and cumulative nature of real-world data science projects; and (3) Multi-Dimensional Evaluation, which provides a holistic assessment across reasoning, code, and results. We systematically evaluate 13 recent advanced agentic LLMs using DSAEval. Our results show that Claude-Sonnet-4.5 achieves the strongest overall performance, MiMo-V2-Pro and GPT-5.2 lead in duration and step efficiency, respectively, and MiMo-V2-Flash is the most cost-effective. We further demonstrate that multimodal perception consistently improves performance on vision-related tasks, with gains ranging from 2.04% to 11.30%. Overall, while current data science agents perform well on structured data and routine data analysis workflows, substantial challenges remain in unstructured domains. Finally, we offer critical insights and outline future research directions.

cs.AI

The global well-posedness for master equations of mean field games of controls

In this manuscript, we establish the global well-posedness for master equations of mean field games of controls, where the interaction is through the joint law of the state and control. Our results are proved under two different conditions: the Lasry-Lions monotonicity and the displacement $\lambda$-monotonicity, both considered in their integral forms. We provide a detailed analysis of both the differential and integral versions of these monotonicity conditions for the corresponding nonseparable Hamiltonian and examine their relation. The proof of global well-posedness relies on the propagation of these monotonicity conditions in their integral forms and a priori uniform Lipschitz continuity of the solution with respect to the measure variable.

math.PR

dHPR: A Distributed Halpern Peaceman--Rachford Method for Non-smooth Distributed Optimization Problems

This paper introduces the distributed Halpern Peaceman--Rachford (dHPR) method, an efficient algorithm for solving distributed convex composite optimization problems with non-smooth objectives, which achieves a non-ergodic $O(1/k)$ iteration complexity regarding Karush--Kuhn--Tucker residual. By leveraging the symmetric Gauss--Seidel decomposition, the dHPR effectively decouples the linear operators in the objective functions and consensus constraints while maintaining parallelizability and avoiding additional large proximal terms, leading to a decentralized implementation with provably fast convergence. The superior performance of dHPR is demonstrated through comprehensive numerical experiments on distributed LASSO, group LASSO, and $L_1$-regularized logistic regression problems.

math.OC

Progressive Bound Strengthening via Doubly Nonnegative Cutting Planes for Nonconvex Quadratic Programs

We introduce a cutting-plane framework for nonconvex quadratic programs (QPs) that progressively tightens convex relaxations. Our approach leverages the doubly nonnegative (DNN) relaxation to compute strong lower bounds and generate separating cuts, which are iteratively added to improve the relaxation. We establish that, at any Karush-Kuhn-Tucker (KKT) point satisfying a second-order sufficient condition, a valid cut can be obtained by solving a linear semidefinite program (SDP), and we devise a finite-termination local search procedure to identify such points. Extensive computational experiments on both benchmark and synthetic instances demonstrate that our approach yields tighter bounds and consistently outperforms leading commercial and academic solvers in terms of efficiency, robustness, and scalability. Notably, on a standard desktop, our algorithm reduces the relative optimality gap to 0.01% on 138 out of 140 instances of dimension 100 within one hour, without resorting to branch-and-bound.

math.OC