arXiv ScienceSearch

arXiv subjects

Mario Ullrich

Publications and source records attributed to Mario Ullrich.

At least 19 recordsLinked to original sources

On bounds between all s-numbers and widths of convex sets

We prove $a_n(S) \le e\,(n+1)\, s_n(S)$ for every s-number sequence $(s_n)$, every bounded linear operator $S$ between normed spaces, and every $n \in \mathbb{N}_0$, where $a_n$ are the approximation numbers, which are the largest s-numbers. This is sharp up to the constant and settles conjectures of Mityagin, Henkin, Carl and Pietsch dating back to 1963. We also extend it to widths of convex sets and discuss optimality there. The proof is elementary.

math.FA

Approximation of Functions: Optimal Sampling and Complexity

We consider approximation or recovery of functions based on a finite number of function evaluations. This is a well-studied problem in optimal recovery, machine learning, and numerical analysis in general, but many fundamental insights were obtained only recently. We discuss different aspects of the information-theoretic limit that appears because of the limited amount of data available, as well as algorithms and sampling strategies that come as close to it as possible. We also discuss (optimal) sampling in a broader sense, allowing other types of measurements that may be nonlinear, adaptive and random, and present several relations between the different settings in the spirit of information-based complexity. We hope that this article provides both, a basic introduction to the subject and a contemporary summary of the current state of research.

math.NA

Constructive discretization and approximation in reproducing kernel Hilbert spaces

We generalize the sparsification algorithm of Batson, Spielman and Srivastava, making one part of the result dimension-independent. In particular, we recover discretization inequalities in $L_2$- and sup-norms on general finite-dimensional subspaces, prove a suitable infinite-dimensional variant, and discuss the implications for the error of least-squares approximation based on samples. This gives a more constructive version of several recently established approximation bounds, some of which relied on the stronger and less constructive result of Marcus, Spielman and Srivastava. We also improve the constants and oversampling factors in these results.

math.NA

Sampling recovery in $L_2$ and other norms

We study the recovery of functions in various norms, including $L_p$ with $1\le p\le\infty$, based on function evaluations. We obtain worst case error bounds for general classes of functions in terms of the best $L_2$-approximation from a given nested sequence of subspaces and the Christoffel function of these subspaces. In the case $p=\infty$, our results imply that linear sampling algorithms are optimal up to a constant factor for many reproducing kernel Hilbert spaces.

math.NA

Noisy nonlinear information and entropy numbers

It is impossible to recover a vector from $\mathbb{R}^m$ with less than $m$ linear measurements, even if the measurements are chosen adaptively. Recently, it has been shown that one can recover vectors from $\mathbb{R}^m$ with arbitrary precision using only $O(\log m)$ continuous (even Lipschitz) adaptive measurements, resulting in an exponential speed-up of continuous information compared to linear information for various approximation problems. In this note, we characterize the quality of optimal (dis-)continuous information that is disturbed by deterministic noise in terms of entropy numbers. This shows that in the presence of noise the potential gain of continuous over linear measurements is limited, but significant in some cases.

math.NA

Sampling projections in the uniform norm

We show that there are sampling projections on arbitrary $n$-dimensional subspaces of $B(D)$ with at most $2n$ samples and norm of order $\sqrt{n}$, where $B(D)$ is the space of complex-valued bounded functions on a set $D$. This gives a more explicit form of the Kadets-Snobar theorem for the uniform norm and improves upon Auerbach's lemma. We discuss consequences for optimal recovery in $L_p$.

math.FA

On the power of adaption and randomization

We present bounds on the maximal gain of adaptive and randomized algorithms over non-adaptive, deterministic ones for approximating linear operators on convex sets. If the sets are additionally symmetric, then our results are optimal. For non-symmetric sets, we unify some notions of $n$-widths and s-numbers, and show their connection to minimal errors. We also discuss extensions to non-linear widths and approximation based on function values, and conclude with a list of open problems.

math.NA

Sparse grids vs. random points for high-dimensional polynomial approximation

We study polynomial approximation on a $d$-cube, where $d$ is large, and compare interpolation on sparse grids, aka Smolyak's algorithm (SA), with a simple least squares method based on randomly generated points (LS) using standard benchmark functions. Our main motivation is the influential paper [Barthelmann, Novak, Ritter: High dimensional polynomial interpolation on sparse grids, Adv. Comput. Math. 12, 2000]. We repeat and extend their theoretical analysis and numerical experiments for SA and compare to LS in dimensions up to 100. Our extensive experiments demonstrate that LS, even with only slight oversampling, consistently matches the accuracy of SA in low dimensions. In high dimensions, however, LS shows clear superiority.

math.NA

Nonlocal techniques for the analysis of deep ReLU neural network approximations

Recently, Daubechies, DeVore, Foucart, Hanin, and Petrova introduced a system of piece-wise linear functions, which can be easily reproduced by artificial neural networks with the ReLU activation function and which form a Riesz basis of $L_2([0,1])$. This work was generalized by two of the authors to the multivariate setting. We show that this system serves as a Riesz basis also for Sobolev spaces $W^s([0,1]^d)$ and Barron classes ${\mathbb B}^s([0,1]^d)$ with smoothness $0<s<1$. We apply this fact to re-prove some recent results on the approximation of functions from these classes by deep neural networks. Our proof method avoids using local approximations and allows us to track also the implicit constants as well as to show that we can avoid the curse of dimension. Moreover, we also study how well one can approximate Sobolev and Barron functions by ANNs if only function values are known.

cs.LG

How many continuous measurements are needed to learn a vector?

One can recover vectors from $\mathbb{R}^m$ with arbitrary precision, using only $\lceil \log_2(m+1)\rceil +1$ continuous measurements that are chosen adaptively. This surprising result is explained and discussed, and we present applications to infinite-dimensional approximation problems.

math.NA

Inequalities between s-numbers

Singular numbers of operators between Hilbert spaces were generalized to Banach spaces by s-numbers (in the sense of Pietsch). This allows for different choices, including approximation, Gelfand, Kolmogorov and Bernstein numbers. Here, we present an elementary proof of a bound between the smallest and the largest s-number.

math.FA

On the power of iid information for linear approximation

This survey is concerned with the power of random information for approximation in the (deterministic) worst-case setting, with special emphasis on information consisting of functionals selected independently and identically distributed (iid) at random on a class of admissible information functionals. We present a general result based on a weighted least squares method and derive consequences for special cases. Improvements are available if the information is ``Gaussian'' or if we consider iid function values for Sobolev spaces. We include open questions to guide future research on the power of random information in the context of information-based complexity.

math.NA

Exponential tractability of $L_2$-approximation with function values

We study the complexity of high-dimensional approximation in the $L_2$-norm when different classes of information are available; we compare the power of function evaluations with the power of arbitrary continuous linear measurements. Here, we discuss the situation when the number of linear measurements required to achieve an error $\varepsilon \in (0,1)$ in dimension $d\in\mathbb{N}$ depends only poly-logarithmically on $\varepsilon^{-1}$. This corresponds to an exponential order of convergence of the approximation error, which often happens in applications. However, it does not mean that the high-dimensional approximation problem is easy, the main difficulty usually lies within the dependence on the dimension $d$. We determine to which extent the required amount of information changes, if we allow only function evaluation instead of arbitrary linear information. It turns out that in this case we only lose very little, and we can even restrict to linear algorithms. In particular, several notions of tractability hold simultaneously for both types of available information.

math.NA

A sharp upper bound for sampling numbers in $L_{2}$

For a class $F$ of complex-valued functions on a set $D$, we denote by $g_n(F)$ its sampling numbers, i.e., the minimal worst-case error on $F$, measured in $L_2$, that can be achieved with a recovery algorithm based on $n$ function evaluations. We prove that there is a universal constant $c\in\mathbb{N}$ such that, if $F$ is the unit ball of a separable reproducing kernel Hilbert space, then \[ g_{cn}(F)^2 \,\le\, \frac{1}{n}\sum_{k\geq n} d_k(F)^2, \] where $d_k(F)$ are the Kolmogorov widths (or approximation numbers) of $F$ in $L_2$. We also obtain similar upper bounds for more general classes $F$, including all compact subsets of the space of continuous functions on a bounded domain $D\subset \mathbb{R}^d$, and show that these bounds are sharp by providing examples where the converse inequality holds up to a constant. The results rely on the solution to the Kadison-Singer problem, which we extend to the subsampling of a sum of infinite rank-one matrices.

math.NA

Function values are enough for $L_2$-approximation: Part II

In the first part we have shown that, for $L_2$-approximation of functions from a separable Hilbert space in the worst-case setting, linear algorithms based on function values are almost as powerful as arbitrary linear algorithms if the approximation numbers are square-summable. That is, they achieve the same polynomial rate of convergence. In this sequel, we prove a similar result for separable Banach spaces and other classes of functions.

math.NA

Function values are enough for $L_2$-approximation

We study the $L_2$-approximation of functions from a Hilbert space and compare the sampling numbers with the approximation numbers. The sampling number $e_n$ is the minimal worst case error that can be achieved with $n$ function values, whereas the approximation number $a_n$ is the minimal worst case error that can be achieved with $n$ pieces of arbitrary linear information (like derivatives or Fourier coefficients). We show that \[ e_n \,\lesssim\, \sqrt{\frac{1}{k_n} \sum_{j\geq k_n} a_j^2}, \] where $k_n \asymp n/\log(n)$. This proves that the sampling numbers decay with the same polynomial rate as the approximation numbers and therefore that function values are basically as powerful as arbitrary linear information if the approximation numbers are square-summable. Our result applies, in particular, to Sobolev spaces $H^s_{\rm mix}(\mathbb{T}^d)$ with dominating mixed smoothness $s>1/2$ and we obtain \[ e_n \,\lesssim\, n^{-s} \log^{sd}(n). \] For $d>2s+1$, this improves upon all previous bounds and disproves the prevalent conjecture that Smolyak's (sparse grid) algorithm is optimal.

math.NA

Random sections of ellipsoids and the power of random information

We study the circumradius of the intersection of an $m$-dimensional ellipsoid $\mathcal E$ with semi-axes $σ_1\geq\dots\geq σ_m$ with random subspaces of codimension $n$. We find that, under certain assumptions on $σ$, this random radius $\mathcal{R}_n=\mathcal{R}_n(σ)$ is of the same order as the minimal such radius $σ_{n+1}$ with high probability. In other situations $\mathcal{R}_n$ is close to the maximum $σ_1$. The random variable $\mathcal{R}_n$ naturally corresponds to the worst-case error of the best algorithm based on random information for $L_2$-approximation of functions from a compactly embedded Hilbert space $H$ with unit ball $\mathcal E$. In particular, $σ_k$ is the $k$th largest singular value of the embedding $H\hookrightarrow L_2$. In this formulation, one can also consider the case $m=\infty$, and we prove that random information behaves very differently depending on whether $σ\in \ell_2$ or not. For $σ\notin \ell_2$ random information is completely useless, i.e., $\mathbb E[\mathcal{R}_n] = σ_1$. For $σ\in \ell_2$ the expected radius of random information tends to zero at least at rate $o(1/\sqrt{n})$ as $n\to\infty$. In the important case $σ_k \asymp k^{-α} \ln^{-β}(k+1)$, where $α> 0$ and $β\in\mathbb R$, we obtain that $$ \mathbb E [\mathcal{R}_n(σ)] \asymp \begin{cases} σ_1 & : α<1/2 \,\text{ or }\, β\leqα=1/2 \\ σ_n \, \sqrt{\ln(n+1)} & : β>α=1/2 \\ σ_{n+1} & : α>1/2. \end{cases} $$ In the proofs we use a comparison result for Gaussian processes à la Gordon, exponential estimates for sums of chi-squared random variables, and estimates for the extreme singular values of (structured) Gaussian random matrices. The upper bound is constructive. It is proven for the worst case error of a least squares estimator.

math.FA