arXiv ScienceSearch

arXiv · 2008.01011

Phase Transitions in Rate Distortion Theory and Deep Learning

Abstract

Rate distortion theory is concerned with optimally encoding a given signal class $\mathcal{S}$ using a budget of $R$ bits, as $R\to\infty$. We say that $\mathcal{S}$ can be compressed at rate $s$ if we can achieve an error of $\mathcal{O}(R^{-s})$ for encoding $\mathcal{S}$; the supremal compression rate is denoted $s^\ast(\mathcal{S})$. Given a fixed coding scheme, there usually are elements of $\mathcal{S}$ that are compressed at a higher rate than $s^\ast(\mathcal{S})$ by the given coding scheme; we study the size of this set of signals. We show that for certain "nice" signal classes $\mathcal{S}$, a phase transition occurs: We construct a probability measure $\mathbb{P}$ on $\mathcal{S}$ such that for every coding scheme $\mathcal{C}$ and any $s >s^\ast(\mathcal{S})$, the set of signals encoded with error $\mathcal{O}(R^{-s})$ by $\mathcal{C}$ forms a $\mathbb{P}$-null-set. In particular our results apply to balls in Besov and Sobolev spaces that embed compactly into $L^2(Ω)$ for a bounded Lipschitz domain $Ω$. As an application, we show that several existing sharpness results concerning function approximation using deep neural networks are generically sharp. We also provide quantitative and non-asymptotic bounds on the probability that a random $f\in\mathcal{S}$ can be encoded to within accuracy $\varepsilon$ using $R$ bits. This result is applied to the problem of approximately representing $f\in\mathcal{S}$ to within accuracy $\varepsilon$ by a (quantized) neural network that is constrained to have at most $W$ nonzero weights and is generated by an arbitrary "learning" procedure. We show that for any $s >s^\ast(\mathcal{S})$ there are constants $c,C$ such that, no matter how we choose the "learning" procedure, the probability of success is bounded from above by $\min\big\{1,2^{C\cdot W\lceil\log_2(1+W)\rceil^2 -c\cdot\varepsilon^{-1/s}}\big\}$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Philipp Grohs, Andreas Klotz, Felix Voigtlaender. 2020-08-03. Phase Transitions in Rate Distortion Theory and Deep Learning. https://arxiv.org/abs/2008.01011

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Typical dynamical properties of operators on $\ell_p$

We investigate the typical dynamical properties of hypercyclic operators in $\mathcal{L}_M(X)$, the set of all bounded linear operators on $X$ whose norms are at most $M$, when $X=\ell_p$, $1< p<\infty$. We show that, with respect to SOT$^*$, a typical operator $T\in \mathcal{L}_M(X)$ is weakly mixing, is weakly disjoint from a given hypercyclic operator $S$, is not topologically ergodic, and satisfies $(T,T^2,\dotsc,T^k)$ is disjoint hypercyclic for any $k\geq 2$. We also study the typical dynamical properties for the concrete family $\mathcal{M}=\{I+B_w\in \mathcal{L}(X)\colon w\in c_0(\mathbb{Z})\}$, endowed with the norm topology, where $B_w$ is a bilateral weighted backward shift.

math.FA

A bi-Lipschitz characterization of strong minimum-attainment for Lipschitz maps

We completely characterize the denseness of strongly minimum-attaining Lipschitz functions, a minimum analogue for strongly norm-attaining Lipschitz functions, in terms of bi-Lipschitz embeddings. More precisely, our main result shows that the set of strongly minimum-attaining Lipschitz functions defined on a complete metric space $M$ fails the denseness if and only if $M$ is bi-Lipschitz equivalent to a subset of $\mathbb{R}$ with positive Lebesgue measure, or equivalently, if $M$ admits a bi-Lipschitz embedding into $\mathbb{R}$ and $M$ has positive 1-dimensional Hausdorff measure. As a consequence, we provide an isometric characterization of the pure 1-unrectifiability of $M$ in terms of strongly minimum-attaining Lipschitz maps defined on bi-Lipschitz copies of closed subsets of $M$. Several counterexamples showing that the main result cannot be naturally extended to the vector-valued setting are also presented.

math.FA

On weak dominance of t-conorms over t-norms

The weak dominance of aggregation operators, particularly between triangular norms (t-norms) and triangular conorms (t-conorms), has attracted considerable attention in aggregation operator theory. While several characterizations have been obtained for Archimedean and continuous cases, a general criterion for continuous t-conorms over continuous t-norms remains to be fully clarified. In this paper, we provide a complete characterization of a continuous t-conorm weakly dominating a continuous t-norm. We first reduce the problem for ordinal sum operators to that for their single Archimedean components, and then express the weak dominance condition entirely in terms of the additive generators of these components. Our approach covers both strict and nilpotent cases uniformly, and recovers the known results for Archimedean operators as a special case.

math.FA