arXiv ScienceSearch

arXiv subjects

Alex Dytso

Publications and source records attributed to Alex Dytso.

At least 19 recordsLinked to original sources

On the Evolution of the Capacity-Achieving Input Support for the Amplitude-Constrained AWGN Channel

We consider an additive white Gaussian noise channel subject to a peak-amplitude constraint and study how the support of the capacity-achieving input distribution (CAID) evolves as the amplitude constraint $A$ varies. Although the CAID is known to be unique, symmetric, discrete, and finitely supported, the structure of its support transitions has remained largely unresolved. We show that the origin is the only possible degenerate support point and the only possible inactive contact point of the KKT function. We then establish continuity properties of the optimal distribution and the KKT function and prove that every nonzero support point is locally stable: under small changes in $A$, it persists as a unique nearby atom whose location and probability mass vary continuously. Combining these results, we prove that, locally, the support cardinality at a nearby amplitude is either unchanged or larger by one and that any such increase can occur only at the origin. Consequently, all local changes in support cardinality are confined to the origin: locally, the only possible transition mechanisms are the appearance or disappearance of the point at the origin and the splitting or merging of the origin into a symmetric pair.

cs.IT

The Binomial Channel: On Capacity, Optimal Inputs, and Beta-Binomial Approximation

We study the binomial channel with input alphabet $[0,1]$ and output alphabet ${0,\ldots,n}$. We investigate its capacity and the structure of the capacity-achieving input and output distributions. Since the output alphabet is finite whereas the input alphabet is continuous, different input distributions may induce the same output distribution; hence, uniqueness and support properties of optimal inputs do not follow from strict concavity arguments. We first establish structural properties of the capacity-achieving input distribution. In particular, we show that it is discrete, unique, symmetric around $1/2$, and contains the endpoints ${0,1}$ in its support. We also derive location constraints and bounds on the probability masses of support points, and improve the Witsenhausen-type upper bound on the support size from order $n$ to order $n/2$. We derive explicit nonasymptotic upper and lower bounds on the capacity $C(n)$. These bounds imply $C(n)=\frac{1}{2}\log(\frac{n\pi}{2e})+o(1).$ The lower bound is obtained by evaluating the mutual information at the reference input $X_r\sim \mathrm{Beta}(1/2,1/2)$, which induces a beta-binomial output distribution, while the upper bound follows from a minimax redundancy construction. Finally, we prove an improved lower bound on the support size of the capacity-achieving input distribution. We show that the beta-binomial output induced by $X_r$ is asymptotically optimal and close to the capacity-achieving output distribution in relative entropy and $\chi^2$ divergence. We also prove a finite-mixture approximation lower bound showing that the beta-binomial output cannot be approximated too accurately by binomial mixtures with few components. Combining these results yields a support-size lower bound of order $\Omega(\sqrt{n\log\log n})$, with explicit constants. Numerical results illustrate the capacity bounds and optimal input.

cs.IT

An Improved Lower Bound on Support Size of Capacity-Achieving Inputs for the Binomial Channel: Extended version

We study the binomial channel and the structure of its capacity-achieving input and output distributions. It is known that the capacity-achieving input distribution is discrete and supported on finitely many points. The best previously known bounds show that the support size of the capacity-achieving distribution is lower-bounded by a term of order $\sqrt n$ and upper-bounded by a term of order $n/2$, where $n$ is the number of trials. In this work, we derive a new lower bound on the support size of order $\sqrt{n\log\log n}$, up to explicit constants. The proof consists of three main steps. First, we derive new upper and lower bounds on the capacity with a gap that vanishes as $n\to\infty$, which yields $C(n)=\frac12\log\frac{n\pi}{2e}+o(1)$. Second, we show that the Beta-binomial output distribution induced by the reference input $X_r\sim\mathrm{Beta}(1/2,1/2)$ is asymptotically optimal: it approaches the capacity-achieving output distribution in relative entropy and, after a comparison step, in $\chi^2$ divergence. Third, we prove a quantitative $\chi^2$ approximation lower bound showing that this Beta-binomial output cannot be approximated too well by the output induced by a $K$-point input. Combining these ingredients forces the capacity-achieving input distribution to have at least order $\sqrt{n\log\log n}$ mass points.

cs.IT

Sub-Gaussian Concentration and Entropic Normality of the Maximum Likelihood Estimator

It is well known that, under standard regularity conditions, the maximum likelihood estimator (MLE) satisfies a central limit theorem and converges in distribution to a Gaussian random variable as the sample size grows. This paper strengthens this classical result by developing several stronger forms of asymptotic normality for the normalized MLE. With additional assumptions on the score, we first establish sub-Gaussian tail bounds and convergence of all moments for the normalized estimation error. We then prove an entropic central limit theorem for a smoothed version of the estimator, showing convergence in relative entropy to the limiting Gaussian law. When the Fisher information of the normalized estimate is bounded, or its density has bounded first derivative, we further show that the smoothing can be removed, yielding entropic normality of the MLE itself. The proofs develop auxiliary tools that may be of independent interest, including exponential consistency bounds, high-moment estimates, and entropy-control arguments for the estimator.

cs.IT

Support Size of $\varepsilon$-Capacity-Achieving Inputs for the Amplitude-Constrained AWGN Channel

We study the amplitude-constrained additive white Gaussian noise (AWGN) channel from the perspective of near-optimal input distributions. While it is known that the capacity-achieving input is discrete with finitely many mass points, the precise scaling of its support size as a function of the amplitude constraint remains an open problem. In this work, we instead consider the minimal support size required to achieve capacity up to an $\varepsilon$-gap. We introduce the quantity $K_\varepsilon(A)$, defined as the smallest support size among discrete inputs supported on $[-A,A]$ that achieves mutual information within $\varepsilon$ of capacity. We show that this relaxed formulation is significantly more tractable and admits sharp characterizations across different regimes of $\varepsilon$. In particular, when $\varepsilon$ decays polynomially with $A$, i.e., $\varepsilon = A^{-\beta}$ for $\beta \geq 1$, we establish that $K_\varepsilon(A) = \Theta(A\sqrt{\log A})$. For exponentially small gaps, we obtain bounds of order between $A\sqrt{\log A}$ and $A^{3/2}$. Our approach combines approximation-theoretic bounds for Gaussian mixtures with information-theoretic control of entropy via $\chi^2$-divergence, together with a wrapping argument that relates the problem to approximating the uniform distribution on the circle. Beyond the technical results, our framework provides a conceptual explanation for the variety of scaling laws observed in prior numerical studies, showing that these correspond to different regimes of $\varepsilon$-optimality rather than intrinsic properties of the exact optimizer.

cs.IT

$\alpha$-Mutual Information for the Gaussian Noise Channel

In this paper, we study Sibson's $\alpha$-mutual information in the context of the additive Gaussian noise channel. While the classical case $\alpha = 1$ is well understood and admits deep connections to estimation-theoretic quantities, such as the minimum mean-square error (MMSE) and Fisher information, many of the corresponding structural properties for general $\alpha$ remain less explored. Our goal is to develop a systematic understanding of $\alpha$-mutual information in the Gaussian noise setting and to identify which properties extend beyond the Shannon case. To this end, we establish several regularity properties, including finiteness conditions, continuity with respect to the signal-to-noise ratio (SNR) and the input distribution, and strict concavity/convexity properties that ensure uniqueness in associated optimization problems. A central contribution is the development of an $\alpha$-I-MMSE relationship, generalizing the classical identity by relating the derivative of $\alpha$-mutual information with respect to SNR to the MMSE evaluated under appropriately tilted distributions. This connection further leads to a generalized de Bruijn identity and new estimation-theoretic representations of R\'enyi entropy and differential R\'enyi entropy. We also characterize the low- and high-SNR behavior. In the low-SNR regime, the first-order behavior depends only on the input variance. In the high-SNR regime, for discrete inputs, $\alpha$-mutual information converges to the R\'enyi entropy of order $1/\alpha$, while for general inputs we connect it to $\alpha$-information dimension. Overall, our results show that many fundamental relationships between information and estimation extend beyond the Shannon setting, in a form involving $\alpha$-tilted distributions.

cs.IT

Functional Properties of the Focal-Entropy

The focal-loss has become a widely used alternative to cross-entropy in class-imbalanced classification problems, particularly in computer vision. Despite its empirical success, a systematic information-theoretic study of the focal-loss remains incomplete. In this work, we adopt a distributional viewpoint and study the focal-entropy, a focal-loss analogue of the cross-entropy. Our analysis establishes conditions for finiteness, convexity, and continuity of the focal-entropy, and provides various asymptotic characterizations. We prove the existence and uniqueness of the focal-entropy minimizer, describe its structure, and show that it can depart significantly from the data distribution. In particular, we rigorously show that the focal-loss amplifies mid-range probabilities, suppresses high-probability outcomes, and, under extreme class imbalance, induces an over-suppression regime in which very small probabilities are further diminished. These results, which are also experimentally validated, offer a theoretical foundation for understanding the focal-loss and clarify the trade-offs that it introduces when applied to imbalanced learning tasks.

cs.IT

An Improved Lower Bound on Cardinality of Support of the Amplitude-Constrained AWGN Channel

We study the amplitude-constrained additive white Gaussian noise channel. It is well known that the capacity-achieving input distribution for this channel is discrete and supported on finitely many points. The best known bounds show that the support size of the capacity-achieving distribution is lower-bounded by a term of order $A$ and upper-bounded by a term of order $A^2$, where $A$ denotes the amplitude constraint. It was conjectured in [1] that the linear scaling is optimal. In this work, we establish a new lower bound of order $A\sqrt{\log A}$, improving the known bound and ruling out the conjectured linear scaling. To obtain this result, we quantify the fact that the capacity-achieving output distribution is close to the uniform distribution in the interior of the amplitude constraint. Next, we introduce a wrapping operation that maps the problem to a compact domain and develop a theory of best approximation of the uniform distribution by finite Gaussian mixtures. These approximation bounds are then combined with stability properties of capacity-achieving distributions to yield the final support-size lower bound.

cs.IT

Functional uniqueness and stability of Gaussian priors in optimal L1 estimation

We study when optimal Bayesian estimators under Gaussian noise are approximately linear, and what this implies about the underlying prior distribution. Consider the classical model \(Y = X + Z\), where \(Z\) is Gaussian and independent of \(X\). It is well known that under squared-error loss, the conditional mean \(\mathbb{E}[X|Y]\) is a linear function of \(Y\) if and only if the prior is Gaussian. Much less is understood under absolute-error loss, where the optimal estimator is the conditional median and standard orthogonality-based tools no longer apply. Recent work has established that, in the Gaussian noise model, the Gaussian prior is also the unique distribution that induces an exactly linear conditional median. In this paper, we move beyond exact characterizations and develop a quantitative stability theory: if the optimal estimator is approximately linear, must the prior be close to Gaussian? For the \(L_2\) setting, we derive explicit rates showing that near-linearity of the conditional mean forces the prior to be close to Gaussian in the Levy metric. For the \(L_1\) setting, we develop a functional-analytic framework based on Hermite expansions and adjoint operators, establishing that approximate linearity of the conditional median implies proximity to the Gaussian family.

cs.IT

A Rate-Distortion Bound for ISAC

This paper addresses the fundamental performance limits of Integrated Sensing and Communication (ISAC) systems by introducing a novel converse bound based on rate-distortion theory. This rate-distortion bound (RDB) overcomes the restrictive regularity conditions of classical estimation theory, such as the Bayesian Cram\'er-Rao Bound (BCRB). The proposed framework is broadly applicable, holding for arbitrary parameter distributions and distortion measures, including mean-squared error and probability of error. The bound is proved to be tight in the high sensing noise regime and can be strictly tighter than the BCRB in the low sensing noise regime. The RDB's utility is demonstrated on two challenging scenarios: Nakagami fading channel estimation, where it provides a valid bound even when the BCRB is inapplicable, and a binary occupancy detection task, showcasing its versatility for discrete sensing problems. This work provides a powerful and general tool for characterizing the ultimate performance tradeoffs in ISAC systems.

cs.IT

Linearity-Inducing Priors for Poisson Parameter Estimation Under $L^{1}$ Loss

We study prior distributions for Poisson parameter estimation under $L^1$ loss. Specifically, we construct a new family of prior distributions whose optimal Bayesian estimators (the conditional medians) can be any prescribed increasing function that satisfies certain regularity conditions. In the case of affine estimators, this family is distinct from the usual conjugate priors, which are gamma distributions. Our prior distributions are constructed through a limiting process that matches certain moment conditions. These results provide the first explicit description of a family of distributions, beyond the conjugate priors, that satisfy the affine conditional median property; and more broadly for the Poisson noise model they can give any arbitrarily prescribed conditional median.

math.ST

Lossy Source Coding with Focal Loss

Focal loss has recently gained significant popularity, particularly in tasks like object detection where it helps to address class imbalance by focusing more on hard-to-classify examples. This work proposes the focal loss as a distortion measure for lossy source coding. The paper provides single-shot converse and achievability bounds. These bounds are then used to characterize the distortion-rate trade-off in the infinite blocklength, which is shown to be the same as that for the log loss case. In the non-asymptotic case, the difference between focal loss and log loss is illustrated through a series of simulations.

cs.IT

Estimation Error: Distribution and Pointwise Limits

In this paper, we examine the distribution and convergence properties of the estimation error $W = X - \hat{X}(Y)$, where $\hat{X}(Y)$ is the Bayesian estimator of a random variable $X$ from a noisy observation $Y = X +\sigma Z$ where $\sigma$ is the parameter indicating the strength of noise $Z$. Using the conditional expectation framework (that is, $\hat{X}(Y)$ is the conditional mean), we define the normalized error $\mathcal{E}_\sigma = \frac{W}{\sigma}$ and explore its properties. Specifically, in the first part of the paper, we characterize the probability density function of $W$ and $\mathcal{E}_\sigma$. Along the way, we also find conditions for the existence of the inverse functions for the conditional expectations. In the second part, we study pointwise (i.e., almost sure) convergence of $\mathcal{E}_\sigma$ as $\sigma \to 0$ under various assumptions about the noise and the underlying distributions. Our results extend some of the previous limits of $\mathcal{E}_\sigma$ as $\sigma \to 0$ studied under the $L^2$ convergence, known as the \emph{mmse dimension}, to the pointwise case.

cs.IT

Generalized Linear Models with 1-Bit Measurements: Asymptotics of the Maximum Likelihood Estimator

This work establishes regularity conditions for consistency and asymptotic normality of the multiple parameter maximum likelihood estimator(MLE) from censored data, where the censoring mechanism is in the form of $1$-bit measurements. The underlying distribution of the uncensored data is assumed to belong to the exponential family, with natural parameters expressed as a linear combination of the predictors, known as generalized linear model (GLM). As part of the analysis, the Fisher information matrix is also derived for both censored and uncensored data, which helps to quantify the impact of censoring and assess the performance of the MLE. The choice of GLM allows one to consider a variety of practical examples where 1-bit estimation is of interest. In particular, it is shown how the derived results can be used to analyze two practically relevant scenarios: the Gaussian model with both unknown mean and variance, and the Poisson model with an unknown mean.

math.ST

A Comprehensive Study on Ziv-Zakai Lower Bounds on the MMSE

This paper explores Bayesian lower bounds on the minimum mean squared error (MMSE) that belong to the Ziv-Zakai (ZZ) family. The ZZ technique relies on connecting the bound to an M-ary hypothesis testing problem. Three versions of the ZZ bound (ZZB) exist: the first relies on the so-called valley-filling function (VFF), the second omits the VFF, and the third, i.e., the single-point ZZB (SZZB), uses a single point maximization. The first part of this paper provides the most general version of the bounds. First, it is shown that these bounds hold without any assumption on the distribution of the estimand. Second, the SZZB bound is extended to an M-ary setting and a version of it for the multivariate case is provided. In the second part, general properties of the bounds are provided. First, it is shown that all the bounds tensorize. Second, a complete characterization of the high-noise asymptotic is provided, which is used to argue about the tightness of the bounds. Third, the low-noise asymptotic is provided for mixed-input distributions and Gaussian additive noise channels. Specifically, in the low-noise, it is shown that the SZZB is not always tight. In the third part, the tightness of the bounds is evaluated. First, it is shown that in the low-noise regime the ZZB bound without the VFF is tight for mixed-input distributions and Gaussian additive noise channels. Second, for discrete inputs, the ZZB with the VFF is shown to be always sub-optimal, and equal to zero without the VFF. Third, unlike for the ZZB, an example is shown for which the SZZB is tight to the MMSE for discrete inputs. Fourth, sufficient and necessary conditions for the tightness of the bounds are provided. Finally, some examples are shown in which the bounds in the ZZ family outperform other well-known Bayesian bounds, i.e., the Cram\'er-Rao bound and the maximum entropy bound.

cs.IT

Multivariate Priors and the Linearity of Optimal Bayesian Estimators under Gaussian Noise

Consider the task of estimating a random vector $X$ from noisy observations $Y = X + Z$, where $Z$ is a standard normal vector, under the $L^p$ fidelity criterion. This work establishes that, for $1 \leq p \leq 2$, the optimal Bayesian estimator is linear and positive definite if and only if the prior distribution on $X$ is a (non-degenerate) multivariate Gaussian. Furthermore, for $p > 2$, it is demonstrated that there are infinitely many priors that can induce such an estimator.

math.ST

On $2 \times 2$ MIMO Gaussian Channels with a Small Discrete-Time Peak-Power Constraint

A multi-input multi-output (MIMO) Gaussian channel with two transmit antennas and two receive antennas is studied that is subject to an input peak-power constraint. The capacity and the capacity-achieving input distribution are unknown in general. The problem is shown to be equivalent to a channel with an identity matrix but where the input lies inside and on an ellipse with principal axis length $r_p$ and minor axis length $r_m$. If $r_p \le \sqrt{2}$, then the capacity-achieving input has support on the ellipse. A sufficient condition is derived under which a two-point distribution is optimal. Finally, if $r_m < r_p \le \sqrt{2}$, then the capacity-achieving distribution is discrete.

cs.IT

Data-Driven Estimation of the False Positive Rate of the Bayes Binary Classifier via Soft Labels

Classification is a fundamental task in many applications on which data-driven methods have shown outstanding performances. However, it is challenging to determine whether such methods have achieved the optimal performance. This is mainly because the best achievable performance is typically unknown and hence, effectively estimating it is of prime importance. In this paper, we consider binary classification problems and we propose an estimator for the false positive rate (FPR) of the Bayes classifier, that is, the optimal classifier with respect to accuracy, from a given dataset. Our method utilizes soft labels, or real-valued labels, which are gaining significant traction thanks to their properties. We thoroughly examine various theoretical properties of our estimator, including its consistency, unbiasedness, rate of convergence, and variance. To enhance the versatility of our estimator beyond soft labels, we also consider noisy labels, which encompass binary labels. For noisy labels, we develop effective FPR estimators by leveraging a denoising technique and the Nadaraya-Watson estimator. Due to the symmetry of the problem, our results can be readily applied to estimate the false negative rate of the Bayes classifier.

cs.LG