arXiv ScienceSearch

arXiv subjects

Martin Kroll

Publications and source records attributed to Martin Kroll.

18 recordsLinked to original sources

On public and private binary classification with metric space valued predictors

We consider the problem of binary classification in a framework where the predictor $X$ takes values in an arbitrary separable metric space $\mathcal X$ and the label $Y$ values in $\{ \pm 1 \}$. In the first part of this work, we assume that one has direct access to an i.i.d. sample $(X_1,Y_1),\ldots,(X_n,Y_n)$ from the unknown distribution of the pair $(X,Y)$. We derive a convergence rate for the Proto-NN classifier which was recently introduced as a classifier in the presence of metric space-valued predictors. In the second part of the paper, we reconsider the same problem under an additional privacy constraint. More precisely, we work in the framework of local differential privacy where one assumes that the data $(X_1,Y_1),\ldots,(X_n,Y_n)$ cannot be directly observed but only a privatised surrogate obtained through a suitable mechanism satisfying the privacy constraint is available. The statistician should select an optimal privacy mechanism from the class of all mechanism that guarantee local differential privacy. Our method of choice is to add Laplace distributed noise to both a set of in Proto-NN classifier using the privatised data only is universally consistent. Finally, a rate of convergence for the privatised Proto-NN classifier is derived.

math.ST

Asymptotic equivalence of non-parametric regression with spherical regressors and Gaussian white noise

We study the asymptotic behavior of both spherical $t$-designs and random uniform designs as the set of sampling points in non-parametric regression with spherical regressors of arbitrary dimension. We show that the corresponding regression experiments are asymptotically equivalent, in the sense of Le Cam, to the same sequence of Gaussian white noise experiments as the sample size tends to infinity. More precisely, global asymptotic equivalence is established over spherical Sobolev balls (for both the fixed and the random uniform design case) and over spherical Besov balls (for the fixed design case). We also derive a matching non-equivalence result showing the sharpness of the imposed smoothness assumptions for any fixed choice of design points.

math.ST

Optimal Designs for Regression on Lie Groups

We consider a linear regression model with complex-valued response and predictors from a compact and connected Lie group. The regression model is formulated in terms of eigenfunctions of the Laplace-Beltrami operator on the Lie group. We show that the normalized Haar measure is an approximate optimal design with respect to all Kiefer's $\Phi_p$-criteria. Inspired by the concept of $t$-designs in the field of algebraic combinatorics, we then consider so-called $\lambda$-designs in order to construct exact $\Phi_p$-optimal designs for fixed sample sizes in the considered regression problem. In particular, we explicitly construct $\Phi_p$-optimal designs for regression models with predictors in the Lie groups $\mathrm{SU}(2)$ and $\mathrm{SO}(3)$, the groups of $2\times 2$ unitary matrices and $3\times 3$ orthogonal matrices with determinant equal to $1$, respectively. We also discuss the advantages of the derived theoretical results in a concrete biological application.

math.ST

Towards active learning: A stopping criterion for the sequential sampling of grain boundary degrees of freedom

Many materials processes and properties depend on the anisotropy of the energy of grain boundaries, i.e.~on the fact that this energy is a function of the five geometric degrees of freedom (DOF) of the interface. To access this parameter space in an efficient way and to discover energy cusps in unexplored regions, a method was recently established, which combines atomistic simulations with statistical methods 10.1002/adts.202100615. This sequential sampling technique is now extended in the spirit of an active learning algorithm by adding a criterion to decide when the sampling has advanced enough to stop. In this instance, two parameters to analyse the sampling results on the fly are introduced: the number of cusps, which correspond to the most interesting and important regions of the energy landscape, and the maximum change of energy between two sequential iterations. Monitoring these two quantities provides valuable insight into how the subspaces are energetically structured. The combination of both parameters provides the necessary information to evaluate the sampling of the 2D subspaces of grain boundary plane inclinations of even non-periodic, low angle grain boundaries. With a reasonable number of data points in the initial design, only a few appropriately chosen sequential iterations already improve the accuracy of the sampling substantially and unknown cusps can be found within a few additional sequential steps.

cond-mat.mtrl-sci

On rate optimal private regression under local differential privacy

We consider the problem of estimating a regression function from anonymized data in the framework of local differential privacy. We propose a novel partitioning estimate of the regression function, derive a rate of convergence for the excess prediction risk over H\"older classes, and prove a matching lower bound. In contrast to the existing literature on the problem the so-called strong density assumption on the design distribution is obsolete.

math.ST

Efficient prediction of grain boundary energies from atomistic simulations via sequential design

Data based materials science is the new promise to accelerate materials design. Especially in computational materials science, data generation can easily be automatized. Usually, the focus is on processing and evaluating the data to derive rules or to discover new materials, while less attention is being paid on the strategy to generate the data. In this work, we show that by a sequential design of experiment scheme, the process of generating and learning from the data can be combined to discover the relevant sections of the parameter space. Our example is the energy of grain boundaries as a function of their geometric degrees of freedom, calculated via atomistic simulations. The sampling of this grain boundary energy space, or even subspaces of it, represents a challenge due to the presence of deep cusps of the energy, which are located at irregular intervals of the geometric parameters. Existing approaches to sample grain boundary energy subspaces therefore either need a huge amount of datapoints or a~priori knowledge of the positions of these cusps. We combine statistical methods with atomistic simulations and a sequential sampling technique and compare this strategy to a regular sampling technique. We thereby demonstrate that this sequential design is able to sample a subspace with a minimal amount of points while finding unknown cusps automatically.

cond-mat.mtrl-sci

Multivariate density estimation from privatised data: universal consistency and minimax rates

We revisit the classical problem of nonparametric density estimation but impose local differential privacy constraints. Under such constraints, the original multivariate data $X_1,\ldots,X_n \in \mathbb{R}^d$ cannot be directly observed, and all estimators are functions of the randomised output of a suitable privacy mechanism. The statistician is free to choose the form of the privacy mechanism, and in this work we propose to add Laplace distributed noise to a discretisation of the location of a vector $X_i$. Based on these randomised data, we design a novel estimator of the density function, which can be viewed as a privatised version of the well-studied histogram density estimator. Our theoretical results include universal pointwise consistency and strong universal $L_1$-consistency. In addition, a convergence rate over classes of Lipschitz functions is derived, which is complemented by a matching minimax lower bound. We illustrate the trade-off between data utility and privacy by means of a small simulation study.

math.ST

Asymptotic equivalence for nonparametric regression with dependent errors: Gauss-Markov processes

For the class of Gauss-Markov processes we study the problem of asymptotic equivalence of the nonparametric regression model with errors given by the increments of the process and the continuous time model, where a whole path of a sum of a deterministic signal and the Gauss-Markov process can be observed. In particular we provide sufficient conditions such that asymptotic equivalence of the two models holds for functions from a given class, and we verify these for the special cases of Sobolev ellipsoids and H\"older classes with smoothness index $> 1/2$ under mild assumptions on the Gauss-Markov process at hand. To derive these results, we develop an explicit characterization of the reproducing kernel Hilbert space associated with the Gauss-Markov process, that hinges on a characterization of such processes by a property of the corresponding covariance kernel introduced by Doob. In order to demonstrate that the given assumptions on the Gauss-Markov process are in some sense sharp we also show that asymptotic equivalence fails to hold for the special case of Brownian bridge. Our findings demonstrate that the well-known asymptotic equivalence of the Gaussian white noise model and the nonparametric regression model with i.i.d. standard normal errors can be extended to a result treating general Gauss-Markov noises in a unified manner.

math.ST

Adaptive spectral density estimation by model selection under local differential privacy

We study spectral density estimation under local differential privacy. Anonymization is achieved through truncation followed by Laplace perturbation. We select our estimator from a set of candidate estimators by a penalized contrast criterion. This estimator is shown to attain nearly the same rate of convergence as the best estimator from the candidate set. A key ingredient of the proof are recent results on concentration of quadratic forms in terms of sub-exponential random variables obtained in arXiv:1903.05964. We illustrate our findings in a small simulation study.

math.ST

Pointwise adaptive kernel density estimation under local approximate differential privacy

We consider non-parametric density estimation in the framework of local approximate differential privacy. In contrast to centralized privacy scenarios with a trusted curator, in the local setup anonymization must be guaranteed already on the individual data owners' side and therefore must precede any data mining tasks. Thus, the published anonymized data should be compatible with as many statistical procedures as possible. We suggest adding Laplace noise and Gaussian processes (both appropriately scaled) to kernel density estimators to obtain approximate differential private versions of the latter ones. We obtain minimax type results over Sobolev classes indexed by a smoothness parameter $s>1/2$ for the mean squared error at a fixed point. In particular, we show that taking the average of private kernel density estimators from $n$ different data owners attains the optimal rate of convergence if the bandwidth parameter is correctly specified. Notably, the optimal convergence rate in terms of the sample size $n$ is $n^{-(2s-1)/(2s+1)}$ under local differential privacy and thus deteriorated to the rate $n^{-(2s-1)/(2s)}$ which holds without privacy restrictions. Since the optimal choice of the bandwidth parameter depends on the smoothness $s$ and is thus not accessible in practice, adaptive methods for bandwidth selection are necessary and must, in the local privacy framework, be performed directly on the anonymized data. We address this problem by means of a variant of Lepski's method tailored to the privacy setup and obtain general oracle inequalities for private kernel density estimators. In the Sobolev case, the resulting adaptive estimator attains the optimal rate of convergence at least up to extra logarithmic factors.

math.ST

Local differential privacy: Elbow effect in optimal density estimation and adaptation over Besov ellipsoids

We address the problem of non-parametric density estimation under the additional constraint that only privatised data are allowed to be published and available for inference. For this purpose, we adopt a recent generalisation of classical minimax theory to the framework of local $\alpha$-differential privacy and provide a lower bound on the rate of convergence over Besov spaces $B^s_{pq}$ under mean integrated $\mathbb L^r$-risk. This lower bound is deteriorated compared to the standard setup without privacy, and reveals a twofold elbow effect. In order to fulfil the privacy requirement, we suggest adding suitably scaled Laplace noise to empirical wavelet coefficients. Upper bounds within (at most) a logarithmic factor are derived under the assumption that $\alpha$ stays bounded as $n$ increases: A linear but non-adaptive wavelet estimator is shown to attain the lower bound whenever $p \geq r$ but provides a slower rate of convergence otherwise. An adaptive non-linear wavelet estimator with appropriately chosen smoothing parameters and thresholding is shown to attain the lower bound within a logarithmic factor for all cases.

math.ST

Rate optimal estimation of quadratic functionals in inverse problems with partially unknown operator and application to testing problems

We consider the estimation of quadratic functionals in a Gaussian sequence model where the eigenvalues are supposed to be unknown and accessible through noisy observations only. Imposing smoothness assumptions both on the signal and the sequence of eigenvalues, we develop a minimax theory for this problem. We propose a truncated series estimator and show that it attains the optimal rate of convergence if the truncation parameter is chosen appropriately. Consequences for testing problems in inverse problems are equally discussed: in particular, the minimax rates of testing for signal detection and goodness-of-fit testing are derived.

math.ST

Nonparametric Poisson regression from independent and weakly dependent observations by model selection

We consider the non-parametric Poisson regression problem where the integer valued response $Y$ is the realization of a Poisson random variable with parameter $\lambda(X)$. The aim is to estimate the functional parameter $\lambda$ from independent or weakly dependent observations $(X_1,Y_1),\ldots,(X_n,Y_n)$ in a random design framework. First we determine upper risk bounds for projection estimators on finite dimensional subspaces under mild conditions. In the case of Sobolev ellipsoids the obtained rates of convergence turn out to be optimal. The main part of the paper is devoted to the construction of adaptive projection estimators of $\lambda$ via model selection. We proceed in two steps: first, we assume that an upper bound for $\Vert \lambda \Vert_\infty$ is known. Under this assumption, we construct an adaptive estimator whose dimension parameter is defined as the minimizer of a penalized contrast criterion. Second, we replace the known upper bound on $\Vert \lambda \Vert_\infty$ by an appropriate plug-in estimator of $\Vert \lambda \Vert_\infty$. The resulting adaptive estimator is shown to attain the minimax optimal rate up to an additional logarithmic factor both in the independent and the weakly dependent setup. Appropriate concentration inequalities for Poisson point processes turn out to be an important ingredient of the proofs. We illustrate our theoretical findings by a short simulation study and conclude by indicating directions of future research.

math.ST

Interplay of Fluorescence and Phosphorescence in Organic Biluminescent Emitters

Biluminescent organic emitters show simultaneous fluorescence and phosphorescence at room temperature. So far, the optimization of the room temperature phosphorescence (RTP) in these materials has drawn the attention of research. However, the continuous wave operation of these emitters will consequently turn them into systems with vastly imbalanced singlet and triplet populations, which is due to the respective excited state lifetimes. This study reports on the exciton dynamics of the biluminophore NPB (N,N-di(1-naphthyl)-N,N-diphenyl-(1,1-biphenyl)-4,4-diamine). In the extreme case, the singlet and triplet exciton lifetimes stretch from 3 ns to 300 ms, respectively. Through sample engineering and oxygen quenching experiments, the triplet exciton density can be controlled over several orders of magnitude allowing to studying exciton interactions between singlet and triplet manifolds. The results show, that singlet-triplet annihilation reduces the overall biluminescence efficiency already at moderate excitation levels. Additionally, the presented system represents an illustrative role model to study excitonic effects in organic materials.

physics.chem-ph

Nonparametric intensity estimation from noisy observations of a Poisson process under unknown error distribution

We consider the nonparametric estimation of the intensity function of a Poisson point process in a circular model from indirect observations $N_1,\ldots,N_n$. These observations emerge from hidden point process realizations with the target intensity through contamination with additive error. In case that the error distribution can only be estimated from an additional sample $Y_1,\ldots,Y_m$ we derive minimax rates of convergence with respect to the sample sizes $n$ and $m$ under abstract smoothness conditions and propose an orthonormal series estimator which attains the optimal rate of convergence. The performance of the estimator depends on the correct specification of a dimension parameter whose optimal choice relies on smoothness characteristics of both the intensity and the error density. We propose a data-driven choice of the dimension parameter based on model selection and show that the adaptive estimator attains the minimax optimal rate.

math.ST

Automated Cryptanalysis of Bloom Filter Encryptions of Health Records

Privacy-preserving record linkage with Bloom filters has become increasingly popular in medical applications, since Bloom filters allow for probabilistic linkage of sensitive personal data. However, since evidence indicates that Bloom filters lack sufficiently high security where strong security guarantees are required, several suggestions for their improvement have been made in literature. One of those improvements proposes the storage of several identifiers in one single Bloom filter. In this paper we present an automated cryptanalysis of this Bloom filter variant. The three steps of this procedure constitute our main contributions: (1) a new method for the detection of Bloom filter encrytions of bigrams (so-called atoms), (2) the use of an optimization algorithm for the assignment of atoms to bigrams, (3) the reconstruction of the original attribute values by linkage against bigram sets obtained from lists of frequent attribute values in the underlying population. To sum up, our attack provides the first convincing attack on Bloom filter encryptions of records built from more than one identifier.

cs.CR

A graph theoretic linkage attack on microdata in a metric space

Certain methods of analysis require the knowledge of the spatial distances between entities whose data are stored in a microdata table. For instance, such knowledge is necessary and sufficient to perform data mining tasks such as nearest neighbour searches or clustering. However, when inter-record distances are published in addition to the microdata for research purposes, the risk of identity disclosure has to be taken into consideration again. In order to tackle this problem, we introduce a flexible graph model for microdata in a metric space and propose a linkage attack based on realistic assumptions of a data snooper's background knowledge. This attack is based on the idea of finding a maximum approximate common subgraph of two vertex-labelled and edge-weighted graphs. By adapting a standard argument from algorithmic graph theory to our setup, this task is transformed to the maximum clique detection problem in a corresponding product graph. Using a toy example and experimental results on simulated data show that publishing even approximate distances could increase the risk of identity disclosure unreasonably.

cs.CR