arXiv ScienceSearch

arXiv subjects

Tim Kutta

Publications and source records attributed to Tim Kutta.

At least 19 recordsLinked to original sources

Online Change-Point Monitoring for Object-valued Time Series

We develop closed- and open-end procedures for monitoring changes in the marginal distribution of object-valued time series. The method combines two distance-based Hilbert-space embeddings, a monitoring-time-dependent projection, and self-normalization. It is computable entirely from pairwise distances, does not require long-run variance estimation, and admits exact recursive updates. Under weak temporal dependence, the closed-end null limit is a pivotal Brownian functional. For monitoring over an unbounded horizon, we introduce a growing monitoring-time weight, establish a maximal inequality and uniform remote-tail control, and derive an open-end pivotal limit together with an equivalent fixed-interval representation for critical-value simulation. Both procedures are consistent against fixed marginal changes and retain the $n^{-1/2}$ and $n^{-1/4}$ local detection boundaries associated with the linear and quadratic projection signals. Simulations with distribution-valued time series illustrate favorable numerical properties, and an application to monthly stock-return distributions illustrates the usefulness of our procedure. Further simulations with graph-valued time series and an application to spatial point processes of seismic activity are reported in the supplementary material.

math.ST

Change Point Detection and Localization in High-Dimensional Time Series

We present new inference tools for change point detection in high-dimensional time series. We discuss two distinct statistical applications: First, sequential change point testing in an incoming data-stream. Second, retrospective localization of multiple changes, with confidence intervals at a globally controlled error level. Test statistics are built on the maximum norm to generate power against sparse and asynchronous changes. Both problems are tackled by related multiscale statistics that search for changes in the data at many different levels of resolution. For fixed dimension, our statistical approaches can be validated using traditional H\"olderian invariance principles. In this paper, we present the high-dimensional analogue: H\"older-Gauss-approximations, which can be (roughly) interpreted as the Gaussian approximation for a H\"older-norm of the high-dimensional partial sum process. Such approximations are of interest beyond change point detection and can be used for other problems such as for stationarity testing in high dimensions. We evaluate finite-sample performance in a simulation study and give an application to air contamination due to wildfires in California, which occurs asynchronously across a panel of measuring stations.

stat.ME

Simultaneous Inference for Partially Observed Functional Time Series

Functional data analysis (FDA) provides statistical methods for analyzing samples of time-continuous stochastic processes. Measurements often arise in the form of sensor data for a key scientific variable. The practical problem of irregular sensor disruptions has fostered interest in analyzing partially observed random functions. Specifically, this paper is motivated by a time series of intermittently missing pollution data with dependence along pollution paths and missingness patterns. To allow statistical analysis, we develop the first inference methods for dependent, partially observed functional time series. Existing methods were not appropriate for this task, because they heavily rely on the independence of the data functions. Mathematically, we model data on the space of bounded functions equipped with the supremum norm. This allows simultaneous inference across the entire functional domain, including simultaneous confidence bands -- something existing Hilbert-space-based methods cannot provide. To study non-stationary trends along the time series, we extend state-of-the-art multiscale inference methods (originally developed for scalar data) to partially observed functions. The key application of the latter methods is testing for excessive pollution levels in inner cities. Our approach combines state-of-the-art Gaussian approximations with stochastic process theory. Interestingly, it also improves existing results for fully observed functional time series by avoiding a functional CLT.

stat.ME

Does PCA Work for Rough Functional Data?

Functional data analysis is concerned with the analysis of infinite-dimensional data functions. Functional principal component analysis (FPCA) is a key method to obtain finite-dimensional summaries. Consistency of FPCA has been theoretically established for sufficiently regular data functions. However, empirical evidence shows that FPCA can become severely inconsistent when the underlying functions are too rough. This paper provides the first theoretical explanation for this phenomenon. We propose a model that explicitly captures the roughness of functional data and allows us to quantify the resulting bias of FPCA, depending on the functional roughness. The model undergoes a phase transition marking the point at which FPCA becomes entirely uninformative. Based on these probabilistic results, we discuss diagnostic tests for informative principal components. As an additional contribution, we derive results on spectral statistics that may serve as a foundation for goodness-of-fit tests for rough functional data. Mathematically, our approach combines recent advances in random matrix theory and generic chaining with tools from FDA. We illustrate the effects of roughness on FPCA using simulations, as well as climate and environmental datasets.

stat.ME

Sequential Auditing for f-Differential Privacy

We present new auditors to assess Differential Privacy (DP) of an algorithm based on output samples. Such empirical auditors are common to check for algorithmic correctness and implementation bugs. Most existing auditors are batch-based or targeted toward the traditional notion of $(\varepsilon,\delta)$-DP; typically both. In this work, we shift the focus to the highly expressive privacy concept of $f$-DP, in which the entire privacy behavior is captured by a single tradeoff curve. Our auditors detect violations across the full privacy spectrum with statistical significance guarantees, which are supported by theory and simulations. Most importantly, and in contrast to prior work, our auditors do not require a user-specified sample size as an input. Rather, they adaptively determine a near-optimal number of samples needed to reach a decision, thereby avoiding the excessively large sample sizes common in many auditing studies. This reduction in sampling cost becomes especially beneficial for expensive training procedures such as DP-SGD. Our method supports both whitebox and blackbox settings and can also be executed in one-run frameworks.

cs.CR

Iterative Data-Consistent Inversion with Multiple Push-forward Constraints

A foundational challenge in uncertainty quantification involves estimating a probability measure on the space of uncertain parameters such that its push-forward through a computational model matches an observed probability measure on the output data associated with quantities of interest (QoI). When multiple, distinct sets of observational data are available, the desired parameter measure should simultaneously satisfy multiple push-forward constraints associated with various subsets of the QoI. In this work, we present a convergent measure-theoretic framework for solving this problem based on an iterative application of Data-Consistent Inversion (DCI). We first rigorously establish the theoretical optimality of the DCI solution to the standard problem, proving that it minimizes the $f$-divergence over the space of all possible pullback measures that satisfy the push-forward constraint. This optimality property provides the foundation for our iterative DCI scheme, which is shown to converge to a solution of the multiple push-forward constraint problem. This iterative solution minimizes the cumulative $f$-divergence across all constraints and, under uniform initializations, represents the maximal entropy solution (the I-projection) onto the intersection of the solution sets. We provide a rigorous convergence analysis for the proposed method and demonstrate its practical utility through numerical examples, including a high-dimensional parameter space governed by partial differential equations, where the iterative approach robustly avoids the complexities associated with approximating high-dimensional joint observed measures.

math.OC

Inference for Forecasting Accuracy: Pooled versus Individual Estimators in High-dimensional Panel Data

Panels with large time $(T)$ and cross-sectional $(N)$ dimensions are a key data structure in social sciences and other fields. A central question in panel data analysis is whether to pool data across individuals or to estimate separate models. Pooled estimators typically have lower variance but may suffer from bias, creating a fundamental trade-off for optimal estimation. We develop a new inference method to compare the forecasting performance of pooled and individual estimators. Specifically, we propose a confidence interval for the difference between their forecasting errors and establish its asymptotic validity. Our theory allows for complex temporal and cross-sectional dependence in the model errors and covers scenarios where $N$ can be much larger than $T$-including the independent case under the classical condition $N/T^2 \to 0$. The finite-sample properties of the proposed method are examined in an extensive simulation study.

stat.ME

Multiscale Change Point Detection for Functional Time Series

We study the problem of detecting and localizing multiple changes in the mean parameter of a Banach space-valued time series. The goal is to construct a collection of narrow confidence intervals, each containing at least one (or exactly one) change, with globally controlled error probability. Our approach relies on a new class of weighted scan statistics, called H\"older-type statistics, which allow a smooth trade-off between efficiency (enabling the detection of closely spaced, small changes) and robustness (against heavier tails and stronger dependence). For Gaussian noise, maximum weighting can be applied, leading to a generalization of optimality results known for scalar, independent data. Even for scalar time series, our approach is advantageous, as it accommodates broad classes of dependency structures and non-stationarity. Its primary advantage, however, lies in its applicability to functional time series, where few methods exist and established procedures impose strong restrictions on the spacing and magnitude of changes. We obtain general results by employing new Gaussian approximations for the partial sum process in H\"older spaces. As an application of our general theory, we consider the detection of distributional changes in a data panel. The finite-sample properties and applications to financial datasets further highlight the merits of our method.

math.ST

TWIN: Two window inspection for online change point detection

We propose a new class of sequential change point tests, both for changes in the mean parameter and in the overall distribution function. The methodology builds on a two-window inspection scheme (TWIN), which aggregates data into symmetric samples and applies strong weighting to enhance statistical performance. The detector yields logarithmic rather than polynomial detection delays, representing a substantial reduction compared to state-of-the-art alternatives. Delays remain short, even for late changes, where existing methods perform worst. Moreover, the new procedure also attains higher power than current methods across broad classes of local alternatives. For mean changes, we further introduce a self-normalized version of the detector that automatically cancels out temporal dependence, eliminating the need to estimate nuisance parameters. The advantages of our approach are supported by asymptotic theory, simulations and an application to monitoring COVID19 data. Here, structural breaks associated with new virus variants are detected almost immediately by our new procedures. This indicates potential value for the real-time monitoring of future epidemics. Mathematically, our approach is underpinned by new exponential moment bounds for the global modulus of continuity of the partial sum process, which may be of independent interest beyond change point testing.

math.ST

Monitoring Violations of Differential Privacy over Time

Auditing differential privacy has emerged as an important area of research that supports the design of privacy-preserving mechanisms. Privacy audits help to obtain empirical estimates of the privacy parameter, to expose flawed implementations of algorithms and to compare practical with theoretical privacy guarantees. In this work, we investigate an unexplored facet of privacy auditing: the sustained auditing of a mechanism that can go through changes during its development or deployment. Monitoring the privacy of algorithms over time comes with specific challenges. Running state-of-the-art (static) auditors repeatedly requires excessive sampling efforts, while the reliability of such methods deteriorates over time without proper adjustments. To overcome these obstacles, we present a new monitoring procedure that extracts information from the entire deployment history of the algorithm. This allows us to reduce sampling efforts, while sustaining reliable outcomes of our auditor. We derive formal guarantees with regard to the soundness of our methods and evaluate their performance for important mechanisms from the literature. Our theoretical findings and experiments demonstrate the efficacy of our approach.

cs.CR

Monitoring Time Series for Relevant Changes

We consider the problem of sequentially testing for changes in the mean parameter of a time series, compared to a benchmark period. Most tests in the literature focus on the null hypothesis of a constant mean versus the alternative of a single change at an unknown time. Yet in many applications it is unrealistic that no change occurs at all, or that after one change the time series remains stationary forever. We introduce a new setup, modeling the sequence of means as a piecewise constant function with arbitrarily many changes. Instead of testing for a change, we ask whether the evolving sequence of means, say $(\mu_n)_{n \geq 1}$, stays within a narrow corridor around its initial value, that is, $\mu_n \in [\mu_1-\Delta, \mu_1+\Delta]$ for all $n \ge 1$. Combining elements from multiple change point detection with a H\"older-type monitoring procedure, we develop a new online monitoring tool. A key challenge in both construction and proof of validity is that the risk of committing a type-I error after any time $n$ fundamentally depends on the unknown future of the time series. Simulations support our theoretical results and we present two real-world applications: (1) healthcare monitoring, with a focus on blood glucose tracking, and (2) political consensus analysis via citizen opinion polls.

stat.ME

Monitoring for a Phase Transition in a Time Series of Wigner Matrices

We develop methodology and theory for the detection of a phase transition in a time-series of high-dimensional random matrices. In the model we study, at each time point \( t = 1,2,\ldots \), we observe a deformed Wigner matrix \( \mathbf{M}_t \), where the unobservable deformation represents a latent signal. This signal is detectable only in the supercritical regime, and our objective is to detect the transition to this regime in real time, as new matrix--valued observations arrive. Our approach is based on a partial sum process of extremal eigenvalues of $\mathbf{M}_t$, and its theoretical analysis combines state-of-the-art tools from random-matrix-theory and Gaussian approximations. The resulting detector is self-normalized, which ensures appropriate scaling for convergence and a pivotal limit, without any additional parameter estimation. Simulations show excellent performance for varying dimensions. Applications to pollution monitoring and social interactions in primates illustrate the usefulness of our approach.

math.ST

Prokhorov Metric Convergence of the Partial Sum Process for Reconstructed Functional Data

Motivated by applications in functional data analysis, we study the partial sum process of sparsely observed, random functions. A key novelty of our analysis are bounds for the distributional distance between the limit Brownian motion and the entire partial sum process in the function space. To measure the distance between distributions, we employ the Prokhorov and Wasserstein metrics. We show that these bounds have important probabilistic implications, including strong invariance principles and new couplings between the partial sums and their Gaussian limits. Our results are formulated for weakly dependent, nonstationary time series in the Banach space of d-dimensional, continuous functions. Mathematically, our approach rests on a new, two-step proof strategy: First, using entropy bounds from empirical process theory, we replace the function-valued partial sum process by a high-dimensional discretization. Second, using Gaussian approximations for weakly dependent, high-dimensional vectors, we obtain bounds on the distance. As a statistical application of our coupling results, we validate an open-ended monitoring scheme for sparse functional data. Existing probabilistic tools were not appropriate for this task.

math.ST

General-Purpose $f$-DP Estimation and Auditing in a Black-Box Setting

In this paper we propose new methods to statistically assess $f$-Differential Privacy ($f$-DP), a recent refinement of differential privacy (DP) that remedies certain weaknesses of standard DP (including tightness under algorithmic composition). A challenge when deploying differentially private mechanisms is that DP is hard to validate, especially in the black-box setting. This has led to numerous empirical methods for auditing standard DP, while $f$-DP remains less explored. We introduce new black-box methods for $f$-DP that, unlike existing approaches for this privacy notion, do not require prior knowledge of the investigated algorithm. Our procedure yields a complete estimate of the $f$-DP trade-off curve, with theoretical guarantees of convergence. Additionally, we propose an efficient auditing method that empirically detects $f$-DP violations with statistical certainty, merging techniques from non-parametric estimation and optimal classification theory. Through experiments on a range of DP mechanisms, we demonstrate the effectiveness of our estimation and auditing procedures.

cs.CR

Detection of a structural break in intraday volatility pattern

We develop theory leading to testing procedures for the presence of a change point in the intraday volatility pattern. The new theory is developed in the framework of Functional Data Analysis. It is based on a model akin to the stochastic volatility model for scalar point-to-point returns. In our context, we study intraday curves, one curve per trading day. After postulating a suitable model for such functional data, we present three tests focusing, respectively, on changes in the shape, the magnitude and arbitrary changes in the sequences of the curves of interest. We justify the respective procedures by showing that they have asymptotically correct size and by deriving consistency rates for all tests. These rates involve the sample size (the number of trading days) and the grid size (the number of observations per day). We also derive the corresponding change point estimators and their consistency rates. All procedures are additionally validated by a simulation study and an application to US stocks.

stat.ME

Testing separability for continuous functional data

Analyzing the covariance structure of data is a fundamental task of statistics. While this task is simple for low-dimensional observations, it becomes challenging for more intricate objects, such as multivariate functions. Here, the covariance can be so complex that just saving a non-parametric estimate is impractical and structural assumptions are necessary to tame the model. One popular assumption for space-time data is separability of the covariance into purely spatial and temporal factors. In this paper, we present a new test for separability in the context of dependent functional time series. While most of the related work studies functional data in a Hilbert space of square integrable functions, we model the observations as objects in the space of continuous functions equipped with the supremum norm. We argue that this (mathematically challenging) setup enhances interpretability for users and is more in line with practical preprocessing. Our test statistic measures the maximal deviation between the estimated covariance kernel and a separable approximation. Critical values are obtained by a non-standard multiplier bootstrap for dependent data. We prove the statistical validity of our approach and demonstrate its practicability in a simulation study and a data example.

stat.ME

Lower Bounds for R\'enyi Differential Privacy in a Black-Box Setting

We present new methods for assessing the privacy guarantees of an algorithm with regard to R\'enyi Differential Privacy. To the best of our knowledge, this work is the first to address this problem in a black-box scenario, where only algorithmic outputs are available. To quantify privacy leakage, we devise a new estimator for the R\'enyi divergence of a pair of output distributions. This estimator is transformed into a statistical lower bound that is proven to hold for large samples with high probability. Our method is applicable for a broad class of algorithms, including many well-known examples from the privacy literature. We demonstrate the effectiveness of our approach by experiments encompassing algorithms and privacy enhancing methods that have not been considered in related works.

cs.CR

Validating Approximate Slope Homogeneity in Large Panels

Statistical inference for large data panels is omnipresent in modern economic applications. An important benefit of panel analysis is the possibility to reduce noise and thus to guarantee stable inference by intersectional pooling. However, it is wellknown that pooling can lead to a biased analysis if individual heterogeneity is too strong. In classical linear panel models, this trade-off concerns the homogeneity of slope parameters, and a large body of tests has been developed to validate this assumption. Yet, such tests can detect inconsiderable deviations from slope homogeneity, discouraging pooling, even when practically beneficial. In order to permit a more pragmatic analysis, which allows pooling when individual heterogeneity is sufficiently small, we present in this paper the concept of approximate slope homogeneity. We develop an asymptotic level $\alpha$ test for this hypothesis, that is uniformly consistent against classes of local alternatives. In contrast to existing methods, which focus on exact slope homogeneity and are usually sensitive to dependence in the data, the proposed test statistic is (asymptotically) pivotal and applicable under simultaneous intersectional and temporal dependence. Moreover, it can accommodate the realistic case of panels with large intersections. A simulation study and a data example underline the usefulness of our approach.

stat.ME