arXiv Science⌕ Search

arXiv · 2609.32478

Agnostic Smoothed Online Regression with Adversarial Responses

Abstract

We study smoothed online prediction with bounded adversarial responses. This widely studied framework bridges i.i.d. sampling and adversarial covariate selection through a smoothness parameter $\textsf{C}_{\textsf{cov}}$, which bounds conditional covariate densities relative to a fixed, unknown base measure. We propose \textsc{Hedge-Cover}, an information-theoretic algorithm that achieves sublinear regret $\widetilde{O}(\sqrt{\text{Pdim}(\mathcal{F}) \textsf{C}_{\textsf{cov}} T})$ for function classes with bounded pseudo-dimension. The algorithm aggregates a carefully constructed family of experts using \textsc{Hedge}, with a prior that links regret to the number of disagreements between a consistent selector and a target function. We bound this number by exploiting covariate smoothness. This answers an open problem posed in \cite{blanchard2025agnostic} on the minimax optimal adaptive regret of the smoothed online regression problem. We establish a matching lower bound for the class of linear predictors. The main intricacy of the lower bound lies in explicitly constructing a challenging sequential covariate distribution supported on mutually orthogonal hyperplanes. This construction may be of independent technical interest. Finally, we revisit the well-specified setting and quantify the effect of response noise. For conditionally $ν^2$-subGaussian responses, we extend the existing lower bound under realizable responses by showing that the minimax expected regret is $Ω((1\vee ν)\sqrt{(\textsf{C}_{\textsf{cov}}-1)dT})$ for a function class of VC dimension $d$. A corresponding upper bound for ERM matches this dependence on $ν$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xuanyu Chen, Yue Yu. 2026-09-26. Agnostic Smoothed Online Regression with Adversarial Responses. https://arxiv.org/abs/2609.32478

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Adaptively Truncated Signature-based Logistic Regression for Semi-parametric Functional Classification

Modern sensing technologies routinely generate multi-dimensional functional data, often accompanied by scalar covariates, motivating statistical methods that can jointly accommodate nonlinear functional effects, complex cross-component dependencies, and irregular sampling patterns. Existing functional logistic regression approaches rely on basis expansions whose performance is critically sensitive to the choice of basis, smoothing parameters, and sampling design; meanwhile, current signature-based methods require heuristic selection of the signature truncation order. We propose Adaptively Truncated Signature-based Logistic Regression (ATSLR), a semi-parametric classification framework that integrates path-signature representations with a data-driven procedure for selecting the signature truncation order via penalized empirical risk minimization. Our method is supported by rigorous theoretical guarantees, including the existence of an optimal truncation order, consistency of its adaptive estimator, convergence rates for the classifier risk, a finite and computable search bound, and a general error-propagation analysis for irregular sampling. Experiments on synthetic and real-world datasets demonstrate that our framework consistently outperforms traditional functional classifiers and fixed-order signature baselines in accuracy, robustness, and interpretability. Our results highlight the practical and theoretical value of integrating rough path theory with adaptive model complexity control.

stat.ML↗

Design of Experiment for Discovering Directed Mixed Graph

We study the design of interventions for causal discovery in simple structural causal models whose causal graphs are directed mixed graphs (DMGs) that may contain directed cycles and bidirected edges representing latent confounding. In such case, observational conditional-independence (CI) information may not identify even the graph skeleton, while CI alone cannot generally detect a bidirected edge coexisting with a directed edge. To this end, we propose a stage-wise framework based on tailored separating systems. Separating-system interventions first recover descendant relations and strongly connected components (SCCs). The SCCs are then ordered by ancestry, and an SCC-Anc separating system recovers the directed subgraph. Given this subgraph, further systems use CI tests interpreted through $d$- or $σ$-separation to recover non-adjacent bidirected edges, whose endpoints share no directed edge, and do-see comparisons to recover those coexisting with exactly one directed edge. Under our assumptions, the framework recovers the directed subgraph and every bidirected edge except double-adjacent ones, whose endpoints are connected by a directed edge in each direction. We develop algorithms for unrestricted and $M$-bounded settings, with each experiment targeting at most $M$ variables in the latter. For recovering the directed subgraph and non-adjacent bidirected edges, our upper bounds on the number and maximum size of experiments match corresponding worst-case lower bounds up to logarithmic factors.

stat.ML↗

Synthetic data for ratemaking: imputation-based methods vs adversarial networks and autoencoders

Actuarial ratemaking depends on high-quality data, yet access to such data is often limited by the cost of obtaining new data, privacy concerns, etc. In this paper, we explore synthetic-data generation as a potential solution to these issues. In addition to generative methods previously studied in the actuarial literature, we explore and benchmark another class of approaches based on Multivariate Imputation by Chained Equations (MICE). In a comparative study using an open-source dataset, MICE-based models are evaluated against other generative models like Variational Autoencoders and Conditional Tabular Generative Adversarial Networks. We assess how well synthetic data preserves the original marginal distributions of variables as well as the multivariate relationships among covariates. The consistency between Generalized Linear Models (GLMs) trained on synthetic data with GLMs trained on the original data is also investigated. Furthermore, we assess the ease of use of each generative approach and study the impact of generically augmenting original data with synthetic data on the estimation of GLMs for predicting claim counts. Our results highlight the potential of MICE-based methods in creating high-fidelity tabular data while offering lower implementation complexity compared to deep generative models.

stat.ML↗