arXiv Science⌕ Search

arXiv · 2609.31019

Ensembles of Exactly Solved Subsamples for Clusterwise Regression: Trimming Without a Trimming Level

Abstract

Clusterwise least squares partitions regression data into K groups with separate linear fits. We study an ensemble whose base learner is exact: solve the problem to global optimality on each of B random subsamples of size m << n, extend each solution by nearest-surface assignment, align the labels, and combine the replicates by vote or by selection. Each replicate is then an empirical K-quantizer on m points, and the ensemble admits an exact analysis. A vote is correct at a unit once the probability of a clean subsample times the clean-data replicate accuracy exceeds one half, whatever the contaminating values; iterating the ensemble on the units it has not flagged gives a variant that estimates the trimming level rather than requiring it. With up to 20% of gross outliers in the response its worst-case accuracy was 0.89, against 0.80 for trimmed alternation at the true contamination fraction and less at every fixed level tried. Conditionally on the data the replicates are i.i.d., so the vote converges exponentially fast in B to the plurality partition of the replicate law, agreeing with the criterion minimiser outside a boundary set whose size depends on m and the micro-solver, not on B. For the two-group location model, subsamples of order 1/pi_min make a replicate right more often than wrong above a separation threshold. An O(n^3) enumeration gives the exact minimiser for two groups and one covariate, against which the theory is checked. On clean data the ensemble loses to multistart alternation at equal cost.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Samir Orujov. 2026-09-25. Ensembles of Exactly Solved Subsamples for Clusterwise Regression: Trimming Without a Trimming Level. https://arxiv.org/abs/2609.31019

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

glmSTARMA -- An R-Package for fitting autoregressive spatio-temporal models following generalized linear models

The R package glmSTARMA implements autoregressive models for spatio-temporal data at fixed locations, with time-invariant spatial dependency structure. We rely on generalized linear models methodology and unify several approaches for the analysis of spatial count time series. Such models allow the (conditional) mean of the response to depend on past observations, lagged (conditional) expectations, and covariates. The response can be a continuous or a discrete random variable. Additionally, the package develops inference for double generalized linear models, allowing the dispersion parameter(s) of the marginal distributions to be modeled similarly to the mean process. This is a new capability which introduces, for example, spatio-temporal volatility models, such as space-time GARCH processes, and count time series models with spatio-temporal overdispersion and underdispersion. We provide functions for model estimation, simulation, inference, and prediction. Its use is illustrated by data examples.

stat.CO↗

Scientific Data Analysis for Class-Informatics in Computational Taxonomy

In this A.I. era, Computational Taxonomy is proposed to study complex systems by analyzing their databases under taxonomic hierarchies abiding the Principle of Science by providing "good explanations". Comparisons among branches or classes are carried out by Scientific Data Analysis (SDA) paradigm that explores all potential associative patterns, including interacting effects of all high orders, and evaluate finite sample precisions for all information pieces individually by effectively making use of all variables' categorical nature. Under each comparison, all confirmed information pieces are collected and displayed along row-axis of a heatmap with all involved study-subjects on the column-axis. Each comparison's heatmap individually characterizes participating classes and study-subjects and simultaneously provides a scientific basis for outlier detection upon all non-participants. All these heatmaps then collectively constitutes so-called Class-informatics that offers good explanations based on characteristic of all classes and study-subjects. Computational Taxonomy's Class-informatics indeed resolves multiple fundamental issues: Tukey's more than 60 years outlier detection problem, issue of self-correction annotation, and a crucial check on assumption of information-content equality between testing and training data sets in Machine Learning. A showcase of Computational Taxonomy is exclusively illustrated on Iris data.

stat.CO↗

Contrast-Aware Annotation for Statistical Graphics: The ggtwotone Package for ggplot2

Readable annotations are essential for interpreting statistical graphics, yet conventional single-color lines and labels can lose visibility when they cross backgrounds with heterogeneous luminance. We introduce a contrast-aware annotation framework implemented in the ggtwotone R package for ggplot2. The framework combines dual-stroke rendering, adaptive text-color selection, and perceptually guided highlight palettes to improve annotation visibility across light and dark regions. A shared contrast-adjustment engine supports WCAG- and APCA-based contrast criteria, reducing the need for manual color adjustment. The package provides contrast-aware geoms for segments, curves, paths, mathematical functions, regression overlays, and text while remaining compatible with standard ggplot2 workflows. A simulation-based evaluation across heterogeneous background colors demonstrates improved worst-case contrast for dual-stroke annotations and adaptive text selection, while also identifying trade-offs in perceptual separation as the number of requested highlight colors increases. Applications to statistical graphics and scientific images illustrate the framework in practical visualization settings. Together, these tools provide a reproducible approach for incorporating contrast considerations directly into graphical annotation.

stat.CO↗