arXiv · 2610.08227
How Many Independent Samples Does a Satellite Image Contain? Generalization Bounds for Spatially Dependent Data
Abstract
Machine learning classifiers for remote sensing imagery are typically evaluated as though every pixel were an independent sample. Spatial autocorrelation violates this assumption, since neighboring pixels carry redundant information which inflates sample sizes. How many independent samples does a satellite image actually contain? For an $n \times n$ image whose spatial correlation persists over a range of $r$ pixels, the effective sample size is $Θ(n^2/r^2)$, not $n^2$. We prove this as a finite-sample upper bound for classifiers on spatially correlated data, and show via a matching lower bound that the rate is tight, and no algorithm can do better. We extend the results to images with directional correlation and spatially varying correlation structure. Our result justifies spatial cross-validation since block holdout with separation proportional to the correlation range achieves optimal generalization guarantees, while random holdout can underestimate confidence interval widths by a factor proportional to $r$. We validate the theory on synthetic data and satellite image tiles from three sensors (Landsat 8, Sentinel-2, and Sentinel-1).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Robin Young. 2026-10-06. How Many Independent Samples Does a Satellite Image Contain? Generalization Bounds for Spatially Dependent Data. https://doi.org/10.1109/tgrs.2026.3739158
Cite the original work for its findings. Save a collection to share your selection of sources.