arXiv · 2610.00101
Weighted Data Selection: Sharp Upper-Half and Five-Dimensional Laws
Abstract
How much risk does a small reweighted training support retain? For finite weighted least squares with the minimum-norm learner, we prove the exact law $Γ_d(n)=3-n/d$ throughout $\lceil3d/2\rceil\leq n\leq2d-1$. The guarantee covers every observed feature rank and uses selections that preserve the full feature span. Balanced simplex anchors reduce dimension; positive-weight lifting and independent-line compression close the risk bound. Shifted coordinate pairs attain the matching lower bound. The complete dataset-level upper bound and sharpness construction are verified in Lean 4. At the smaller budget $(d,n)=(5,6)$, we also prove $Γ_5(6)=11/5$, matching the simplex-block prediction from $5=3+2$ over arbitrary interacting configurations. Circuit covers, comparison second moments, and circuit-plane probabilities give the sharp excess $6/5$, while polar-face geometry resolves shared rank-three circuits. The general simplex-block frontier connects these laws within the intermediate-budget selection problem.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhongxuan Liu, Hongzhi Wang. 2026-09-08. Weighted Data Selection: Sharp Upper-Half and Five-Dimensional Laws. https://arxiv.org/abs/2610.00101
Cite the original work for its findings. Save a collection to share your selection of sources.