arXiv · 2609.29549
StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection
Abstract
Post-training pipelines must select one language-model policy from many checkpoints, prompts, and decoding rules. Mean evaluator scores can conceal rare failures, whereas simultaneous candidate-wise confidence bounds can be unnecessarily conservative. We introduce StepCOPS, which uses an independent proposal split to nominate one lower-tail floor per candidate, exact binomial tests on a fresh certification split, and Holm's step-down procedure to certify a set of floors. With probability at least $1-δ$, every certified floor, including the largest floor used for policy selection, is below its candidate's population lower $α$-quantile. This guarantee assumes i.i.d. evaluation units while allowing arbitrary within-unit dependence across candidates. Across 24 predeclared configurations and 11 benchmarks, StepCOPS obtains 96.4% selected-policy coverage over 500 paired trials, raises the certified floor by 1.5 points over both proposal-Bonferroni and exact COPS, remains 0.6 points below the large-reference jury oracle, and abstains in 2.4% of trials. Shadow-judge, benchmark-native, artifact, and leave-one-judge-out audits characterize the proxy boundary: the guarantee applies to the fixed jury score, not directly to human safety.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma. 2026-08-26. StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection. https://arxiv.org/abs/2609.29549
Cite the original work for its findings. Save a collection to share your selection of sources.