Neutral Is Not Free: Evaluating Downside Risk in Neutral Launches
Evaluating "neutral launches" (e.g., infrastructure upgrades) using traditional confidence interval overlap is flawed: it is dangerously permissive with scarce data and excessively restrictive with abundant data. To resolve this, this paper introduces Expected Bayesian Loss (EBL), a continuous metric that quantifies both the probability and expected severity of metric degradation. Computable directly from standard frequentist estimates, EBL explicitly penalizes empirical noise and high-variance experiments. Validated against expert decisions, EBL provides experimentation platforms with a rigorous, tunable guardrail that aligns statistical safety with institutional risk appetite.