Which Reranking Conclusions Survive the Answer Interface? A Prospective Finite-Orbit Audit
Rerankers are increasingly evaluated through downstream language-model answers. This raises a retrieval-measurement question: if only the reader's answer interface changes, should we reach the same conclusion about BM25 versus BGE-v2-m3? We prospectively audit their claim-paired effect on RAGuard and FEVER with four readers. Retrieval policies, evidence, claims, and context depth remain fixed while semantic-to-label binding, A/B versus X/Y vocabulary, and option order form eight task-equivalent interfaces. We ask whether the estimated retrieval-policy effect, its ordering, or selection value changes. None of the six confirmatory settings showed statistically certified interface variation above the prespecified 0.015 materiality threshold, and none showed a certified reversal of the BM25-BGE ordering. Selector disagreement reaches 33.5% in one environment, yet none of eight environments establishes the prespecified material held-out value difference. These results do not support broad replicated instability, but they do not prove universal invariance: five settings remain too uncertain to satisfy the prespecified higher-order equivalence condition. They show why retrieval evaluations should separate policy level, interface stability, policy ordering, and selection value. When stability is unverified, a uniform average over the enumerated interfaces with explicit variation bounds avoids privileging one interface.