arXiv · 2609.32677
When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds
Abstract
Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks policies well overall. We formalize this gap as Improvement Fidelity, which asks whether proxy improvements preserve the sign and ordering of deployment improvements over the updates an improvement process actually proposes. We show that global policy accuracy need not guarantee update fidelity: operator shift and deployment response can create update-level errors, while candidate margins determine whether those errors change the replacement decision. We introduce PIVOT-KG, a paired, decision-aware validator that allocates scarce high-fidelity evaluation according to the expected reduction in selection regret per unit cost. Across 90 held-out roots in Leduc, Kuhn, and Melting Pot, proxy and deployment optimal sets are disjoint in 51 cases. In an eight-candidate HighwayEnv stress test, PIVOT-KG reduces mean improvement-selection regret from 0.0435 under the exact Uniform validation rule to 0.0055 at the primary budget. Together, these results show why reliable self-improvement should evaluate proposed improvements in the worlds they induce, while providing a practical rule for allocating scarce deployment evidence when it can affect the replacement decision.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ke Wang, Zijie Zhao, Zhiyi Yuan, Changlun Li. 2026-09-26. When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds. https://arxiv.org/abs/2609.32677
Cite the original work for its findings. Save a collection to share your selection of sources.