arXiv · 2609.33052
BudgetVerify: Budget-Tiered Verification for Financial QA
Abstract
Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute. We propose BudgetVerify, a budget-tiered generator-verifier framework that routes each generated answer to one of three verification tiers: no verification, lightweight check-and-revise, or higher-cost solve-first-then-compare verification. The router is trained from offline correctness and token-cost outcomes and, at test time, selects a verification tier using information available before verification, including the question, context statistics, the generated answer, and associated generator metadata. The selected tier either returns the generated answer directly or invokes the corresponding verifier. Across six commercial and open-weight base models, BudgetVerify consistently produces more efficient accuracy-cost Pareto frontiers than fixed verification policies by selectively allocating stronger verification only when it is useful. Although absolute performance varies across models, these efficiency gains and the resulting qualitative frontier shape are consistent across generator models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Janet Jenq, Hongda Shen. 2026-09-27. BudgetVerify: Budget-Tiered Verification for Financial QA. https://arxiv.org/abs/2609.33052
Cite the original work for its findings. Save a collection to share your selection of sources.