arXiv · 2609.32885
Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets
Abstract
We study whether model upgrades improve probability estimates for prediction-market questions. Our Resolved Market Forecasting (RMF) benchmark contains 3,000 resolved binary questions across nine domains, on which we evaluate six Claude and Qwen model variants using a question-only, zero-shot protocol. We assess Brier scores relative to an empirical base-rate predictor, examine their Murphy decomposition, and compare models through paired differences, with results stratified by event category and timing relative to training cutoffs. In the reported post-cutoff stratum, the four Claude models achieve Brier scores of 0.183-0.192, improving on the base-rate reference by 0.024-0.033. Qwen 32B does not significantly outperform that reference, although its paired Brier is 0.024 lower than that of the 7B checkpoint. The evaluated Claude version and tier upgrades yield no significant improvement. Within-model differences across event categories exceed the observed differences among Claude variants. These results show why forecasting scores should be interpreted alongside simple probability baselines and question composition: under this protocol, newer versions or higher model tiers do not consistently produce more accurate probabilities.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuanbo Li, Zekun Li, Xiaoyan cong. 2026-09-26. Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets. https://arxiv.org/abs/2609.32885
Cite the original work for its findings. Save a collection to share your selection of sources.