arXiv Science⌕ Search

arXiv · 2610.04558

ClimateBench v2.0: Probabilistic Climate Model Benchmarking

Abstract

We present ClimateBench v2, a standardized protocol for evaluating climate models on diagnostics expected to be informative for their skill in projecting mid-century regional temperature and precipitation changes. The protocol is designed to evaluate any physics-based, data-driven, or hybrid climate model on equal footing using a common set of observational and out-of-distribution tests. We define three tiers of evaluation. Tier I establishes physical credibility through entry-ticket tests of energy conservation, coupled (co-)variability, and basic forced responses. Tier II scores models against post-2015 observations of surface temperature, precipitation, radiative fluxes, sea ice, and key modes of variability using fair CRPS as the primary probabilistic score, complemented by distributional and ensemble-consistency diagnostics. Tier III tests out-of-distribution generalization through paleoclimate simulations spanning the Last Interglacial, Last Glacial Maximum, and Mid-Holocene, and through perfect-model experiments in which data-driven models must predict the future climate of existing Earth system models from historical data alone. We reserve all observational data after 2015 for testing, and submissions must include multiple ensemble members to enable probabilistic evaluation. This reservation exploits a new opportunity provided by the decade of observations accumulated since the end of the CMIP6 historical experiment, which constitutes an out-of-sample record of forced climate change (and internal variability) for the current generation of models, and we quantify, in an idealized setting, the information it carries about mid-century warming. We provide the evaluation code, observational reference datasets, and perfect-model training data as an open benchmark to drive measurable progress in climate projection across all modeling approaches.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Duncan Watson-Parris, Willa Tobin, Aytaç Paçal, Manuel Schlund, V. Balaji, Kevin Bowman, Chris Bretherton, Peter M. Caldwell, Will Chapman, William D. Collins, Gregory S. Elsaesser, Pierre Gentine, Helene Hewitt, Stephan Hoyer, Ralph Keeling, Nikolay Koldunov, David M. Lawrence, Christian Lessig, Daniel J. Lunt, J. David Neelin, Mike Pritchard, Sarah Purkey, Gavin Schmidt, Tapio Schneider, Michael Schulz, Tiffany Shaw, Isla R. Simpson, Graeme Stephens, Aneesh C. Subramanian, Joao Teixeira, Jessica Tierney, Andrew I. L. Williams, Laure Zanna, Veronika Eyring, Rose Yu. 2026-10-03. ClimateBench v2.0: Probabilistic Climate Model Benchmarking. https://arxiv.org/abs/2610.04558

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Can we create a `race to the top' for weather forecasts to inform smallholder farmer decisions?

Artificial-intelligence weather prediction (AIWP) models have made it possible to produce high-quality tailored forecasts with limited computational resources. This advance has the potential to benefit hundreds of millions of farmers in low- and middle-income countries who lack access to forecasts of critical weather phenomena. However, it can be difficult for key stakeholders to evaluate forecast quality, risking a "race to the bottom" as cheap but low-quality forecasts crowd out forecasts that would benefit farmers. We propose a set of principles and protocols for evaluating agriculturally-relevant forecasts as a starting point for standards that would let forecasters credibly convey their forecasts' quality.

physics.ao-ph↗

Artificial intelligence pathways from weather to climate

Deep learning has made rapid advances in weather forecasting: autoregressive models trained on atmospheric reanalyses now rival dynamical models across nowcasting, medium-range, and subseasonal-to-seasonal lead times, producing well-calibrated ensemble forecasts at reduced cost. We review these advances and consider their extension to climate horizons, where the challenge shifts from initial-condition skill to producing reliable statistical responses under altered forcings. AI-powered climate prediction systems must produce credible forced responses to drivers (e.g., greenhouse gases, land-use change) typically outside the observed record. We propose two minimum requirements for AI in climate modeling: (i) external forcing agents must enter explicitly enough to support interventions in which they vary independently; and (ii) robustness must be stress-tested in out-of-distribution regimes, including extremes and counterfactual trajectories. Using leading AI autoregressive emulators and hybrid physics-AI models, we identify development and coupling challenges. Comparing the reported throughput of these models with that of GPU-ported dynamical models highlights how AI can reduce time-to-solution by advancing only the target variables at the required resolution and using longer time steps, rather than integrating a full high-frequency, multivariate state. Diverse AI downscaling strategies can partially substitute for explicit fine-scale resolution, paving the way toward inexpensive local hazard assessment across prediction horizons.

physics.ao-ph↗

Generative and deterministic deep learning models comparison for fine-scale precipitation retrievals from infrared brightness temperature

Accurate precipitation estimation at fine spatial scales is critical for hydrology, agriculture, and climate studies. Infrared brightness temperatures from geostationary satellites offer excellent temporal coverage over continental-scale domains. However, because these measurements primarily characterize cloud-top properties rather than precipitation processes near the surface, their correlation with rainfall intensity remains limited, making quantitative precipitation estimation challenging. In this study, we conduct a systematic inter-comparison of state-of-the-art deep learning models for high-resolution precipitation retrieval from Meteosat Second Generation infrared brightness temperatures over metropolitan France. These models include deterministic U-Nets, transformer-based architectures, conditional GANs, and diffusion models. We construct a curated dataset spanning 2008--2023, combining M{é}t{é}o-France radar mosaics as reference with multi-channel infrared observations, and design preprocessing and sampling strategies to address the heavy-tailed, intermittent nature of rainfall. Our results show that deterministic models provide robust mean estimates and excel in pixel-wise accuracy, but systematically underestimate extreme precipitation. In contrast, generative models better capture the full precipitation distribution, including rare and heavy rainfall events, producing more realistic spatial structures at the cost of reduced pixel-wise fidelity. These results highlight a trade-off between pixel-wise accuracy and precipitation variability, showing that generative approaches are advantageous for extreme-event detection and probabilistic applications. This work establishes a reproducible framework for evaluating infrared- based precipitation retrieval methods and provides guidance for designing models that balance precision, variability, and extreme-event representation.

physics.ao-ph↗