arXiv ScienceSearch

arXiv · 2606.02508

AI and physics-based weather forecasting: A comparative study

Abstract

In the last few years, AI-based models have become the centre of attention in weather forecasting due to their increasing accuracy and efficiency. Pioneering among weather services, ECMWF has developed its Artificial Intelligence Forecasting System (AIFS) model, which was first to provide data-driven ensemble forecasts in June 2024. Since July 2025, the AIFS ensemble model has been operational and runs in parallel with ECMWF's physics-based Integrated Forecasting System (IFS), which is considered the gold standard in weather prediction. The new AIFS model can generate forecasts ten times faster than the classical numerical weather prediction model, while consuming approximately a thousand times less energy. We present the results of our systematic assessment of the performance of the IFS and AIFS models by comparing the accuracy of raw and post-processed medium-range 10-m wind-speed ensemble forecasts generated operationally by the two models for the period between July and November 2025 for more than 9000 synoptic observation stations across the globe. The post-processed case involves the parametric ensemble model output statistics (EMOS) as well as the non-parametric quantile regression (QR) approach to correct any systematic inaccuracies in the raw forecasts. The predictive performance of raw IFS ensemble forecasts proves to be substantially superior to the skill of the raw AIFS predictions for all investigated forecast horizons. As expected, post-processing significantly improves the skill of both IFS and AIFS predictions, and, across most verification metrics, EMOS is superior to QR, especially for short lead times. Compared to the raw ensemble, the differences in skill between the matching IFS and AIFS predictions are substantially decreased by post-processing and are mostly significant at short lead times, when the IFS forecasts outperform their AIFS counterparts.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mátyás Kocsis, Sándor Baran. 2026-06-01. AI and physics-based weather forecasting: A comparative study. https://arxiv.org/abs/2606.02508

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Uncertainty-Aware Missing-Data Multimodal Latent for Fetal-Growth Analysis

Objective: Routine third-trimester examination yields fetal biometry, maternal, Doppler and fetal-cardiac measurements, acquired at clinical discretion and therefore often incomplete. We show that these four measurement blocks are close to mutually uninformative, and that this single property determines what a representation of them can impute, what it can audit, and what fetal size alone cannot indicate. Methods: A linear-Gaussian factor model (K = 8 by parallel analysis, VARIMAX-rotated) was fitted to 25 measurements in four blocks from 977 fetuses (169 SGA, 61 severe; 77 LGA). Posterior precision sums contributions from observed measurements only, so missing values are marginalized rather than imputed. Data quality was screened using the standardized residual between each measurement and its reconstruction. Results: Predicting any one block from the other three gives an out-of-fold R2 of 0.023. The representation is a continuous growth spectrum with no cluster structure (Hartigan dip p = 0.99, gap statistic k = 1, three-cluster silhouette 0.07) ordering fetuses by birthweight centile (Spearman rho = 0.55). Among 169 SGA fetuses the haemodynamic redistribution axis separated the 25 adverse outcomes (AUC 0.70, 0.585-0.808) where measured size did not (0.60, 0.451-0.738). With Doppler censored, the marginalized interval covered held-out measurements in 97% and 93% of cases against nominal 95% and 90%. The reconstruction residual flagged 38 of 977 records, 36 confirmed transcription errors in the registry. Conclusion: Marginalizing missing measurements yields a representation whose uncertainty reflects the available data, and whose reconstruction residual doubles as a data-quality screen. Because the blocks are nearly independent, confirmed flags are within-block errors, and a synthetic benchmark gives the coupling needed before cross-block detection becomes available.

stat.AP

Auditing the Global Carbon Budget: Exploring the 2024--2025 Vintage Shift

The Global Carbon Budget (GCB), the community reference dataset for the carbon cycle, is reissued annually. The 2025 release introduces several adjustments to the published series that we compare with prior releases starting in 2017. On a common 1959-2016 sample, the mean of the GCB budget imbalance jumps from within +/-0.17 GtC/yr of zero for every vintage 2017-2024 to 0.61 GtC/yr in 2025, the only vintage whose 95% confidence interval for the imbalance mean excludes zero. The size of the imbalance changes much less: its mean absolute value rises from 0.61 to 0.76 GtC/yr. It is the mean, the quantity the budget identity constrains, that moves. We document and explore this shift in two ways. First, we conduct a model-free analysis, where we attribute the shift to a new adjustment that places the published land sink 0.40 GtC/yr below its ensemble mean (the average of the underlying models), a smaller adjustment in the ocean sink in the opposite direction, and a change in the composition of the bookkeeping ensemble. Second, we consider a dynamic statistical GCB model augmented with climate covariates. Its parameters are estimated for every GCB vintage 2017-2025. The coefficients of atmospheric concentrations in the sink equations shift in opposite directions on the 2025 issue, mirroring the model-free findings. A constant in the budget equation, statistically unnecessary in every vintage from 2017-2024, is required in 2025 and is estimated at -0.59 (0.09) GtC/yr. There is a persistent drifting imbalance across the entire sample in the budget equation. Each of the three adjustments is documented in the 2025 release and rests on evidence about the component it corrects. Their joint effect is a budget that closes over the last ten years and carries a mean imbalance of 0.61 GtC/yr over the full record. We argue that this cost to the full sample outweighs the gain on the last ten years.

stat.AP

Bridging Network Psychometrics and Artificial Intelligence: An Ising-Potts Model with LLM-Derived Weights

The Potts model extends the Ising model to multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of ratings and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses on pairwise agreement among ratings and assigns category-specific positive weights, making it suited for multi-category scoring reliability. We evaluate the model on three constructed-response datasets spanning a corpus of K=14,466 short answers on a three-level rubric and two AERA essay prompts of roughly 1,200-1,400 responses on four-point rubrics. We compare three strategies for sharpening the similarity signal: top-K pruning, min-max normalization with a power transformation, and ColBERT late-interaction similarities. Top-K pruning, which replaces the dense similarity graph with a sparse local network of strongest semantic neighbors, consistently yields the highest accuracy and Cohen's kappa, and the selected neighborhoods are always a small fraction of the corpus. Power tuning consistently ranks second, while ColBERT is competitive on longer essay prompts and adds little on short answers. Across all settings, most misclassifications occur between adjacent score levels, confirming that the model preserves the ordinal structure of scoring rubrics without imposing rigid assumptions. These findings suggest that LLM-derived similarities, combined with a parsimonious Potts formulation and a sparse local graph, offer a robust and interpretable framework for reliability auditing in educational assessment. We discuss extensions to multiple raters and hierarchical rating designs.

stat.AP