arXiv ScienceSearch

arXiv · 2608.29818

Decarbonising price formation: unit-level evidence on battery storage and the imbalance price in the GB Balancing Mechanism

Abstract

Renewables now dominate Great Britain's generation mix but rarely occupy the marginal price-setting position, which raises the question of which flexible technologies translate a renewable-rich system into real-time price formation. This study reconstructs the price-ranked edge of the eligible bid or offer stack in the GB Balancing Mechanism across 50,684 half-hourly Settlement Periods from 2023 to 2025 and attributes it to individual Balancing Mechanism Units, separating long-system bid-active from short-system offer-active conditions and retaining co-marginal ties. Batteries rose from 0.8% to 36.2% of bid-active and from 3.2% to 26.9% of offer-active marginality, displacing combined-cycle gas on the bid side and pumped storage on the offer side. Capacity-normalised marginal capture rose on both sides, and unit-quarter fixed effects show that an additional 100 MW of Balancing Mechanism active capacity was associated with a 0.55 and 0.48 percentage-point increase in quarterly bid- and offer-side marginal share. By 2025, battery actions were around £9 to 10/MWh more favourable than volume-matched non-battery alternatives drawn from the same ladder, yet batteries were under-represented by 12.9 percentage points in the highest-priced 5% of short-system periods. Individual batteries remain consistent with price-taking, while the fleet has become endogenous to routine balancing price formation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Robert Dalton, Aidan O'Sullivan. 2026-08-30. Decarbonising price formation: unit-level evidence on battery storage and the imbalance price in the GB Balancing Mechanism. https://arxiv.org/abs/2608.29818

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Uncertainty-Aware Missing-Data Multimodal Latent for Fetal-Growth Analysis

Objective: Routine third-trimester examination yields fetal biometry, maternal, Doppler and fetal-cardiac measurements, acquired at clinical discretion and therefore often incomplete. We show that these four measurement blocks are close to mutually uninformative, and that this single property determines what a representation of them can impute, what it can audit, and what fetal size alone cannot indicate. Methods: A linear-Gaussian factor model (K = 8 by parallel analysis, VARIMAX-rotated) was fitted to 25 measurements in four blocks from 977 fetuses (169 SGA, 61 severe; 77 LGA). Posterior precision sums contributions from observed measurements only, so missing values are marginalized rather than imputed. Data quality was screened using the standardized residual between each measurement and its reconstruction. Results: Predicting any one block from the other three gives an out-of-fold R2 of 0.023. The representation is a continuous growth spectrum with no cluster structure (Hartigan dip p = 0.99, gap statistic k = 1, three-cluster silhouette 0.07) ordering fetuses by birthweight centile (Spearman rho = 0.55). Among 169 SGA fetuses the haemodynamic redistribution axis separated the 25 adverse outcomes (AUC 0.70, 0.585-0.808) where measured size did not (0.60, 0.451-0.738). With Doppler censored, the marginalized interval covered held-out measurements in 97% and 93% of cases against nominal 95% and 90%. The reconstruction residual flagged 38 of 977 records, 36 confirmed transcription errors in the registry. Conclusion: Marginalizing missing measurements yields a representation whose uncertainty reflects the available data, and whose reconstruction residual doubles as a data-quality screen. Because the blocks are nearly independent, confirmed flags are within-block errors, and a synthetic benchmark gives the coupling needed before cross-block detection becomes available.

stat.AP

Auditing the Global Carbon Budget: Exploring the 2024--2025 Vintage Shift

The Global Carbon Budget (GCB), the community reference dataset for the carbon cycle, is reissued annually. The 2025 release introduces several adjustments to the published series that we compare with prior releases starting in 2017. On a common 1959-2016 sample, the mean of the GCB budget imbalance jumps from within +/-0.17 GtC/yr of zero for every vintage 2017-2024 to 0.61 GtC/yr in 2025, the only vintage whose 95% confidence interval for the imbalance mean excludes zero. The size of the imbalance changes much less: its mean absolute value rises from 0.61 to 0.76 GtC/yr. It is the mean, the quantity the budget identity constrains, that moves. We document and explore this shift in two ways. First, we conduct a model-free analysis, where we attribute the shift to a new adjustment that places the published land sink 0.40 GtC/yr below its ensemble mean (the average of the underlying models), a smaller adjustment in the ocean sink in the opposite direction, and a change in the composition of the bookkeeping ensemble. Second, we consider a dynamic statistical GCB model augmented with climate covariates. Its parameters are estimated for every GCB vintage 2017-2025. The coefficients of atmospheric concentrations in the sink equations shift in opposite directions on the 2025 issue, mirroring the model-free findings. A constant in the budget equation, statistically unnecessary in every vintage from 2017-2024, is required in 2025 and is estimated at -0.59 (0.09) GtC/yr. There is a persistent drifting imbalance across the entire sample in the budget equation. Each of the three adjustments is documented in the 2025 release and rests on evidence about the component it corrects. Their joint effect is a budget that closes over the last ten years and carries a mean imbalance of 0.61 GtC/yr over the full record. We argue that this cost to the full sample outweighs the gain on the last ten years.

stat.AP

Bridging Network Psychometrics and Artificial Intelligence: An Ising-Potts Model with LLM-Derived Weights

The Potts model extends the Ising model to multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of ratings and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses on pairwise agreement among ratings and assigns category-specific positive weights, making it suited for multi-category scoring reliability. We evaluate the model on three constructed-response datasets spanning a corpus of K=14,466 short answers on a three-level rubric and two AERA essay prompts of roughly 1,200-1,400 responses on four-point rubrics. We compare three strategies for sharpening the similarity signal: top-K pruning, min-max normalization with a power transformation, and ColBERT late-interaction similarities. Top-K pruning, which replaces the dense similarity graph with a sparse local network of strongest semantic neighbors, consistently yields the highest accuracy and Cohen's kappa, and the selected neighborhoods are always a small fraction of the corpus. Power tuning consistently ranks second, while ColBERT is competitive on longer essay prompts and adds little on short answers. Across all settings, most misclassifications occur between adjacent score levels, confirming that the model preserves the ordinal structure of scoring rubrics without imposing rigid assumptions. These findings suggest that LLM-derived similarities, combined with a parsimonious Potts formulation and a sparse local graph, offer a robust and interpretable framework for reliability auditing in educational assessment. We discuss extensions to multiple raters and hierarchical rating designs.

stat.AP