arXiv ScienceSearch

arXiv subjects

Ramon Winterhalder

Publications and source records attributed to Ramon Winterhalder.

At least 19 recordsLinked to original sources

VERaiPHY -- Validation & Evaluation for Robust AI in PHYsics

Modern machine learning is leading to substantial gains in precision, flexibility, and computational efficiency in fundamental physics. Statistical validation, uncertainty quantification, and robustness assessment are less systematically addressed. The VERaiPHY initiative (Validation & Evaluation for Robust AI in PHYsics) is a series of articles developed within the PHYSTAT programme, aimed at establishing statistical standards for the development, evaluation, and deployment of ML techniques. Each article focuses on a specific methodological domain from a statistics perspective and clarifies statistical questions, tests, and the interpretation of results. This opening article establishes the probabilistic, statistical, and machine learning foundations that the later contributions assume, together with the notation used throughout.

hep-ph

Unifying Generative Models with Path Integrals

We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action. Its Martin-Siggia-Rose-Janssen-de~Dominicis (MSRJD) form separates free from interacting probability flows and opens them to diagrammatic perturbation theory. The expansion yields a one-loop correction to deterministic samplers at no stochastic-sampling cost, which we validate on solvable and nonlinear drifts, where it reduces a 53 % tree-level error to 1.6 %. Imperfect learned scores enter as insertions and yield a response-weighted score-matching objective, and symmetry-equivariant drift design becomes an operator expansion with EFT power counting.

cs.LG

The Living Guide of Machine Learning for Particle Physics

We started the Living Review of Machine Learning for Particle Physics (HEP-ML Living Review) in 2020 as a community-maintained, near-comprehensive bibliography of machine learning in particle physics. The field was then growing faster than any single researcher could follow, finding the relevant papers was hard, and a structured, continuously updated reference paid off immediately. Since then the literature has grown by more than an order of magnitude, the methods reach far beyond the classification and generation tasks of the early years, and the community has built its own ecosystem of topic-specific reviews, benchmark papers, and software frameworks. The original model no longer serves this field well, and we can no longer sustain it. We therefore change direction. We freeze the Living Review as an archival reference covering the literature up to 1 June 2026, where it remains a stable record of the first phase of HEP-ML. A new resource, the HEP-ML Living Guide, replaces it. It does not list everything. It curates, it annotates, and it points readers to foundational and representative work, so that researchers can find their way into a mature and rapidly diversifying field. In this article we explain why we make this change and how the new resource works.

hep-ph

Neural Control Variates at LO and NLO

We employ neural control variates to minimize the range of event weights and avoid negative weights for phase-space integration and event generation. A signed control variate, built from two normalizing flows, fulfills both tasks. Combined with neural importance sampling, it significantly reduces the computational cost of LO and NLO predictions. For the NLO case, our conditional neural control variate can be viewed as a trainable subtraction term, complementing the established physics subtraction schemes for enhanced sampling performance.

hep-ph

Interpreting Parton Distributions with Shapley Values

We show that Shapley values can be used to trace how individual parton distributions (PDFs) shape the theory predictions for high-energy observables computed from them. This provides a tool for assessing the impact of data on PDFs when determining them, and the impact of PDF uncertainties when using the PDFs to compute collider observables. The Shapley value is computed by treating the regression of PDFs from data as a cooperative game. The PDFs are the players, and the reward is the likelihood ($\chi^2$) that characterizes the agreement between data and the predictions obtained from a given PDF, with theory and methodology held fixed. The method is agnostic to the way PDFs have been determined in the first place: for PDFs determined with a black-box AI model it may be used in order to explain the behavior of the model, and for PDFs determined using a fixed parametrization it may be used in order to expose the features and potential limitations of the parametrization. We find that the method recovers known expectations about which data constrain which PDFs in a global fit, while placing them on a more quantitative footing. We demonstrate its effectiveness in two ways. We uncover an unexpected loss of sensitivity of the gluon PDF at intermediate $x$, with potential implications for BSM searches and the gluon fusion Higgs cross section. We also show that the method can be used to improve the hyperparameter optimization procedure currently used by the NNPDF collaboration.

hep-ph

The Monte Carlo Ecosystem in High-Energy Physics: A Primer

Monte Carlo event generators are the central interface between theoretical calculations and experimental measurements in collider physics. Over several decades, a comprehensive and highly modular ecosystem of tools has developed around them, encompassing matrix-element calculations, parton showers, hadronisation models, and their integration with detector simulation, event-level analysis and statistical inference. While these tools are ubiquitous in modern research, the conceptual scope and technical structure of the full simulation chain can be challenging to navigate, particularly for researchers entering the field. In this primer, we provide a structured and up-to-date overview of the high-energy physics Monte Carlo ecosystem, focusing primarily on event-generator methodologies and their role within the broader collider workflow. We discuss the conceptual foundations of modern generators, the computational and organisational challenges of large-scale simulations, and the principles that enable interoperability and reproducibility across theory and experiment. We also examine the evolving computing landscape and sustainability considerations that will shape the future development of these tools. Aimed primarily at early-stage doctoral researchers while serving as a reference for the broader community, this article seeks to clarify architecture, methodology, and long-term trajectory of Monte Carlo event generation in collider physics.

hep-ph

Uncertainty in Physics and AI: Taxonomy, Quantification, and Validation

Reliable uncertainty quantification is essential for the use of machine learning in physics, where scientific discoveries depend on validated probabilistic statements. We provide a structured overview of uncertainty quantification in ML for physics, introducing a unified taxonomy of uncertainty and clarifying the interpretation of predictive and inference uncertainties across frequentist and Bayesian frameworks. We discuss principled validation tools, including coverage, calibration, bias tests, and proper scoring rules, and illustrate them with simple regression and classification examples.

stat.ML

MadNIS at NLO

We combine fast amplitude surrogates with neural importance sampling to accelerate NLO calculations. For virtual corrections, a learned ratio to the Born matrix element with calibrated uncertainties guarantees reliable precision across phase space. For real emission, we stick to the standard FKS subtraction and train sector-conditioned surrogates of the regularized integrands away from divergences. MadNIS then uses multi-channel mappings and FKS sectors as conditions. We validate our approach for electron-positron scattering to three and four jets and find significant speed-ups and variance reduction in the integration.

hep-ph

MadSpace -- Event Generation for the Era of GPUs and ML

MadSpace is a new modular phase-space and event-generation library written in C++ with native GPU support via CUDA and HIP. It provides a unified compute-graph-based framework for phase-space construction, adaptive and neural importance sampling, and event unweighting. It includes a wide range of mappings, from the standard MadGraph multi-channel phase space to optimized normalizing flows with analytic inverse transformations. All components operate on batches of events and support end-to-end on-device workflows. A high-level Python interface enables seamless integration with machine-learning libraries such as PyTorch.

hep-ph

Amplitude Surrogates for Multi-Jet Processes

Accurate and efficient amplitude predictions are essential for precision studies of multi-jet processes at the LHC. We introduce a novel neural network architecture that predicts multi-jet amplitudes by leveraging the Catani-Seymour factorization scheme and related lower-jet amplitudes, requiring the network to learn only a correction factor. This hybrid approach combines theoretical factorization with a data-driven ansatz, enabling fast and scalable amplitude predictions. Our networks also estimate the accuracy of each prediction, allowing us to selectively use results that meet a predefined accuracy threshold. In the context of leading-order event generation, this approach achieves speed-up factors of up to 20 while maintaining all observables at the percent-level accuracy.

hep-ph

FASTColor -- Full-color Amplitude Surrogate Toolkit for QCD

High-multiplicity events remain a bottleneck for LHC simulations due to their computational cost. We present a ML-surrogate approach to accelerate matrix element reweighting from leading-color (LC) to full-color (FC) accuracy, building on recent advancements in LC event generation. Comparing a variety of modern network architectures for representative QCD processes, we achieve speed-up of around a factor two over the current LC-to-FC baseline. We also show how transformers learn and exploit underlying symmetries, to improve generalization. Given the gained trust in trained networks and developments in learned uncertainties, the LC-to-FC approach will eventually benefit further from not needing a final classic unweighting step.

hep-ph

Amplitude Uncertainties Everywhere All at Once

Ultra-fast, precise, and controlled amplitude surrogates are essential for future LHC event generation. First, we investigate the noise reduction and biases of network ensembles and outline a new method to learn well-calibrated systematic uncertainties for them. We also establish evidential regression as a sampling-free method for uncertainty quantification. In a second part, we tackle localized disturbances for amplitude regression and demonstrate that learned uncertainties from Bayesian networks, ensembles, and evidential regression all identify numerical noise or gaps in the training data.

hep-ph

The Physics Behind ML-based Quark-Gluon Taggers

Jet taggers provide an ideal testbed for applying explainability techniques to powerful ML tools. For theoretically and experimentally challenging quark-gluon tagging, we first identify the leading latent features that correlate strongly with physics observables, both in a linear and a non-linear approach. Next, we show how Shapley values can assess feature importance, although the standard implementation assumes independent inputs and can lead to distorted attributions in the presence of correlations. Finally, we use symbolic regression to derive compact formulas to approximate the tagger output.

hep-ph

BitHEP -- The Limits of Low-Precision ML in HEP

The increasing complexity of modern neural network architectures demands fast and memory-efficient implementations to mitigate computational bottlenecks. In this work, we evaluate the recently proposed BitNet architecture in HEP applications, assessing its performance in classification, regression, and generative modeling tasks. Specifically, we investigate its suitability for quark-gluon discrimination, SMEFT parameter estimation, and detector simulation, comparing its efficiency and accuracy to state-of-the-art methods. Our results show that while BitNet consistently performs competitively in classification tasks, its performance in regression and generation varies with the size and type of the network, highlighting key limitations and potential areas for improvement.

hep-ph

Accurate Surrogate Amplitudes with Calibrated Uncertainties

Neural networks for LHC physics have to be accurate, reliable, and controlled. Using neural surrogates for the prediction of loop amplitudes as a use case, we first show how activation functions are systematically tested with Kolmogorov-Arnold Networks. Then, we train neural surrogates to simultaneously predict the target amplitude and an uncertainty for the prediction. We disentangle systematic uncertainties, learned by a well-defined likelihood loss, from statistical uncertainties, which require the introduction of Bayesian neural networks or repulsive ensembles. We test the coverage of the learned uncertainties using pull distributions to quantify the calibration of cutting-edge neural surrogates.

hep-ph

Differentiable MadNIS-Lite

Differentiable programming opens exciting new avenues in particle physics, also affecting future event generators. These new techniques boost the performance of current and planned MadGraph implementations. Combining phase-space mappings with a set of very small learnable flow elements, MadNIS-Lite, can improve the sampling efficiency while being physically interpretable. This defines a third sampling strategy, complementing VEGAS and the full MadNIS.

hep-ph

Full and approximated NLO predictions for like-sign W-boson scattering at the LHC

We report on a recent calculation of next-to-leading-order (NLO) QCD and electroweak corrections to like-sign W-boson scattering at the Large Hadron Collider, including all partonic channels and W-boson decays in the process $pp \to e^+ \nu_e \mu^+ \nu_\mu jj + X$. The calculation is implemented in the Monte Carlo integrator Bonsay and comprises the full tower of NLO contributions of the orders $\alpha_s^3\alpha^4$, $\alpha_s^2\alpha^5$, $\alpha_s\alpha^6$, and $\alpha^7$. Our numerical results confirm and extend previous results, in particular the occurrence of large purely electroweak corrections of the order of $\sim-12\%$ for integrated cross sections, which get even larger in distributions. We construct a "VBS approximation" for the NLO prediction based on partonic channels and gauge-invariant (sub)matrix elements potentially containing the vector-boson scattering (VBS) subprocess and on resonance expansions of the Wdecays. The VBS approximation reproduces the full NLO predictions within $\sim1.5\%$ in the most important regions of phase space. Moreover, we discuss results from different versions of "effective vector-boson approximations" at leading order, based on the collinear emission of W bosons of incoming (anti)quarks. However, owing to the only mild collinear enhancement and the design of VBS analysis cuts, the quality of this approximation turns out to be only qualitative at the LHC.

hep-ph

The MadNIS Reloaded

In pursuit of precise and fast theory predictions for the LHC, we present an implementation of the MadNIS method in the MadGraph event generator. A series of improvements in MadNIS further enhance its efficiency and speed. We validate this implementation for realistic partonic processes and find significant gains from using modern machine learning in event generators.

hep-ph