arXiv ScienceSearch

arXiv · 2503.19051

LHCb Stripping Project: Continuing to Fully and Efficiently Utilize Legacy Data

Abstract

The LHCb collaboration continues to heavily utilize the Run 1 and Run 2 legacy datasets well into Run 3. As the operational focus shifts from the legacy data to the live Run 3 samples, it is vital that a sustainable and efficient system is in place to allow analysts to continue to profit from the legacy datasets. The LHCb Stripping project is the user-facing offline data-processing stage that allows analysts to select their physics candidates of interest simply using a Python-configurable architecture. After physics selections have been made and validated, the full legacy datasets are then reprocessed in small time windows known as Stripping campaigns. Stripping campaigns at LHCb are characterized by a short development window with a large portion of collaborators, often junior researchers, directly developing a wide variety of physics selections; the most recent campaign dealt with over 900 physics selections. Modern organizational tools, such as GitLab Milestones, are used to track all of the developments and ensure the tight schedule is adhered to by all developers across the physics working groups. Additionally, continuous integration is implemented within GitLab to run functional tests of the physics selections, monitoring rates and timing of the different algorithms to ensure operational conformity. Outside of these large campaigns the project is also subject to nightly builds, ensuring the maintainability of the software when parallel developments are happening elsewhere.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nathan Grieser, Federico Leo Redi, Eduardo Rodrigues, Niladri Sahoo, Shuqi Sheng, Nicole Skidmore, Mark Smith, Aravind Venkateswaran. 2025-03-24. LHCb Stripping Project: Continuing to Fully and Efficiently Utilize Legacy Data. https://arxiv.org/abs/2503.19051

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Search for the decays $B_{(s)}^0\to J/ψγ$ at LHCb

A search for the rare decays $B_{(s)}^0\to J/ψγ$ is performed with proton-proton collision data collected by the LHCb experiment, corresponding to integrated luminosities of $3~\rm{fb}^{-1}$ at centre-of-mass energies of 7 and 8 TeV, and $6~\rm{fb}^{-1}$ at 13 TeV. Assuming no contribution from $B^0\to J/ψγ$ decay, an upper limit is set on the branching fraction $\mathcal{B}(B_{s}^0\to J/ψγ)<2.9\times10^{-6}$ at the 90% confidence level. If instead no contribution from $B_{s}^0\to J/ψγ$ decay is assumed, the limit is $\mathcal{B}(B^0\to J/ψγ)<2.5\times10^{-6}$ at the 90% confidence level. These results supersede the previous LHCb results, with the limit for $B_{s}^0\to J/ψγ$ improved by a factor of 2.5.

hep-ex

Atmospheric Neutrino Oscillations: the Full Picture

We present the first combined oscillation analysis of recent atmospheric neutrino datasets, featuring data from Super-Kamiokande, IceCube-DeepCore, and KM3NeT/ORCA together with reactor data from Daya Bay. Such combinations have long been considered infeasible outside experimental collaborations; we demonstrate that a unified physics model can simultaneously describe all datasets with no significant parameter tensions. Fitting 839\,048 events across 1536 ins with 91 parameters, our combined analysis yields competitive measurements of the neutrino mixing parameters, and prefers the Normal over the Inverted Mass Ordering at $3σ$ significance.

hep-ex

AgentRivet: an automated system for producing Rivet routines from journal publications

Particle physics collider experiments provide Rivet routines as part of the analysis preservation strategy for model-independent measurements. Rivet is a C++ toolkit that allow new theoretical models to be compared to the measurements, thus aiding the development and tuning of Monte Carlo event generators as well as searches for physics beyond the Standard Model. However, analysis coverage is known to be incomplete, with only 39% of measurements having documented and publicly available Rivet routines. In this article, we design and implement an automated workflow based on Large Language Models with the goal of providing the missing routines. This multi-step workflow, referred to as AgentRivet, extracts the physics analysis information from published papers and writes the missing Rivet routines, with intermediate code- and physics- reviews as part of an autonomous quality control. We report the results obtained using commercial Large Language Models, provided by OpenAI, Anthropic, and Google, for two recent measurements from the ATLAS and CMS experiments. We find that AgentRivet produces competent Rivet routines with few syntax errors. The physics fidelity of the routines is reasonable and follows the explanations given in the relevant publications. Nevertheless, physics-implementation issues do arise and are investigated using the artefacts produced by AgentRivet. The majority of physics implementation issues arise from subtle-but-ambiguous definitions in the given publication, although some models struggle to implement complex observables even when clear definitions are given.

hep-ex