Abstract Just as data sets from next-generation experiments grow, processing requirements for physics analysis become more computationally demanding, necessitating performance optimizations for RooFit. One possibility to speed-up minimization and add stability is the use of Automatic Differentiation (AD). Unlike for numerical differentiation, the computation cost scales linearly with the number of parameters, making AD particularly appealing for statistical models with many parameters. In this paper, we report on one possible way to implement AD in RooFit. Our approach is to add a facility to generate C++ code for a full RooFit model automatically. Unlike the original RooFit model, this generated code is free of virtual function calls and other RooFit-specific overhead. In particular, this code is then used to produce the gradient automatically with Clad. Clad is a source transformation AD tool implemented as a plugin to the clang compiler, which automatically generates the derivative code for input C++ functions. We show results demonstrating the improvements observed when applying this code generation strategy to HistFactory and other commonly used RooFit models. HistFactory is the subcomponent of RooFit that implements binned likelihood models with probability densities based on histogram templates. These models frequently have a very large number of free parameters and are thus an interesting first target for AD support in RooFit.
Statistical models in high-energy physics formally encode the relationship between observed data, physics parameters of interest, and experimental and theoretical uncertainties. Likelihood-based inference is the central tool for precision measurements, effective field theory fits, and cross-analysis combinations. Consequently, there is an increasing need for machine-readable, descriptive, and portable model representations. Existing formats such as ROOT workspaces, pyhf JSON, and CMS DataCards provide valuable capabilities but remain tied to specific software stacks and offer no universal standard for exchange, validation, or long-term preservation. We introduce HS3, the High-Energy Physics Statistics Serialization Standard, an implementation-agnostic, human-readable, and extensible serialization format for statistical models. HS3 is designed such that new statistical constructs can be incorporated through backward-compatible extensions, while inference procedures and implementation-specific execution details remain the responsibility of downstream frameworks. HS3 represents likelihoods as computational graphs composed of named distributions, functions, datasets, domains, and analysis prescriptions. It supports binned and unbinned likelihoods as well as hierarchical composite models. HS3 is convertible from and to ROOT/RooFit and is a superset of pyhf. We describe the design principles, structure, and semantics of HS3 and summarize existing implementations in C++, Python, and Julia. We also present early applications to public likelihoods on HEPData, cross-framework validation, and reproducibility efforts. HS3 provides a foundation for FAIR (Findable, Accessible, Interoperable, Reusable), long-lived statistical models at the LHC and beyond. The standard is intended to serve the broader scientific community and to evolve over time for application across a wide range of domains.
With the growing datasets of HE(N)P experiments, statistical analysis becomes more computationally demanding, requiring improvements in existing statistical analysis algorithms and software. One way forward is to use Machine Learning (ML) techniques to approximate the otherwise untractable likelihood ratios. Likelihood fits in HEP are often done with RooFit, a C++ framework for statistical modeling that is part of ROOT. This contribution demonstrates how learned likelihood ratios can be used in RooFit analyses, showcasing new RooFit features that were developed for that purpose. Since ML models are often created with Python libraries, this necessitated new RooFit “pythonizations”, e.g. for using Python functions as RooFit functions in general. Some of these pythonizations were only possible by a major PyROOT upgrade that was undertaken this year. Therefore, this contribution will also summarize the new PyROOT features, showcasing the new features of both RooFit and PyROOT that benefit the users of the latest ROOT versions.
ROOT is a software toolkit at the core of LHC experiments and High Energy and Nuclear Physics (HENP) collaborations worldwide, widely used by the community and in continuous development with it. The package is available through many channels that cater different types of users with different needs. This ranges from software releases (called LCG stacks) provided via a distributed file system developed at CERN (CVMFS) for all HENP users to benefit, to pre-built binaries available on the three major platforms (Linux, MacOS, Windows), to more specialised packaging systems such as Homebrew, Snap, conda. The last example is one of the main systems to distribute software to a Python user base, particularly beneficial for complex environments with realworld scientific applications in mind such as those found in HENP. Nonetheless, a conda package installer tool can’t be assumed to exist on a machine if Python is installed. This in turn adds an additional heavyweight dependency for both the user and the packager. Instead, the standard Python implementation defaults to using pip as a package installer. This technology, together with the Python Package Index (PyPI), distributes many Python packages and has the advantage of providing a lightweight path to downstream development of a package with some upstream Python dependencies. This contribution highlights the steps required towards making pip install ROOT possible, demonstrating its availability as an early-stage release, and discussing some of the unique challenges of delivering a highly-performant multi-language software via the standard Python packaging system.
With the growing datasets of current and next-generation HighEnergy and Nuclear Physics (HEP/NP) experiments, statistical analysis has become more computationally demanding. These increasing demands elicit improvements and modernizations in existing statistical analysis software. One way to address these issues is to improve parameter estimation performance and numeric stability using Automatic Differentiation (AD). AD’s computational efficiency and accuracy are superior to the preexisting numerical differentiation techniques, and it offers significant performance gains when calculating the derivatives of functions with a large number of inputs, making it particularly appealing for statistical models with many parameters. For such models, many HEP/NP experiments use RooFit, a toolkit for statistical modeling and fitting that is part of ROOT. In this paper, we report on the effort to support the AD of RooFit likelihood functions. Our approach is to extend RooFit with a tool that generates overheadfree C++ code for a full likelihood function built from RooFit functional models. Gradients are then generated using Clad, a compiler-based source-codetransformation AD tool, using this C++ code. We present our results from applying AD to the entire minimization pipeline and profile likelihood calculations of several RooFit and HistFactory models at the LHC-experiment scale. We show significant reductions in calculation time and memory usage for the minimization of such likelihood functions. We also elaborate on this approach’s current limitations and explain our plans for the future.
RooFit is a toolkit for statistical modeling and fitting, and together with RooStats it is used for measurements and statistical tests by most experiments in particle physics, particularly the LHC experiments. As the LHC program progresses, physics analysis becomes more computationally demanding. Therefore, RooFit development in recent years is focused on modernizing RooFit, improving its ease of use, and on performance optimization. This paper presents the new RooFit vectorized computation mode, which supports calculations on the GPU. Additionally, we discuss new features in the upcoming ROOT 6.26 release, highlighting the new pythonizations in particular.
RooFit is a toolkit for statistical modeling and fitting used by most experiments in particle physics. Just as data sets from next-generation experiments grow, processing requirements for physics analysis become more computationally demanding, necessitating performance optimizations for RooFit. One possibility to speed-up minimization and add stability is the use of Automatic Differentiation (AD). Unlike for numerical differentiation, the computation cost scales linearly with the number of parameters, making AD particularly appealing for statistical models with many parameters. In this paper, we report on one possible way to implement AD in RooFit. Our approach is to add a facility to generate C++ code for a full RooFit model automatically. Unlike the original RooFit model, this generated code is free of virtual function calls and other RooFit-specific overhead. In particular, this code is then used to produce the gradient automatically with Clad. Clad is a source transformation AD tool implemented as a plugin to the clang compiler, which automatically generates the derivative code for input C++ functions. We show results demonstrating the improvements observed when applying this code generation strategy to HistFactory and other commonly used RooFit models. HistFactory is the subcomponent of RooFit that implements binned likelihood models with probability densities based on histogram templates. These models frequently have a very large number of free parameters and are thus an interesting first target for AD support in RooFit.
The B_{c}^{+} meson is observed for the first time in heavy ion collisions. Data from the CMS detector are used to study the production of the B_{c}^{+} meson in lead-lead (Pb-Pb) and proton-proton (pp) collisions at a center-of-mass energy per nucleon pair of sqrt[s_{NN}]=5.02 TeV, via the B_{c}^{+}→(J/ψ→μ^{+}μ^{-})μ^{+}ν_{μ} decay. The B_{c}^{+} nuclear modification factor, derived from the Pb-Pb-to-pp ratio of production cross sections, is measured in two bins of the trimuon transverse momentum and of the Pb-Pb collision centrality. The B_{c}^{+} meson is shown to be less suppressed than quarkonia and most of the open heavy-flavor mesons, suggesting that effects of the hot and dense nuclear matter created in heavy ion collisions contribute to its production. This measurement sets forth a promising new probe of the interplay of suppression and enhancement mechanisms in the production of heavy-flavor mesons in the quark-gluon plasma.
As the field of high energy physics moves to an era of precision measurements its models become ever more complex and so do the challenges for computational frameworks that intend to fit these models to data. This report describes two computational optimizations with which RooFit intends to address this challenge: parallelization and batched computations. For the former, a problem-agnostic parallelization framework was devised with generality in mind such that it could be seamlessly applied at various stages of the existing Minuit2 minimization routine. In the results shown in this report parallelization was applied at the gradient calculation stage. The batched computations approach on the other hand required an overhaul of the current manner in which RooFit prepares its computatational graph for the evaluation of likelihoods. This report includes initial benchmarks of the batched computations strategy run on a CPU with vector instructions. Both strategies show significant performance improvements and the parallelization approach at its current state also proves to be robust enough to consistently fit state of the art physics models to real LHC data. Future developments are targeted towards combining both technologies in RooFit in a production-ready state in a ROOT release in the near future.
A search is presented for heavy bosons decaying to Z($\nu\bar{\nu}$)V(qq'), where V can be a W or a Z boson. A sample of proton-proton collision data at $\sqrt{s} =$ 13 TeV was collected by the CMS experiment during 2016-2018. The data correspond to an integrated luminosity of 137 fb$^{-1}$. The event categorization is based on the presence of high-momentum jets in the forward region to identify production through weak vector boson fusion. Additional categorization uses jet substructure techniques and the presence of large missing transverse momentum to identify W and Z bosons decaying to quarks and neutrinos, respectively. The dominant standard model backgrounds are estimated using data taken from control regions. The results are interpreted in terms of radion, W' boson, and graviton models, under the assumption that these bosons are produced via gluon-gluon fusion, Drell-Yan, or weak vector boson fusion processes. No evidence is found for physics beyond the standard model. Upper limits are set at 95% confidence level on various types of hypothetical new bosons. Observed (expected) exclusion limits on the masses of these bosons range from 1.2 to 4.0 (1.1 to 3.7) TeV.
The second workshop on the HEP Analysis Ecosystem took place 23-25 May 2022 at IJCLab in Orsay, to look at progress and continuing challenges in scaling up HEP analysis to meet the needs of HL-LHC and DUNE, as well as the very pressing needs of LHC Run 3 analysis. The workshop was themed around six particular topics, which were felt to capture key questions, opportunities and challenges. Each topic arranged a plenary session introduction, often with speakers summarising the state-of-the art and the next steps for analysis. This was then followed by parallel sessions, which were much more discussion focused, and where attendees could grapple with the challenges and propose solutions that could be tried. Where there was significant overlap between topics, a joint discussion between them was arranged. In the weeks following the workshop the session conveners wrote this document, which is a summary of the main discussions, the key points raised and the conclusions and outcomes. The document was circulated amongst the participants for comments before being finalised here.
This document discusses the state, roadmap, and risks of the foundational components of ROOT with respect to the experiments at the HL-LHC (Run 4 and beyond). As foundational components, the document considers in particular the ROOT input/output (I/O) subsystem. The current HEP I/O is based on the TFile container file format and the TTree binary event data format. The work going into the new RNTuple event data format aims at superseding TTree, to make RNTuple the production ROOT event data I/O that meets the requirements of Run 4 and beyond.
Production cross sections of Υ(1S), Υ(2S), and Υ(3S) states decaying into μ^+μ^- in proton-lead (pPb) collisions are reported using data collected by the CMS experiment at √(s_NN) = 5.02 TeV. A comparison is made with corresponding cross sections obtained with pp data measured at the same collision energy and scaled by the Pb nucleus mass number. The nuclear modification factor for Υ(1S) is found to be R_pPb(Υ(1S)) = 0.806 ± 0.024 (stat) ± 0.059 (syst). Similar results for the excited states indicate a sequential suppression pattern, such that R_pPb(Υ(1S)) R_pPb(Υ(2S)) R_pPb(Υ(3S)). The suppression is much less pronounced in pPb than in PbPb collisions, and independent of transverse momentum p_T^Υ and center-of-mass rapidity y_CM^Υ of the individual Υ state in the studied range p_T^Υ 30 GeV/c and | y_CM^Υ| 1.93. Models that incorporate sequential suppression of bottomonia in pPb collisions are in better agreement with the data than those which only assume initial-state modifications.
A search for new heavy resonances decaying to pairs of bosons ($WW$, $WZ$, or $WH$) is presented. The analysis uses data from proton-proton collisions collected with the CMS detector at a center-of-mass energy of 13 TeV, corresponding to an integrated luminosity of $137\text{ }\text{ }{\mathrm{fb}}^{\ensuremath{-}1}$. One of the bosons is required to be a W boson decaying to an electron or muon and a neutrino, while the other boson is required to be reconstructed as a single jet with mass and substructure compatible with a quark pair from a $W$, $Z$, or Higgs boson decay. The search is performed in the resonance mass range between 1.0 and 4.5 TeV and includes a specific search for resonances produced via vector boson fusion. The signal is extracted using a two-dimensional maximum likelihood fit to the jet mass and the diboson invariant mass distributions. No significant excess is observed above the estimated background. Model-independent upper limits on the production cross sections of spin-0, spin-1, and spin-2 heavy resonances are derived as functions of the resonance mass and are interpreted in the context of bulk radion, heavy vector triplet, and bulk graviton models. The reported bounds are the most stringent to date.
ROOT is high energy physics' software for storing and mining data in a statistically sound way, to publish results with scientific graphics. It is evolving since 25 years, now providing the storage format for more than one exabyte of data; virtually all high energy physics experiments use ROOT. With another significant increase in the amount of data to be handled scheduled to arrive in 2027, ROOT is preparing for a massive upgrade of its core ingredients. As part of a review of crucial software for high energy physics, the ROOT team has documented its R&D plans for the coming years.
ROOT is high energy physics' software for storing and mining data in a statistically sound way, to publish results with scientific graphics. It is evolving since 25 years, now providing the storage format for more than one exabyte of data; virtually all high energy physics experiments use ROOT. With another significant increase in the amount of data to be handled scheduled to arrive in 2027, ROOT is preparing for a massive upgrade of its core ingredients. As part of a review of crucial software for high energy physics, the ROOT team has documented its R&D plans for the coming years.
Jets containing a prompt J$/\psi$ meson are studied in lead-lead collisions at a nucleon-nucleon center-of-mass energy of 5.02 TeV, using the CMS detector at the LHC. Jets are selected to be in the transverse momentum range of 30 $\lt p_\mathrm{T} \lt$ 40 GeV. The J$/\psi$ yield in these jets is evaluated as a function of the jet fragmentation variable $z$, the ratio of the J$/\psi$$p_\mathrm{T} $ to the jet $p_\mathrm{T}$. The nuclear modification factor, $R_\mathrm{AA}$, is then derived by comparing the yield in lead-lead collisions to the corresponding expectation based on proton-proton data, at the same nucleon-nucleon center-of-mass energy. The suppression of the J$/\psi$ yield shows a dependence on $z$, indicating that the interaction of the J$/\psi$ with the quark-gluon plasma formed in heavy ion collisions depends on the fragmentation that gives rise to the J$/\psi$ meson.
Angular distributions of the decay B$^+$ $\to$ K$^*$(892)$^+\mu^+\mu^-$ are studied using events collected with the CMS detector in $\sqrt{s}=$ 8 TeV proton-proton collisions at the LHC, corresponding to an integrated luminosity of 20.0 fb$^{-1}$. The forward-backward asymmetry of the muons and the longitudinal polarization of the K$^*$(892)$^+$ meson are determined as a function of the square of the dimuon invariant mass. These are the first results from this exclusive decay mode and are in agreement with a standard model prediction.