As a joint effort from various communities involved in the Worldwide LHC Computing Grid, the Operational Intelligence project aims at increasing the level of automation in computing operations and reducing human interventions. The distributed computing systems currently deployed by the LHC experiments have proven to be mature and capable of meeting the experimental goals, by allowing timely delivery of scientific results. However, a substantial number of interventions from software developers, shifters, and operational teams is needed to efficiently manage such heterogenous infrastructures. Under the scope of the Operational Intelligence project, experts from several areas have gathered to propose and work on "smart" solutions. Machine learning, data mining, log analysis, and anomaly detection are only some of the tools we have evaluated for our use cases. In this community study contribution, we report on the development of a suite of operational intelligence services to cover various use cases: workload management, data management, and site operations.
Managing the data of scientific projects is an increasingly complicated challenge, which was historically met by developing experiment-specific solutions. However, the ever-growing data rates and requirements of even small experiments make this approach very difficult, if not prohibitive. In recent years, the scientific data management system Rucio has evolved into a successful open-source project that is now being used by many scientific communities and organisations. Rucio is incorporating the contributions and expertise of many scientific projects and is offering common features useful to a diverse research community. This article describes the recent experiences in operating Rucio, as well as contributions to the project, by ATLAS, Belle II, CMS, ESCAPE, IGWN, LDMX, Folding@Home, and the UK’s Science and Technology Facilities Council (STFC).
The ATLAS Experiment at the LHC generates petabytes of data that is distributed among 160 computing sites all over the world and is processed continuously by various central production and user analysis tasks. The popularity of data is typically measured as the number of accesses and plays an important role in resolving data management issues: deleting, replicating, moving between tapes, disks and caches. These data management procedures were still carried out in a semi-manual mode and now we have focused our efforts on automating it, making use of the historical knowledge about existing data management strategies. In this study we describe sources of information about data popularity and demonstrate their consistency. Based on the calculated popularity measurements, various distributions were obtained. Auxiliary information about replication and task processing allowed us to evaluate the correspondence between the number of tasks with popular data executed per site and the number of replicas per site. We also examine the popularity of user analysis data that is much less predictable than in the central production and requires more indicators than just the number of accesses.
For the last 10 years, the ATLAS Distributed Computing project has based its monitoring infrastructure on a set of custom designed dashboards provided by CERN. This system functioned very well for LHC Runs 1 and 2, but its maintenance has progressively become more difficult and the conditions for Run 3, starting in 2021, will be even more demanding; hence a more standard code base and more automatic operations are needed. A new infrastructure has been provided by CERN, based on InfluxDB as the data store and Grafana as the display environment. ATLAS has adapted and further developed its monitoring tools to use this infrastructure for data and workflow management monitoring and accounting dashboards, expanding the range of previous possibilities with the aim to achieve a single, simpler, environment for all monitoring applications. This document describes these tools and the data flows for monitoring and accounting.
Transparent use of commercial cloud resources for scientific experiments is a hard problem. In this article, we describe the first steps of the Data Ocean R&D collaboration between the high-energy physics experiment ATLAS together with Google Cloud Platform, to allow seamless use of Google Compute Engine and Google Cloud Storage for physics analysis. We start by describing the three preliminary use cases that were identified at the beginning of the project. The following sections then detail the work done in the data management system Rucio and the workflow management systems PanDA and Harvester to interface Google Cloud Platform with the ATLAS distributed computing environment, and show the results of the integration tests. Afterwards, we describe the setup and results from a full ATLAS user analysis that was executed natively on Google Cloud Platform, and give estimates on projected costs. We close with a summary and and outlook on future work.
Rucio, the distributed data management system of the ATLAS experiment already manages more than 400 Petabytes of physics data on the grid. Rucio was incrementally improved throughout LHC Run-2 and is currently being prepared for the HL-LHC era of the experiment. Next to these improvements the system is currently evolving into a full-scale generic data management system for application beyond ATLAS, or even beyond high-energy physics. This contribution focuses on the development roadmap of Rucio for LHC Run-3, such as event level data management, generic meta-data support and increased usage of networks and tapes. At the same time Rucio is evolving beyond the original ATLAS requirements. This includes additional authentication mechanisms, generic database compatibility, deployment and packaging of the software stack in containers, and a project paradigm shift to a full-scale open source project..
Rucio is an open-source software framework that provides scientific collaborations with the functionality to organize, manage, and access their data at scale. The data can be distributed across heterogeneous data centers at widely distributed locations. Rucio was originally developed to meet the requirements of the high-energy physics experiment ATLAS, and now is continuously extended to support the LHC experiments and other diverse scientific communities. In this article, we detail the fundamental concepts of Rucio, describe the architecture along with implementation details, and report operational experience from production usage.
For high-throughput computing the efficient use of distributed computing resources relies on an evenly distributed workload, which in turn requires wide availability of input data that is used in physics analysis. In ATLAS, the dynamic data placement agent C3PO was implemented in the ATLAS distributed data management system Rucio which identifies popular data and creates additional, transient replicas to make data more widely and more reliably available. This proceedings presents studies on the performance of C3PO and the impact it has on throughput rates of distributed computing in ATLAS. Furthermore, results of a study on popularity prediction using machine learning techniques are presented.
A search is presented for the direct pair production of the stop, the supersymmetric partner of the top quark, that decays through an R-parity-violating coupling to a final state with two leptons and two jets, at least one of which is identified as a b-jet. The data set corresponds to an integrated luminosity of 36.1 fb(-1) of proton-proton collisions at a center-of-mass energy of root s = 13 TeV, collected in 2015 and 2016 by the ATLAS detector at the LHC. No significant excess is observed over the Standard Model background, and exclusion limits are set on stop pair production at a 95% confidence level. Lower limits on the stop mass are set between 600 GeV and 1.5 TeV for branching ratios above 10% for decays to an electron or muon and a b-quark.
A search for heavy resonances decaying into a Higgs boson (H) and a new particle (X) is reported, utilizing 36.1 fb(-1) of proton-proton collision data at root s = 13 TeV collected during 2015 and 2016 with the ATLAS detector at the CERN Large Hadron Collider. The particle Xis assumed to decay to a pair of light quarks, and the fully hadronic final state XH -> q (q) over bar 'b (b) over bar is analysed. The search considers the regime of high XH resonance masses, where the X and H bosons are both highly Lorentz-boosted and are each reconstructed using a single jet with large radius parameter. A two-dimensional phase space of XH mass versus X mass is scanned for evidence of a signal, over a range of XH resonance mass values between 1 TeV and 4 TeV, and for X particles with masses from 50 GeV to 1000 GeV. All search results are consistent with the expectations for the background due to Standard Model processes, and 95% CL upper limits are set, as a function of XH and X masses, on the production cross-section of the XH -> q (q) over bar 'b (b) over bar resonance. (c) 2018 The Author(s). Published by Elsevier B.V.
A measurement of the production cross section for two isolated photons in proton-proton collisions at a center-of-mass energy of $\sqrt{s}=8$ TeV is presented. The results are based on an integrated luminosity of 20.2 fb$^{-1}$ recorded by the ATLAS detector at the Large Hadron Collider. The measurement considers photons with pseudorapidities satisfying $|\eta^{\gamma}|<1.37$ or ${1.56<|\eta^{\gamma}|<2.37}$ and transverse energies of respectively $E_{\mathrm{T,1}}^{\gamma}>40$ GeV and $E_{\mathrm{T,2}}^{\gamma}>30$ GeV for the two leading photons ordered in transverse energy produced in the interaction.The background due to hadronic jets and electrons is subtracted using data-driven techniques. The fiducial cross sections are corrected for detector effects and measured differentially as a function of six kinematic observables. The measured cross section integrated within the fiducial volume is $16.8 \pm 0.8$ pb. The data are compared to fixed-order QCD calculations at next-to-leading-order and next-to-next-to-leading-order accuracy as well as next-to-leading-order computations including resummation of initial-state gluon radiation at next-to-next-to-leading logarithm or matched to a parton shower, with relative uncertainties varying from 5% to 20%.
The production of exclusive gamma gamma -> mu(+)mu(-) events in proton-proton collisions at a centre-of-mass energy of 13 TeV is measured with the ATLAS detector at the LHC, using data corresponding to an integrated luminosity of 3.2 fb(-1). The measurement is performed for a dimuon invariant mass of 12 GeV < m(mu+mu-) < 70 GeV. The integrated cross-section is determined within a fiducial acceptance region of the ATLAS detector and differential cross-sections are measured as a function of the dimuon invariant mass. The results are compared to theoretical predictions both with and without corrections for absorptive effects. (c) 2017 The Author(s). Published by Elsevier B.V.
Bose-Einstein correlations between identified charged pions are measured for $p$+Pb collisions at $\sqrt{s_{\mathrm{NN}}}=5.02$ TeV using data recorded by the ATLAS detector at the LHC corresponding to a total integrated luminosity of $28$ $\mathrm{nb}^{-1}$. Pions are identified using ionization energy loss measured in the pixel detector. Two-particle correlation functions and the extracted source radii are presented as a function of collision centrality as well as the average transverse momentum ($k_{\mathrm{T}}$) and rapidity ($y^{\star}_{\pi\pi}$) of the pair. Pairs are selected with a rapidity $-2<y^{\star}_{\pi\pi}<1$ and with an average transverse momentum $0.1<k_{\mathrm{T}}<0.8$ GeV. The effect of jet fragmentation on the two-particle correlation function is studied, and a method using opposite-charge pair data to constrain its contributions to the measured correlations is described. The measured source sizes are substantially larger in more central collisions and are observed to decrease with increasing pair $k_{\mathrm{T}}$. A correlation of the radii with the local charged-particle density is demonstrated. The scaling of the extracted radii with the mean number of participating nucleons is also used to compare a selection of initial-geometry models. The cross-term $R_\mathrm{ol}$ is measured as a function of rapidity, and a nonzero value is observed with $5.1\sigma$ combined significance for $-1<y^{\star}_{\pi\pi}<1$ in the most central events.
The centrality dependence of the mean chargedparticle multiplicity as a function of pseudorapidity is measured in approximately 1 μb−¹ of proton–lead collisions at a nucleon–nucleon centre-of-mass energy of √sNN=5.02 TeV using the ATLAS detector at the Large Hadron Collider. Charged particles with absolute pseudorapidity less than 2.7 are reconstructed using the ATLAS pixel detector. The p + Pb collision centrality is characterised by the total transverse energy measured in the Pb-going direction of the forward calorimeter. The charged-particle pseudorapidity distributions are found to vary strongly with centrality, with an increasing asymmetry between the proton-going and Pb-going directions as the collisions become more central. Three different estimations of the number of nucleons participating in the p+Pb collision have been carried out using the Glauber model as well as two Glauber–Gribov inspired extensions to theGlauber model. Charged-particle multiplicities per participant pair are found to vary differently for these three models, highlighting the importance of including colour fluctuations in nucleon–nucleon collisions in the modelling of the initial state of p + Pb collisions.
Citation Aad, G., B. Abbott, J. Abdallah, S. Abdel Khalek, O. Abdinov, R. Aben, B. Abi, et al. 2015. “Erratum to: Search for production of WW / WZ resonances decaying to a lepton, neutrino and jets in pp collisions at documentclass[12pt]{minimal} usepackage{amsmath} usepackage{wasysym} usepackage{amsfonts} usepackage{amssymb} usepackage{amsbsy} usepackage{mathrsfs} usepackage{upgreek} setlength{oddsidemargin}{-69pt} begin{document}$$sqrt{s}=8$$end{document}s=8 TeV with the ATLAS detector.” The European Physical Journal. C, Particles and Fields 75 (8): 370. doi:10.1140/epjc/s10052-0153593-4. http://dx.doi.org/10.1140/epjc/s10052-015-3593-4.