Summary The Blue Waters system, installed in 2012 at NCSA, has the largest component count of any system Cray has built. Blue Waters includes a mix of dual‐socket CPU (XE) and single‐socket CPU, single GPU (XK) nodes. The primary storage is provided by Cray's Sonexion/ClusterStor Luster storage system delivering 35 PB (raw) storage at 1 TB/s. The statistical failure rates over time for each component including CPU, DIMM, GPU, disk drive, power supply, blower, etc and their impact on higher level failure rates for individual nodes and the systems as a whole are presented in detail, with a particular emphasis on identifying any increases in rate that might indicate the right‐side of the expected bathtub curve has been reached. Strategies employed by NCSA and Cray for minimizing the impact of component failure, such as the preemptive removal of suspect disk drives, are also presented.
Resilience engineering enables a complementary approach to Occupational Health and Safety (OSH) management in complex organizations, in compliance with the systems resilience paradigm.This work intends to analyze the performance in OSH, in a Portuguese local public administration organization, measuring its resilience capacity, evidenced by the four resilience potentials (capabilities to respond, monitor, learn and anticipate), based on the "Resilience Assessment Grid" (RAG).Through a simplified procedure of the Delphi methodology, 38 questions were internally validated, applied to the group of gardeners, and the frequency of responses was analyzed by radar charts.The response potential was the best evaluated, followed by the potential to learn, anticipate and finally the potential to monitor.For each potential improvement needs were identified towards future intervention.RAG application assumes a relevant contribution to the management of OSH, allowing the development of a new vision in accordance with Safety-II.Diagnosing the resilience potential of organizations, in a specific work context, through the perception of workers, allows approaching and understanding the variability of the system in terms of OSH.
We present an analysis of the collection of user support tickets created during nearly nine years of operation of the Blue Waters supercomputer. The analysis is based on information obtained from the Jira ticketing system and its corresponding queues. The paper contains a set of statistics showing, in quantitative form, the distribution of tickets across system areas. It also shows the computed metrics related to management of tickets by our staff. Additionally, we present an analysis, based on Machine Learning and Sentiment Analysis techniques, conducted over the text entered in tickets, targeting detecting trends on users' views and perspectives about the Blue Waters system. This kind of study, which is uncommon in the literature, could provide guidance for operators of future large systems about the expected volume of user support demanded by each system area, and about how to allocate support staff such that users receive the best possible assistance.
Large-scale hydrological models simulate watershed processes with applications in water resources, climate change, land use, and forecast systems. The quality of the simulations mainly depends on calibrating optimal sets of watershed parameters, a time-consuming task that highly demands computational resources from repeated simulations. This work aims at performance optimizations on the MGB ( "Modelo de Grandes Bacias ") hydrological model and the MOCOM-UA (Multi-Objective Complex Evolution) calibration method for two watersheds. The optimizations target state-of-the-art CPU/GPU systems, exploiting techniques that include AVX-512 vectorization, and multi-core (CPU) and many-core (GPU) parallelisms. Significant speedups of up to 20 x (CPU) were achieved for calibration, while the scalability analysis indicated 24 x (CPU) and 65 x (GPU) for simulations with larger problem sizes. The roofline analysis confirmed more effective use of the hardware resources, and the quantitative accuracy evaluation of the optimized implementations reached maximum relative errors of approximately 6% for discharges and objective functions.
Parametric computational modeling of galaxies is a process with a high computational cost. The statistical component of modeling, which may involve model refinements in relation to the source brightness distribution, achieves more satisfactory results when the Bayesian approach is employed. In our research, we use GALaxy PHotometric ATtributes (GALPHAT) as our primary tool for data processing. In the current scenario of cosmology, to be scientifically relevant, this type of modeling must be performed on thousands of galaxies. In this article, we present the study and optimization of solutions based on modern HPC platforms, including a many-core processor, that enable effective processing of that amount of galaxies obtained from Sloan Digital Sky Survey.
HiPC 2020 is the 27th edition of the IEEE International Conference on High Performance Computing, Data, and Analytics.The conference focus is not only HPC but also includes Data Science.Due to the COVID-19 pandemic, this year the conference will be held virtually on December 16, 17, and 18.Each day of the three day event will open with a keynote talk, followed by two one-hour live remote sessions to present the technical program of thirty-three peer reviewed papers.All papers accepted for the conference will be published as part of the proceedings that will be available before, during, and after the week of the scheduled conference.The online publication will include both papers and presentations (slides) for each paper.Access to the online publication is part of the free registration to attend the virtual live sessions.
Hydrological models are extensively used in applications such as water resources, climate change, land use, and forecast systems. The focus of this paper is performance optimization of the MGB hydrological model, which is widely employed to simulate water flows in large-scale watersheds. The optimization strategies that we selected include AVX-512 vectorization, thread-parallelism on multi-core CPUs (OpenMP), and data-parallelism on many-core GPUs (CUDA). We conducted experiments for real-world input datasets on state-of-the-art HPC systems based on Intel's Skylake CPUs and NVIDIA GPUs. In addition, a Roofline model characterization for these datasets confirmed performance improvements of up to 37.5x on the most time-consuming part of the code and 8.6x on the full MGB model. The work proposed herein shows that careful optimizations are needed for hydrological models to achieve a significant fraction of the performance potential in modern processors.
The Brazilian Earth System Model (BESM) is a Global Climate Model (GCM) developed by the Brazilian National Institute for Space Research (INPE). The main purpose of a GCM is to simulate Earth’s climate in a decadal or centennial scale. The simulations usually include representations of the main elements of the Earth, such as atmosphere, ocean, ice and land. Since its first release, BESM has provided support materials for contributions to the Intergovernmental Panel on Climate Change (IPCC). This paper evaluates BESM’s performance and explores optimization possibilities, aiming to speed up the model execution. Our study started with a detailed analysis that characterized the performance of BESM executions on hundreds of processors, which served to reveal the major performance bottlenecks. Next, we worked on schemes to mitigate some of those bottlenecks. The changes made so far resulted on performance gains up to a factor of 4 in some cases, when compared to the way it was previously being executed in production. We also describe ongoing work towards additional performance improvements. Despite presenting results only for BESM, our optimization techniques are applicable to other scientific, multi-physics models as well.
Purpose: To investigate age-related differences in outcomes of critically ill patients with sepsis around the world. Methods: We performed a secondary analysis of data from the prospective ICON audit, in which all adult ( >16 years ) patients admitted to participating ICUs between May 8 and 18, 2012, were included, except admissions for routine postoperative observation. For this sub-analysis, the 10,012 patients with completed age data were included. They were divided into five age groups - <= 50, 51-60, 61-70, 71-80, >80 years. Sepsis was defined as infection plus at least one organ failure. Results: A total of 2963 patients had sepsis, with similar proportions across the age groups (<= 50 = 25.2%: 51-60 = 30.3%; 61-70 = 32.8%; 71-80 = 30.7%; >80 = 30.9%). Hospital mortality increased with age and in patients >80 years was almost twice that of patients <= 50 years (493% vs 25.2%, p < .05). The maximum rate of increase in mortality was about 0.75% per year, occurring between the ages of 71 and 77 years. In multilevel analysis, age > 70 years was independently associated with increased risk of dying. Conclusions: The odds for death in ICU patients with sepsis increased with age with the maximal rate of increase occurring between the ages of 71 and 77 years. (C) 2019 Elsevier Inc. All rights reserved.
The Roofline model gives insights about the performance behavior of applications bounded by either memory or processor limits, providing useful guidelines for performance improvements. This work uses the Roofline model on the analysis of the MGB model that simulates hydrological processes in largescale watersheds. Real-world input data are used to characterize the performance on two multicore architectures, one with only CPUs and one with CPUs/GPU. The MGB model performance is improved with optimizations for better memory use, and also with shared-memory (OpenMP) and GPU (OpenACC) parallelism. CPU performance achieves 42.51 % and 50.17 % of each system’s peak, whereas GPU performance is low due to overheads caused by the MGB model structure.
This paper describes the improvement of computational performance of the BRASIL-SR model. This computational code, which is used for the estimate of incident surface solar radiation, was optimized with an application of OpenMP directives and code modifications for the implementation of data output (I/O) operations using NetCDF format files. We verified that speedup of the code improved with the (I/O) optimizations.Also, it exhibitied a more stable efficiency, which shows a better use of the processors from the supercomputer.
Este artigo apresenta a melhoria de desempenho computacional do modelo BRASIL-SR. Este código computacional, utilizado para a estimativa da irradiação solar incidente na superficie, foi otimizado com a aplicação de diretivas de OpenMP e modificações do código para implementação de operações de entrada e saída de dados (I/O) utilizando arquivos em formato NetCDF. Após estas implementações, verificou-se que o speedup do código foi ainda maior com um novo tratamento de I/O, gerando uma eficiência mais estável, o que mostra uma melhor utilização do supercomputador Santos Dumont.
SummaryTo achieve their mission and goals, HPC centers continually strive to improve the effectiveness of their resources and services to best serve their constituencies. Collectively, the community has learned a great deal about how to manage and operate HPC centers, provide robust and effective services, and develop new communities as well as about other important aspects. Yet, cataloguing best practices to help inform and guide the broader HPC community is not often done. To improve the situation, the Blue Waters project has documented sets of best practices that have been adopted for the deployment and operation over the past five years of the Blue Waters leadership system, a large Cray XE6/XK7 supercomputer at NCSA. Those practices, described in this paper, cover aspects of managing and operating the system and its resources, supporting its users, and expanding the diversity of applications and communities. Although the technical practices are sometimes discussed relative to Cray systems and leadership‐scale systems, we believe that they would benefit the deployment and operation of other large HPC installations as well.
Hydrological models are commonly employed to calculate water flows on rivers and watersheds for the analysis of extreme events in nature. Computations in these models can grow depending on the numerical method, and also on the spatial and temporal resolutions, thus affecting the model efficiency and utility. This work parallelizes the MGB hydrological model on either CPU with OpenMP or GPU with OpenACC, respectively, aiming at the improvement in performance by employing computing resources of an HPC system. An analysis of the sequential and parallel executions is presented together with the runtime, speedup, efficiency, and load balance achieved.
SummaryThe roofline analysis model is a visually intuitive performance model used to understand hardware performance limitations as well as potential benefits of optimizations for science and engineering applications. Intel Advisor has provided a useful roofline analysis feature since its version 2017 update 2, but it is not widely compatible with other compilers and chip‐architectures. As an alternative, we have employed Cray Performance Analysis Tools (CrayPat) that are more flexible for multiple compilers and architectures. First, we present our procedure for measuring a reliable computational intensity for roofline analysis. We performed several numerical studies for validation via manually derived reference data as well as data from Intel Advisor. Second, we provide roofline analysis results on Blue Waters for several HPC benchmarks and sparse linear algebra libraries. In addition, we present an example of roofline‐based performance projection for a future system.
This article describes the main features of the Brazilian Global Atmospheric Model (BAM), analyses of its performance for tropical rainfall forecasting, and its sensitivity to convective scheme and horizontal resolution. BAMis the new global atmospheric model of the Center for Weather Forecasting and Climate Research [Centro de Previsao de Tempo e Estudos Climaticos (CPTEC)], which includes a new dynamical core and state-of-the-art parameterization schemes. BAM's dynamical core incorporates a monotonic two-time-level semi-Lagrangian scheme, which is carried out completely on the model grid for the tridimensional transport of moisture, microphysical prognostic variables, and tracers. The performance of the quantitative precipitation forecasts (QPFs) from two convective schemes, the Grell-Devenyi (GD) scheme and its modified version (GDM), and two different horizontal resolutions are evaluated against the daily TRMM Multisatellite Precipitation Analysis over different tropical regions. Three main results are 1) the QPF skill was improved substantially with GDM in comparison to GD; 2) the increase in the horizontal resolution without any ad hoc tuning improves the variance of precipitation over continents with complex orography, such as Africa and South America, whereas over oceans there are no significant differences; and 3) the systematic errors (dry or wet biases) remain virtually unchanged for 5-day forecasts. Despite improvements in the tropical precipitation forecasts, especially over southeastern Brazil, dry biases over the Amazon and La Plata remain in BAM. Improving the precipitation forecasts over these regions remains a challenge for the future development of the model to be used not only for numerical weather prediction over South America but also for global climate simulations.
Deployment of a large parallel system typically involves several steps of preparation, delivery, installation, testing and acceptance, making such deployments a very complex process. Despite the availability of various petascale systems currently, the steps and lessons from their deployment are rarely described in the literature. This article documents our experiences from the deployment of the sustained petascale Blue Waters system at NCSA. Our presentation is focused on the final deployment steps, where the system was intensively tested and accepted by NCSA. Those experiences and lessons should be useful to guide similarly complex deployments of large systems in the future.
Supercomputers have seen an exponential increase in their size in the last two decades. Such a high growth rate is expected to take us to exascale in the timeframe 2018-2022. But, to bring a productive exascale environment about, it is necessary to focus on several key challenges. One of those challenges is fault tolerance. Machines at extreme scale will experience frequent failures and will require the system to avoid or overcome those failures. Various techniques have recently been developed to tolerate failures. The impact of these techniques and their scalability can be substantially enhanced by a parallel programming model called migratable objects. In this paper, we demonstrate how the migratable-objects model facilitates and improves several fault tolerance approaches. Our experimental results on thousands of cores suggest fault tolerance schemes based on migratable objects have low performance overhead and high scalability. Additionally, we present a performance model that predicts a significant benefit of using migratable objects to provide fault tolerance at extreme scale.