While load balancing in distributed-memory computing has been well-studied, we present an innovative approach to tackle challenges that an electromagnetic application poses due to irregular workloads and tight memory constraints. To this end, we present a unified model for approximating work in a distributed system that combines three key components: computation, communication, and memory. This enables the exploration of complex trade-offs in task placement, such as increased parallelism at the expense of data replication. We then present our new fully distributed load balancing strategy that incorporates this model. To predict workloads for the matrix assembly of the electromagnetics application, we apply machine learning across an ensemble of executions to train a neural network, which makes online predictions for our task-based decomposition, informing the load balancer of the computational loads. Finally, we demonstrate that our approach, when applied to this application, leads to substantial speedups, up to 2.0x, thereby decreasing time-to-solution for the imbalanced execution.
Contact mechanics, or the modeling of the impenetrability of solid objects, is fundamental to computational solid mechanics (CSM) applications yet is oftentimes the most challenging in terms of computational efficiency and performance. These challenges arise from the irregularity and highly dynamic nature of contact simulation, particularly with algorithms designed for distributed memory architectures. First among these challenges is the inherent load imbalance when distributing contact load across compute nodes. This imbalance is highly problem dependent, and relates to the surface area of contact manifolds and the volume around them, rather than the distribution of the mesh over compute nodes, meaning the application load can vary drastically over different phases. The dynamic nature of contact problems motivates the use of distributed asynchronous many-tasking (AMT) frameworks to efficiently handle irregular workloads. In this paper, we present our work on distBVH, a distributed contact solution using the DARMA/vt library for asynchronous tasking that is also capable of running on-node Kokkos-based kernels. We explore how distBVH addresses the various challenges of CSM contact problems. We evaluate the use of many of DARMA/vt's dynamic load balancers and demonstrate how our load balancing approach can provide significant performance improvements on various computational solid mechanics benchmarks. Additionally, we show how our approach can take advantage of DARMA/vt for tasking and efficient on-node kernels using Kokkos to scale over hundreds of processing elements.
While load balancing in distributed-memory computing has been well-studied, we present an innovative approach to this problem: a unified, reduced-order model that combines three key components to describe "work" in a distributed system: computation, communication, and memory. Our model enables an optimizer to explore complex tradeoffs in task placement, such as increased parallelism at the expense of data replication, which increases memory usage. We propose a fully distributed, heuristic-based load balancing optimization algorithm, and demonstrate that it quickly finds close-to-optimal solutions. We formalize the complex optimization problem as a mixed-integer linear program, and compare it to our strategy. Finally, we show that when applied to an electromagnetics code, our approach obtains up to 2.3x speedups for the imbalanced execution.
This paper explores dynamic load balancing algorithms used by asynchronous many-task (AMT), or ‘taskbased’, programming models to optimize task placement for scientific applications with dynamic workload imbalances. AMT programming models use overdecomposition of the computational domain. Overdecompostion provides a natural mechanism for domain developers to expose concurrency and break their computational domain into pieces that can be remapped to different hardware. This paper explores fully distributed load balancing strategies that have shown great promise for exascale-level computing but are challenging to theoretically reason about and implement effectively. We present a novel theoretical analysis of a gossip-based load balancing protocol and use it to build an efficient implementation with fast convergence rates and high load balancing quality. We demonstrate our algorithm in a next-generation plasma physics application (EMPIRE) that induces time-varying workload imbalance due to spatial non-uniformity in particle density across the domain. Our highly scalable, novel load balancing algorithm, achieves over a 3x speedup (particle work) compared to a bulk-synchronous MPI implementation without load balancing.
We present the execution model of Virtual Transport (VT) a new, Asynchronous Many-Task (AMT) runtime system that provides unprecedented integration and interoperability with MPI. We have developed VT in conjunction with large production applications to provide a highly incremental, high-value path to AMT adoption in the dominant ecosystem of MPI applications, libraries, and developers. Our aim is that the `MPI+X' model of hybrid parallelism can smoothly extend to become `MPI+VT +X'. We illustrate a set of design and implementation techniques that have been useful in building VT. We believe that these ideas and the code embodying them will be useful to others building similar systems, and perhaps provide insight to how MPI might evolve to better support them. We motivate our approach with two applications that are adopting VT and have begun to benefit from increased asynchrony and dynamic load balancing.
The goal of this report is to provide a comprehensive status report of the research & development conducted in the context of the DARMA project by the end of the first quarter of fiscal year 2020. It follows in particular [LBS+19] and [PL19].
The goal of this report is to illustrate the use of Sandia's Automatic Report Generator (ARG), when applied to an Electrostatic simulation case run with Sandia's EMPIRE code. It documents the results of a hackathon session that was held at the March 19-22 DOE Workshop Workflow and Hackathon that was held in Livermore, where the co-authors demonstrated ARG's flexibilty by extending it to several aspect of such simulation in less than a day's worth of work. The Explorator component of ARG automatically picks up the case's input deck, hereby determining the data components that the Generator and Assembler components are currently able to document: meta-data, input deck, mesh, and solution fields. The ARG is not yet capable of documenting the particles file created by the simulation, which will require further work.
report includes findings on the generality of DARMAs backend API as well as findings on interoperability with node- level and network-level system libraries. Together, this information provides a clear understanding of the strengths and limitations of the DARMA approach in the context of Sandias ATDM codes, to guide our future research and development in this area.
Base, Colorado and Clear Air Force Station, Alaska for detailed use case SMR suitability analyses. The final report in December 2015 will assess the feasibility of SMRs for energy security and clean energy for Air Force Space Command (AFSPC) installations. This page intentionally left blank. EXECUTIVE SUMMARY During this first phase of the study, the team conducted broad research of existing federal and private sector policy, guidance, regulations, studies, and reports. Through this and other research the team identified major processes, potential impediments and issues, and key considerations that would affect SMR deployment on AFSPC installations. This research provided valuable information for development of use case installation selection criteria. Further, it will aid the development of recommendations for essential changes needed to facilitate successful and effective SMR deployment. With the assistance of members of the Headquarters, AFSPC staff and other DoD and Air Force representatives, the team also worked with major stakeholders to garner their inputs on factors affecting SMR deployment. The team also gathered information from the four US companies that are developing near term light-water SMRs on the operational performance characteristics and commercialization status of their technologies. Further, the team gathered perspectives of SMR deployment scenarios from utilities that are operating nuclear power plants and some that are serving AFSPC installations. This effort also included an assessment of the Lifecycle Cost of Energy for SMRs and preliminary consideration of economic factors that are critical to realistic commercial introduction of SMRs. The team began with the 12 AFSPC installations in the Continental United States and Alaska and applied criteria derived from research and stakeholder interactions to determine the two optimum installations for the use case studies. The AFSPC installations vary from large Air Force Bases (AFB), with multiple missions, fully mature infrastructures, and robust services to small Air Force Stations (AFS) with single missions and limited infrastructures, often in remote locations. The selection criteria evolved into two categories: AFSPC-related criteria (mission priorities, synergistic support capabilities, and installation operations) and siting criteria (available land, seismology, hydrology, population density, proximity to hazardous activities/protected lands). For mission priorities, the team used prioritized space superiority activities approved by the Commander, AFSPC, and inputs from various subject matter experts (SME). The team determined synergistic support capabilities--installation-SMR owner/operator collaboration on key activities such as security, fire protection, and emergency response--from existing host unit documents, SME inputs, and team member familiarity with AFSPC installation capabilities. Installation operations factors were obtained through inputs from applicable Air Force agencies and SMEs. Based on Department of Energy guidance, the team developed values for a Site Selection Evaluation Criteria and submitted them to the Oak Ridge National Laboratory to apply their Oak Ridge Siting Analysis for power Generation Expansion tool. Since data for Clear AFS, AK are not currently included in the siting tool, the team used similar United States Geologic Survey data. Finally, since the study's focus is SMR feasibility versus actual siting, the team considered use case-unique considerations rather than using AFSPC and siting criteria as the only determinants in selecting the use case installations. As a result of the selection process, the team concluded that Schriever AFB, CO and Clear AFS, AK best lend themselves to the more detailed use case feasibility analysis. This conclusion considers the higher priority missions performed on both installations and their generally favorable siting characteristics, but also enables a robust use case comparison of two installations that represent the spectrum of AFSPC installation characteristics: a fully-mature, multi-mission AFB and a more limited capability, single mission, remotely-located AFS. The next phase of the SMR Suitability study will involve performing use case studies of Schriever AFB and Clear AFS. This includes site visits; in-depth interaction with installation SMEs; expanding interaction with DoD, Air Force, AFSPC, and private sector entities; refining commercial business models; and consideration of micro-grids and other technologies that may enhance SMR deployment on DoD installations
In the past few decades, a number of user-level threading and tasking models have been proposed in the literature to address the shortcomings of OS-level threads, primarily with respect to cost and flexibility. Current state-of-the-art user-level threading and tasking models, however, either are too specific to applications or architectures or are not as powerful or flexible. In this paper, we present Argobots, a lightweight, low-level threading and tasking framework that is designed as a portable and performant substrate for high-level programming models or runtime systems. Argobots offers a carefully designed execution model that balances generality of functionality with providing a rich set of controls to allow specialization by end users or high-level programming models. We describe the design, implementation, and performance characterization of Argobots and present integrations with three high-level models: OpenMP, MPI, and colocated I/O services. Evaluations show that (1) Argobots, while providing richer capabilities, is competitive with existing simpler generic threading runtimes; (2) our OpenMP runtime offers more efficient interoperability capabilities than production OpenMP runtimes do; (3) when MPI interoperates with Argobots instead of Pthreads, it enjoys reduced synchronization costs and better latency-hiding capabilities; and (4) I/O services with Argobots reduce interference with colocated applications while achieving performance competitive with that of a Pthreads approach.
Esteban Meneses合作论文数University of Illinois at Urbana-Champaign;Department of Computer Science2