Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making. However, ABMs have traditionally relied on static prior, which prevents the models from adapting to real-time changes. Our research provides a novel approach to addressing this information gap. Large language models (LLMs) offer new opportunities to predict human decision-making. Here, we introduce a scalable Hybrid Agent-based and Language-driven Epidemic (HALE) modeling framework that leverages LLMs to predict human decision-making in an ABM simulation. As a proof-of-concept, we use HALE to simulate COVID-19 and its effects in Salt Lake County, UT.
Simulating large-scale, high-fidelity population health models demands immense computational resources and efficient parallelization strategies. We present the design and performance of Enable (Efficient National-scale Agent-Based Learning Environment), a hybrid CPU-GPU agent-based modeling framework optimized for the Frontier supercomputer. Enable generates contact networks directly from activity schedules, enabling location-based parallel execution without relying on precomputed contact graphs. A GPU device constructs contact networks and uses an efficient load balancing algorithm to assign locations to processors, while CPUs perform the simulation tasks. This design achieves high parallel efficiency by ensuring balanced edge distributions across processors, resulting in uniform execution times. We evaluate both strong and weak scaling using synthetic and real-world datasets generated from UrbanPop, Uber H3, and OSM maps. Scaling performance is studied with city-to national-scale population sizes on up to 1200 GPUs of the Frontier supercomputer. Enable addresses key computational challenges in large-scale high-fidelity agent-based simulations that beset the development of national-scale virtual population health twins.
Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver similar to 50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.
Digital Twin (DT) represents an essential technology in which an operations model of a physical system uses real-time data to predict, monitor, and improve the physical system’s operations. One of the primary objectives of a DT is to inform the physical system of measures to take in response to one or multiple intervening events that change the physical system’s state. The capability to perform various real-time scenario assessments in readiness for such events is an effective use of simulations as DTs, and here, scalable performance-efficient simulation cloning methods become relevant. However, continuous evaluations of simulation clones, each representing a unique cascade of intervening events, are highly challenging due to the constraints of finite memory and an extensive exploration space. This paper reports a novel simulation cloning-based method to continuously evaluate k-tree probabilistic what-if scenarios under finite resource constraints to realize a DT for the power grid.
Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with about 75 while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers about 20
As part of a larger effort, this work-in-progress reports the possible advantages of modifying conventional workflows used to generate labelled training samples and train machine learning (ML) models on them. We compare results from three different workflows using neutron scattering data analysis as the motivating application and report about 20% improvement in speedup, with no appreciable loss of model accuracy, over a baseline workflow.
Phase field (PF) simulations are computationally expensive but remain a key analysis tool to understand the complex mechanisms of additive manufacturing (AM) processes. Each PF simulation-aided analysis requires thousands of node hours on leadership-class supercomputers. One of the main goals of these analyses is the study of microstructure evolution during the build process which begins with the onset of nucleation. Nucleation occurs under certain thermomechanical conditions which are not known a priori and many PF simulations are required to identify ranges of input thermo-mechanical parameters that can result in the onset of nucleation. Since many of the simulations do not result in nucleation, an analysis campaign often ends up wasting tremendous amounts of precious computing resources executing nucleation-absent simulations. The goal of this work is to design and train deep learning models to inform a PF simulation about the likelihood of the occurrence of nucleation in a future simulation time-step based on the state summary over a finite number of past time-steps of a running simulation. If the prediction determines that the running simulation is unlikely to reach nucleation in the allotted time, then its execution is stopped immediately ultimately resulting in vast reduction in wasted computations when accrued over all the PF simulations typically performed in a single or multiple analysis campaign(s). The paper presents the performance of a machine learning pipeline that uses a convolutional neural network (CNN) model to learn an embedding which is then used with a self-attention network to build a multi-task deep learning model to predict the likelihood of nucleation. The model also predicts the input parameters used in a simulation. Performance is compared with a baseline pipeline that uses an off-the-shelf LeNet-5 model to learn the initial embedding. Despite their smaller size, performance results indicate significant improvement in accuracy of the proposed models compared to the larger baseline models.
Neutron scattering is a state-of-the-art experimental technique that allows scientists to probe material structures with atomic resolutions by scattering beams of neutrons from them. Currently, it is common in the literature to solve these inverse problems using loop refinement techniques. The overall time-to-solution and the quality of results from loop refinement methods also depend on the fidelity of the forward model used to generate the Bragg profiles within the loop iterations. Machine Learning-driven methods for structure determination from neutron scattering data is an emerging area of research. An auto-encoder consist of two parts: an encoder and a decoder. Experimental neutron powder diffraction data of barium titanate as a function of temperature was collected on the NOMAD instrument housed in the Spallation Neutron Source at Oak Ridge National Laboratory. Background signals in neutron detectors originate from a variety of sources and need to be subtracted out to improve the signal-to-noise ratio.
Rapid growth in data, computational methods, and computing power is driving a remarkable revolution in what variously is termed machine learning (ML), statistical learning, computational learning, and artificial intelligence. In addition to highly visible successes in machine-based natural language translation, playing the game Go, and self-driving cars, these new technologies also have profound implications for computational and experimental science and engineering, as well as for the exascale computing systems that the Department of Energy (DOE) is developing to support those disciplines. Not only do these learning technologies open up exciting opportunities for scientific discovery on exascale systems, they also appear poised to have important implications for the design and use of exascale computers themselves, including high-performance computing (HPC) for ML and ML for HPC. The overarching goal of the ExaLearn co-design project is to provide exascale ML software for use by Exascale Computing Project (ECP) applications, other ECP co-design centers, and DOE experimental facilities and leadership class computing facilities.
We present in this work a consistent numerical scheme that allows the computation of 3D magnetic fields a nd 3D density profiles and their usage in ion cyclotron range of frequencies (ICRF) coupling simulations. We first utilize the PARVMEC code to compute the 3D free-boundary plasma equilibrium in the ideal magnetohydrodynamic (MHD) approximation. Since the PARVMEC solution is only defined within the last closed flux surface (LCFS), the magnetic field domain is extended to the scrape-off layer (SOL) via the BMW code, which computes a divergence-free magnetic field solution arising from the external conductors' vacuum field a nd t he PARVMEC flux surface currents. This magnetic reconstruction is then used in the EMC3-EIRENE transport code in order to compute 3D density profiles. In the last step, the RAPLICASOL code is utilized to compute the ICRF antenna S-matrices resulting from the 3D density profiles. We exemplify this scheme for the ASDEX Upgrade tokamak. A new implementation of a curved model for the ASDEX Upgrade ICRF 2-strap antenna in RAPLICASOL allows simulations in realistic geometry, without any coordinate transformations.
One of the main goals of neutron data analysis is to determine the internal structure of materials from their neutron scattering profiles. These structures are defined by a crystallographic class label and a set of real-valued parameters specific to that class. Existing structure analysis approaches use computationally expensive loop refinements methods that routinely take days, and even weeks, to complete. Additionally, the outcomes often rely on the fidelity of physical models that are computed during the refinement process. Here, we evaluate the feasibffity of using trained data-driven machine learning models as fast and accurate substitutes for these expensive methods. We report on the efficacies of a variety of ML models, including convolutional neural networks, auto-encoders, random forests and combinations thereof, in addition to techniques such as transfer learning in predicting these structural parameters. Specifically, we evaluate two categories of models which we call class-conditional and integrated. The first relies on a two-stage inference pipeline in which a crystallographic class label is first predicted followed by regression to predict the length/angle parameters. In the second category, the classification and regression tasks are performed as a single learning task. We train these models on synthetically generated data, validate them against experimental observa-tions and show that integrated models outperform their class-conditional counterparts opening up the possibffity of deep learning models as a viable alternative to existing resource-intensive loop refinement methods in neutron data analysis.
Semantic segmentation of images is an important computer vision task that emerges in a variety of application domains such as medical imaging, robotic vision and autonomous vehicles to name a few. While these domain-specific image analysis tasks involve relatively small image sizes (∼ 102 × 102), there are many applications that need to train machine learning models on image data with extents that are orders of magnitude larger (∼ 104 × 104). Training deep neural network (DNN) models on large extent images is extremely memory-intensive and often exceeds the memory limitations of a single graphical processing unit, a hardware accelerator of choice for computer vision workloads. Here, an efficient, sample parallel approach to train U-Net models on large extent image data sets is presented. Its advantages and limitations are analyzed and near-linear strong-scaling speedup demonstrated on 256 nodes (1536 GPUs) of the Summit supercomputer. Using a single node of the Summit supercomputer, an early evaluation of a recently released model parallel framework called GPipe is demonstrated to deliver ∼ 2X speedup in executing a U-Net model with an order of magnitude larger number of trainable parameters than reported before. Performance bottlenecks for pipelined training of U-Net models are identified and mitigation strategies to improve the speedups are discussed. Together, these results open up the possibility of combining both approaches into a unified scalable pipelined and data parallel algorithm to efficiently train U-Net models with very large receptive fields on data sets of ultra-large extent images.
Understanding structural properties of materials and how they relate to its atomic structure, while extremely challenging, is a key scientific quest that has dominated the landscape of materials research for decades. Neutron and X-ray scattering is a state-of-the-art method to investigate material structure on the atomic scale. Traditional methods of processing neutron scattering data to decipher the structure of target materials have relied on computing scattering patterns using physics-based forward models and comparing them with experimentally gathered scattering profiles within a computationally expensive optimization loop. Here, we report an initial design of a data-driven machine learning pipeline for material structure prediction that is computationally faster (once trained) and potentially more accurate. We describe the architecture of the ML pipeline and a preliminary benchmarking study of shallow machine learning models in terms of their prediction accuracy and limitations. We show that material structure prediction from neutron scattering data using shallow learning models is feasible to within 90% prediction accuracy for certain classes of materials but deeper models are required for more general material structure predictions.
Large, spontaneous m/n = 1/1 helical cores are predicted in tokamaks with extended regions of low- or reversed-magnetic shear profiles in a region within the q = 1 surface and an onset condition determined by constant (dp/dρ)/Bt2 along the threshold. These 3D modes occurred frequently in Alcator C-Mod during ramp-up when slow current penetration results in a reversed shear q-profile. The onset and early development of a helical core in C-Mod were simulated using a new 3D time-dependent equilibrium reconstruction, based on the ideal MHD equilibrium code VMEC. The reconstruction used the experimental density, temperature, and soft-X-ray fluctuations. The pressure profile can become hollow due to an inverted, hollow electron temperature profile caused by molybdenum radiation in the plasma core during the current ramp-up phase before the onset of sawteeth, which may also occur in ITER with tungsten. Based on modeling, it is found that a reverse shear q-profile combined with a hollow pressure profile reduces the onset condition threshold, enabling helical core formation from an otherwise axisymmetric equilibrium.
Atom probe tomography (APT) is a material probing technique that has undergone dramatic improvements in its capability to map individual atoms within a material sample resulting in data files with hundreds of millions of atoms. Understanding the nano-structural features hidden in these massive amounts of atomic data is a crucial analysis task for materials scientists. However, fast analysis capabilities for large APT workloads remains a critical bottleneck. In this paper, we present the design, implementation and detailed performance evaluations of a parallel software capable of efficiently performing extremely time-consuming correlation analyses of massive high density APT data. Starting with shared memory implementations to motivate our design choices, we extend the implementation to hybrid architectures keeping realistic APT workloads in mind. Detailed performance analyses of three different parallel implementations of the software are supported by empirical results on a Cray XC30 and a Cray XC40 architecture. Its usefulness is demonstrated by reducing the turnaround time of an end-to-end APT correlation analysis on 100 millions atoms by three orders of magnitude using 2048 MPI ranks on 1024 nodes (24 cores per node) of a Cray XC30. The software reported here equips material scientists for the first time with a high-speed scalable capability for efficient and timely analyses of massive APT data.
Large, spontaneous m/n = 1/1 helical cores are shown to be expected in tokamaks such as ITER with extended regions of low-or reversed-magnetic shear profiles and q near 1 in the core. The threshold for this spontaneous symmetry breaking is determined using VMEC scans, beginning with reconstructed 3D equilibria from DIII-D and Alcator C-Mod based on observed internal 3D deformations. The helical core is a saturated internal kink mode (Wesson 1986 Plasma Phys. Control. Fusion 28 243); its onset threshold is shown to be proportional to (dp/ d rho)/ B-t(2) around q = 1. Below the threshold, applied 3D fields can drive a helical core to finite size, as in DIII-D. The helical core size thereby depends on the magnitude of the applied perturbation. Above it, a small, random 3D kick causes a bifurcation from axisymmetry and excites a spontaneous helical core, which is independent of the kick size. Systematic scans of the q-profile show that the onset threshold is very sensitive to the q-shear in the core. Helical cores occur frequently in Alcator C-Mod during ramp-up when slow current penetration results in a reversed shear q-profile, which is favorable for helical core formation. Finally, a comparison of the helical core onset threshold for discharges from DIII-D, Alcator C-Mod and ITER confirms that while DIII-D is marginally stable, Alcator C-Mod and especially ITER are highly susceptible to helical core formation without being driven by an externally applied 3D magnetic field.
Powered by recent advances in data acquisition technologies, today's state-of-the-art atom probe microscopes yield data sets with sizes ranging from a few million atoms to billions of atoms. Analysis of these atomic data sets within rea-sonable turnaround times is a pressing data analysis challenge for material scientists currently equipped with software systems that do not scale to these massive data sets. Here, we present the shared memory component of a larger ongoing effort to develop a multi-feature data analysis framework capable of analyzing atom probe data of all sizes and scales from desktop multicore machines to large-scale high-performance computing platforms with hybrid (shared and distributed memory) architectures. Our focus here is on a broad class of popular atom probe data analysis methods that rely on core time-consuming k-NN queries. We present a scalable, heuristic algorithm for k-NN queries using three-dimensional range trees. To demonstrate its efficacy, the k-NN algorithm is integrated with two use cases of atom probe data analysis methods and the resulting analysis times are shown to speedup by over 20X on a 32-core Cray XC40 node using workloads up to 8 million atoms, which is already beyond the at-scale capabilities of existing atom probe software. Using this k-NN algorithm, we also introduce a novel parameter estimation method for a class of cluster finding methods, called friends-of-friends (FoF) methods, to completely bypass their expensive pre-processing steps. In each case, we validate the results on a variety of control data sets.