Spatial Dataflow Architectures are an emerging hardware pattern in high-performance computing, whose mesh-connected fixed-memory processing elements are tailored for structured grid kernels with two-dimensional neighborhoods. However, practical multiphysics codes are often computed on unstructured grids, which induce indirect memory accesses and high-dimensional communication patterns, making them infeasible to directly map onto said architectures. This work takes a principled, model-centric approach to partitioning unstructured problems onto spatial dataflow architectures. Through communication and memory modeling, we propose a joint decomposition that considers both the size of the application's fields and its subroutines. In particular, we automate the analysis process of the original code, define a high-dimensional decomposition that minimizes communication via space-filling curves, and apply memory optimization techniques, crucial in this memory-limited environment. We demonstrate mapping the Livermore Unstructured Lagrangian Explicit Shock Hydrodynamics (LULESH) application to the Cerebras Wafer-Scale Engine, showing that larger, unstructured grid codes can still outperform GPUs.
KRAS4a and KRAS4b are important regulators of signaling, and their interactions with the plasma membrane are dynamic and influenced by lipid composition. KRAS 4a and 4b have nearly identical globular domains but differ in their membrane-associated hyper variable region (HVR). The functional distinctions between these isoforms remain unclear, particularly with regards to their dependence on specific lipids and the membrane environment. Previous work showed that the membrane orientation of KRAS4b affects its ability to bind to RAF kinase RBDCRD and that the KRAS-RBDCRD complex adopts different poses on the membrane as well as influences the size and composition of the lipid environment. To model differences between KRAS 4a and 4b protein-lipid interactions, we extended the Multiscale Machine-Learned Modeling Infrastructure (MuMMI) to incorporate continuum simulations in the grand canonical ensemble, enabling sampling across macroscopic, coarse-grained, and all-atom resolutions. Using this framework, we systematically altered PIP2 concentrations, KRAS 4a versus 4b, and RAF RBDCRD complexation to assess impacts on membrane-protein interactions and dynamics. Our results reveal that reducing PIP2 shifts and broadens the membrane orientational preference of both KRAS 4b and 4a, with stronger effects on 4b HVR localization versus 4a. We demonstrate that with depletion of the strong negatively charged PIP2 lipid, the less charged phosphatidylserine replaces PIP2. Our findings highlight similarities and distinctions in the dynamics and lipid dependency of KRAS isoforms and suggest that ordering of the local lipid composition by HVRs is a shared property and key modulator of RAS-mediated signaling at the plasma membrane.
To gain molecular and mechanistic insights into initiation of the RAS-RAF signaling cascade, we developed and used a combination of multiscale simulation and experimental approaches. The influence and impact of the membrane on RAS and RAF proteins is a factor we are just beginning to understand and appreciate in more detail. Molecular simulation is an ideal methodology to further study this complicated relationship between the membrane and associated proteins. Our previous work using Multiscale Machine-learned Modeling Infrastructure investigated different lipid compositions solely around the KRAS4b protein and the interplay between protein behavior and these membrane environments. Multiscale Machine-learned Modeling Infrastructure uses machine learning to couple adjacent simulation scales and has been efficiently scaled across some of the world’s largest high-performance computers. Recently, we have expanded this multiresolution framework to include the all-atom simulation scale and to incorporate the RAF RBDCRD domains. Here, we present the overall analysis results from this new simulation campaign comprising a mixture of RAS and RAF RBDCRD proteins. Approximately 35,000 coarse-grained and 10,000 all-atom molecular dynamics simulations were completed, sampled from a variety of protein/lipid composition configurations that were generated from a micron-scale continuum simulation containing hundreds of copies of the proteins.Our studies suggest that orientations of the RAS-RBDCRD complex on the membrane occupy distinct configurational states, and the spatial patterns of lipid arrangements around these different protein states are unique to each state. The extent and size of lipid “fingerprints” imposed on the membrane by the RAS-RBDCRD protein complex are significantly larger than observed for just the RAS protein on its own. These protein complexes strongly associate, but we do not observe statistically significant preferred protein-protein orientations. These observations indicate that spatial colocalization of RAS-RBDCRD proteins in the same vicinity may be assisted by specific membrane environments, acting to increase the probability of signaling complex formation.
Advances in deep learning and generative modeling have driven interest in data-driven molecule discovery pipelines, whereby machine learning (ML) models are used to filter and design novel molecules without requiring prohibitively expensive first-principles simulations. Although the discovery of novel molecules that extend the boundaries of known chemistry requires accurate out-of-distribution (OOD) predictions, ML models often struggle to generalize OOD. Furthermore, there are currently no systematic benchmarks for molecular OOD prediction tasks. We present BOOM, benchmarks for out-of-distribution molecular property predictions – a benchmark study of property-based out-of-distribution models for common molecular property prediction models. We evaluate more than 140 combinations of models and property prediction tasks to benchmark deep learning models on their OOD performance. Overall, we do not find any existing models that achieve strong OOD generalization across all tasks: even the top performing model exhibited an average OOD error 3x larger than in-distribution. We find that deep learning models with high inductive bias can perform well on OOD tasks with simple, specific properties. Although chemical foundation models with transfer and in-context learning offer a promising solution for limited training data scenarios, we find that current foundation models do not show strong OOD extrapolation capabilities. We perform extensive ablation experiments to highlight how OOD performance is impacted by data generation, pre-training, hyperparameter optimization, model architecture, and molecular representation. We propose that developing ML models with strong OOD generalization is a new frontier challenge in chemical ML model development. This open-source benchmark will be made available on Github.
We present progress in utilizing a machine learning (ML) assisted optimization framework to study the trends in a parameter space defined by spectrally shaped, high-intensity, petawatt-class (8 J, 45 fs) laser pulses interacting with solid targets and give the first simulation-based overview of predicted trends. A neural network (NN) incorporating uncertainty quantification is trained to predict the number of hot electrons generated by the laser–target interaction as a function of pulse shaping parameters. The predictions of this NN serve as the basis function for a Bayesian optimization framework to navigate this space. For post-experimental evaluation, we compare two separate neural network (NN) models. One is based solely on data from experiments, and the other is trained only on ensemble particle-in-cell simulations. Reviewing the predicted and observed trends across the experiment-capable laser parameter search space, we find that both ML models predict a maximal increase in hot electron generation at a level of approximately 12%–18%; however, no statistically significant enhancement was observed in experiments. On direct comparison of the NN models, the average discrepancy is 8.5%, with a maximum of 30%. Since shot-to-shot fluctuations in experiments affect the observations, we evaluate the behavior of our optimization framework by performing virtual experiments that vary the number of repeated observations and the noise levels. Here, we discuss the implications of such a framework for future autonomous exploration platforms in high-repetition-rate experiments.
Communication overhead is a key challenge in distributed deep learning, especially on slower Ethernet interconnects, and given current hardware trends, communication is likely to become a major bottleneck. While gradient compression techniques have been explored for SGD and Adam, the Lion optimizer has the distinct advantage that its update vectors are the output of a sign operation, enabling straightforward quantization. However, simply compressing updates for communication and using techniques like majority voting fails to lead to end-to-end speedups due to inefficient communication algorithms and reduced convergence. We analyze three factors critical to distributed learning with Lion: optimizing communication methods, identifying effective quantization methods, and assessing the necessity of momentum synchronization. Our findings show that quantization techniques adapted to Lion and selective momentum synchronization can significantly reduce communication costs while maintaining convergence. We combine these into Lion Cub, which enables up to 5x speedups in end-to-end training compared to Lion. This highlights Lion's potential as a communication-efficient solution for distributed training.
Data-driven science and technology offer transformative tools and methods to science. This review article highlights the latest development and progress in the interdisciplinary field of data-driven plasma science (DDPS), i.e., plasma science whose progress is driven strongly by data and data analyses. Plasma is considered to be the most ubiquitous form of observable matter in the universe. Data associated with plasmas can, therefore, cover extremely large spatial and temporal scales, and often provide essential information for other scientific disciplines. Thanks to the latest technological developments, plasma experiments, observations, and computation now produce a large amount of data that can no longer be analyzed or interpreted manually. This trend now necessitates a highly sophisticated use of high-performance computers for data analyses, making artificial intelligence and machine learning vital components of DDPS. This article contains seven primary sections, in addition to the introduction and summary. Following an overview of fundamental data-driven science, five other sections cover widely studied topics of plasma science and technologies, i.e., basic plasma physics and laboratory experiments, magnetic confinement fusion, inertial confinement fusion and high-energy-density physics, space and astronomical plasmas, and plasma technologies for industrial and other applications. The final Section before the summary discusses plasma-related databases that could significantly contribute to DDPS. Each primary Section starts with a brief introduction to the topic, discusses the state-of-the-art developments in the use of data and/or data-scientific approaches, and presents the summary and outlook. Despite the recent impressive signs of progress, the DDPS is still in its infancy. This article attempts to offer a broad perspective on the development of this field and identify where further innovations are required.
nuclear plants currently under construction including 10 in China, 8 in India, and 4 in Russia. In the United States, there have been notifications to the Nuclear Regulatory Commission of intentions to apply for combined construction and operating licenses for 27 new units over the next decade. The projected growth in nuclear power has focused increasing attention on issues related to the permanent disposal of nuclear waste, the proliferation of nuclear weapons technologies and materials, and the sustainability of a once-through nuclear fuel cycle. In addition, the effective utilization of nuclear power will require continued improvements in nuclear technology, particularly related to safety and efficiency. In all of these areas, the performance of materials and chemical processes under extreme conditions is a limiting factor. The related basic research challenges represent some of the most demanding tests of our fundamental understanding of materials science and chemistry, and they provide significant opportunities for advancing basic science with broad impacts for nuclear reactor materials, fuels, waste forms, and separations techniques. Of particular importance is the role that new nanoscale characterization and computational tools can play in addressing these challenges. These tools, which include DOE synchrotron X-ray sources, neutron sources, nanoscale science research centers, and supercomputers, offer the opportunity to transform and accelerate the fundamental materials and chemical sciences that underpin technology development for advanced nuclear energy systems. The fundamental challenge is to understand and control chemical and physical phenomena in multi-component systems from femto-seconds to millennia, at temperatures to 1000?C, and for radiation doses to hundreds of displacements per atom (dpa). This is a scientific challenge of enormous proportions, with broad implications in the materials science and chemistry of complex systems. New understanding is required for microstructural evolution and phase stability under relevant chemical and physical conditions, chemistry and structural evolution at interfaces, chemical behavior of actinide and fission-product solutions, and nuclear and thermomechanical phenomena in fuels and waste forms. First-principles approaches are needed to describe f-electron systems, design molecules for separations, and explain materials failure mechanisms. Nanoscale synthesis and characterization methods are needed to understand and design materials and interfaces with radiation, temperature, and corrosion resistance. Dynamical measurements are required to understand fundamental physical and chemical phenomena. New multiscale approaches are needed to integrate this knowledge into accurate models of relevant phenomena and complex systems across multiple length and time scales.
Interdependence across time and length scales is common in biology, where atomic interactions can impact larger-scale phenomenon. Such dependence is especially true for a well-known cancer signaling pathway, where the membrane-bound RAS protein binds an effector protein called RAF. To capture the driving forces that bring RAS and RAF (represented as two domains, RBD and CRD) together on the plasma membrane, simulations with the ability to calculate atomic detail while having long time and large length- scales are needed. The Multiscale Machine-Learned Modeling Infrastructure (MuMMI) is able to resolve RAS/RAF protein-membrane interactions that identify specific lipid-protein fingerprints that enhance protein orientations viable for effector binding. MuMMI is a fully automated, ensemble-based multiscale approach connecting three resolution scales: (1) the coarsest scale is a continuum model able to simulate milliseconds of time for a 1 μm2 membrane, (2) the middle scale is a coarse-grained (CG) Martini bead model to explore protein-lipid interactions, and (3) the finest scale is an all-atom (AA) model capturing specific interactions between lipids and proteins. MuMMI dynamically couples adjacent scales in a pairwise manner using machine learning (ML). The dynamic coupling allows for better sampling of the refined scale from the adjacent coarse scale (forward) and on-the-fly feedback to improve the fidelity of the coarser scale from the adjacent refined scale (backward). MuMMI operates efficiently at any scale, from a few compute nodes to the largest supercomputers in the world, and is generalizable to simulate different systems. As computing resources continue to increase and multiscale methods continue to advance, fully automated multiscale simulations (like MuMMI) will be commonly used to address complex science questions.
The PROBIES diagnostic is a new, highly flexible, imaging and energy spectrometer designed for laser-accelerated protons. The diagnostic can detect low-mode spatial variations in the proton beam profile while resolving multiple energies on a single detector or more. When a radiochromic film stack is employed for “single-shot mode,” the energy resolution of the stack can be greatly increased while reducing the need for large numbers of films; for example, a recently deployed version allowed for 180 unique energy measurements spanning ∼3 to 75 MeV with <0.4 MeV resolution using just 20 films vs 180 for a comparable traditional film and filter stack. When utilized with a scintillator, the diagnostic can be run in high-rep-rate (>Hz rate) mode to recover nine proton energy bins. We also demonstrate a deep learning-based method to analyze data from synthetic PROBIES images with greater than 95% accuracy on sub-millisecond timescales and retrained with experimental data to analyze real-world images on sub-millisecond time-scales with comparable accuracy.
Graph neural networks (GNNs) are a powerful approach for machine learning on graph datasets. Such datasets often consist of millions of modestly-sized graphs, making them well-suited for data-parallel training. However, existing methods show poor scaling due to load imbalances and kernel overheads. We propose an optimized 2D scatter-gather based represen-tation of GNNs that is amenable to distributed, data-parallel training without changing the underlying mathematics of the GNN. By padding graph data to a fixed size on each process, we can simplify data ingestion, make use of efficient compute kernels, equally distribute computation load, and reduce overheads. We benchmark edge-conditioned GNNs with the PCQM4M-LSC and OGB-PPA datasets. Our implementation shows better runtime performance than the state-of-the-art, with a $12\times$ strong-scaling speedup on 16 GPUs and an $89.4\times\ \text{weak}$ -scaling speedup on 100 GPUs.
With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. In this paper, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. In addition to its design, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.
RAS is a signaling protein associated with the cell membrane that is mutated in up to 30% of human cancers. RAS signaling has been proposed to be regulated by dynamic heterogeneity of the cell membrane. Investigating such a mechanism requires nearatomistic detail at macroscopic temporal and spatial scales, which is not possible with conventional computational or experimental techniques. We demonstrate here a multiscale simulation infrastructure that uses machine learning to create a scale-bridging ensemble of over 100,000 simulations of active wild-type KRAS on a complex, asymmetric membrane. Initialized and validated with experimental data (including a new structure of active wild-type KRAS), these simulations represent a substantial advance in the ability to characterize RAS-membrane biology. We report distinctive patterns of local lipid composition that correlate with interfacially promiscuous RAS multimerization. These lipid fingerprints are coupled to RAS dynamics, predicted to influence effector binding, and therefore may be a mechanism for regulating cell signaling cascades.
Composite science workflows are gaining traction to manage the combined effects of (1) extreme hardware heterogeneity in new High Performance Computing (HPC) systems and (2) growing software complexity – effects necessitated by the convergence of traditional HPC with data sciences. Composing, analyzing, and optimizing a composite workflow remains highly challenging as the component technologies are generally developed in isolation and often feature widely varying levels of performance, scalability, and interoperability. In this paper, we propose novel workflow composition and analysis techniques to create and optimize a scalable and effective composite workflow for heterogeneous HPC centers, and define the performance space of variables that impact composite workflow performance. We present PerfFlowAspect, an Aspect Oriented Programming (AOP)-based tool to perform cross-cutting performance analysis of composite workflows and better understand the impact of key performance variables on workflows. Our solution directly addresses AOP concerns that can affect workflow performance and covers the full software lifecycle, ranging from the workflow's initial composition through performance analysis and optimization. We use our science workflow composition techniques to implement the American Heart Association Molecule Screening (AHA MoleS) workflow. Through experimentation, we demonstrate that tuning a single performance variable can improve AHA MoleS workflow performance by a factor of up to 2.45x. Our evaluation suggests that our techniques can significantly enhance the ability of a multi-disciplinary research and development team to create a high performance composite workflow.
Scalable management of user workloads on large-scale supercomputers remains a challenge due to the tradeoff between capturing adequate detail for analysis from various data sources and minimizing overhead. Co-designed frameworks, such as IBM’s Cluster System Management (CSM), provide a unified approach and novel insights for large-scale cluster management. This paper presents a longitudinal study and detailed analysis of a first-of-its-kind dataset collected by CSM from one of the world’s fastest supercomputers – comprised of over 1.4 million jobs on heterogenous nodes over multiple years. Furthermore, by focusing on a case study for power management, we identify the strengths and limitations of current CSM power measurement techniques in production. We present a deep dive into a large-scale scientific workflow, where finer-grained monitoring reveals power fluctuations at megawatt-levels resulting from the dynamic nature of the application, which are not captured by CSM. We make our unique datasets available to the HPC community, and discuss potential mitigation strategies by analyzing both coarse-grained and fine-grained data.
Transformer models have revolutionized the field of Natural Language Processing (NLP) and they achieve state-of-the-art performance in applications like machine translation, question answering, regression, and summarization. However, training Transformers is challenging because of their large memory and compute requirements. The literature contains several approaches to parallelize training, like layer parallelism and pipeline parallelism, but they are optimized to benefit out-of-core models and they don’t exploit the inherent parallelism in Transformer models. Other work uses model parallelism to achieve weak scaling by increasing the model size. In this paper, we propose sub-graph parallelism that provides a significant performance improvement over pure data parallelism with a fixed number of resources, and as an additional technique for strong- and weak-scaling without increasing model capacity. Our technique accelerates the training of Transformer models and we generalize the concept to any neural network with multiple branches. We optimize the communication for sub-graph parallelism and combine it with data parallelism to scale performance up to 1024 GPUs. To decrease communication overheads, we propose a topology-aware scheme that limits inter-node communication. Finally, we empirically compare sub-graph parallelism with pure data parallelism and demonstrate its performance benefits in end-to-end training.
Multiscale simulations are a well-accepted way to bridge the length and time scales required for scientific studies with the solution accuracy achievable through available computational resources. Traditional approaches either solve a coarse model with selective refinement or coerce a detailed model into faster sampling, both of which have limitations. Here, we present a paradigm of adaptive, multiscale simulations that couple different scales using a dynamic-importance sampling approach. Our method uses machine learning to dynamically and exhaustively sample the phase space explored by a macro model using microscale simulations and enables an automatic feedback from the micro to the macro scale, leading to a self-healing multiscale simulation. As a result, our approach delivers macro length and time scales, but with the effective precision of the micro scale. Our approach is arbitrarily scalable as well as transferable to many different types of simulations. Our method made possible a multiscale scientific campaign of unprecedented scale to understand the interactions of RAS proteins with a plasma membrane in the context of cancer research running over several days on Sierra, which is currently the second-most-powerful supercomputer in the world. Tackling scientific problems often requires computational models that bridge several spatial and temporal scales. A new simulation framework employing machine learning, which is scalable and can be used on standard laptops as well as supercomputers, promises exhaustive multiscale explorations.