In this work, we develop two main Machine Learning based approaches to predict the runtime parameters of highly scalable parallel chemistry computations.These approaches employ active and generative learning together with the empirically determined gradient boosted regression tree models chosen among a rich suite of machine learning models. When evaluated on Coupled-Cluster with Singles and Doubles computations, our models achieve a mean absolute error percentage (MAPE) as low as 0.023 and a coefficient of determination as high as 99.9%. Furthermore, when combined with active learning to mitigate the lack of large amounts of training data, our models score a MAPE about 0.2 with 20-25% of the original dataset.
In this work, we develop machine learning (ML) based strategies to predict resources (costs) required for massively parallel chemistry computations to guide application users before they commit to running expensive experiments on a supercomputer. By predicting application execution time, we determine the optimal runtime parameter values such as number of nodes and tile sizes. Two key questions of interest to users are addressed. The first is the shortest-time question, where the user is interested in knowing the parameter values to achieve the shortest execution time for a given problem size and a target supercomputer. The second is the cheapest-run question in which the user is interested in minimizing resource costs, e.g., finding the number of nodes and tile size that minimizes the number of node-hours for a given problem size. We evaluate a rich family of ML models using the collections of runtime parameter values for the Coupled Cluster with Singles and Doubles (CCSD) application run on the Frontier and Aurora supercomputers. Our experiments show that when predicting the total execution time of CCSD iteration, a Gradient Boosting (GB) ML model achieves a Mean Absolute Percentage Error (MAPE) of 0.023 and 0.073 for Aurora and Frontier, respectively. In the case where it is expensive to run experiments to collect data points, Active Learning (AL) achieves a MAPE of about 0.2 with just around 450 experiments collected from Aurora and Frontier.
The transformative impact of modern computational paradigms and technologies, such as high-performance computing (HPC), quantum computing, and cloud computing, has opened up profound new opportunities for scientific simulations. Scalable computational chemistry is one beneficiary of this technological progress. The main focus of this paper is on the performance of various quantum chemical formulations, ranging from low-order methods to high-accuracy approaches, implemented in different computational chemistry packages and libraries, such as NWChem, NWChemEx, Scalable Predictive Methods for Excitations and Correlated Phenomena, ExaChem, and Fermi-Löwdin orbital self-interaction correction on Azure Quantum Elements, Microsoft's cloud services platform for scientific discovery. We pay particular attention to the intricate workflows for performing complex chemistry simulations, associated data curation, and mechanisms for accuracy assessment, which is demonstrated with the Arrows automated workflow for high throughput simulations. Finally, we provide a perspective on the role of cloud computing in supporting the mission of leadership computational facilities.
Polariton chemistry has attracted great attention as a potential route to modify chemical structure, properties, and reactivity through strong interactions between molecular electronic, vibrational, or rovibrational degrees of freedom. A rigorous theoretical treatment of molecular polaritons requires the treatment of matter and photon degrees of freedom on equal quantum mechanical footing. In the limit of molecular electronic strong or ultrastrong coupling to one or a few molecules, it is desirable to treat the molecular electronic degrees of freedom using the tools of ab initio quantum chemistry, yielding an approach we refer to as ab initio cavity quantum electrodynamics, where the photon degrees of freedom are treated at the level of cavity quantum electrodynamics. Here, we present an approach called Cavity Quantum Electrodynamics Complete Active Space Configuration Interaction theory to provide ground- and excited-state polaritonic surfaces with a balanced description of strong correlation effects among electronic and photonic degrees of freedom. This method provides a platform for ab initio cavity quantum electrodynamics when both strong electron correlation and strong light-matter coupling are important, and is an important step towards computational approaches that yield multiple polaritonic potential energy surfaces and couplings that can be leveraged for {ab initio molecular dynamics simulations of polariton chemistry.
The discussion around "safe" programming languages has significantly increased in recent years, and is impacting how governments, industry, and academia plan to develop current and future software products. The White House Office of the National Cyber Director released a report [1] in February 2024 calling on the technical community to work towards proactively reducing attack surfaces in cyberspace, in part, specifically by adopting memory safe programming languages. While the main discourse thus far has been focused on cybersecurity, memory safety issues are also a concern in HPC, where memory related errors can result in wasted execution time, incorrect results, etc. Legacy programming languages in HPC such as C and C++ provide freedom and flexibility with memory management, but requires the developer to guarantee safety. While it is possible to develop "un-safe" code in all programming languages, "memory-safe" languages help guarantee safety by utilizing various compile time and runtime checks and validation systemsIn this paper we introduce Lamellar, an asynchronous tasking and PGAS runtime system for HPC written in Rust, one such "memory-safe" language. We describe the entire Lamellar stack, from network interfaces to safe high-level abstractions such as distributed LamellarArrays and Active Messages. The goal of our runtime is to enable end-users to develop entirely safe Rust code in their applications, limiting the use of any "unsafe" code blocks to rigorously tested code blocks within the runtime itself. We conclude by showing comparable performance against several C, C++, and Chapel implementations of a subset of the BALE kernel suite while maintaining strong memory safety principles.
The power of quantum chemistry to predict the ground and excited state properties of complex chemical systems has driven the development of computational quantum chemistry software, integrating advances in theory, applied mathematics, and computer science. The emergence of new computational paradigms associated with exascale technologies also poses significant challenges that require a flexible forward strategy to take full advantage of existing and forthcoming computational resources. In this context, the sustainability and interoperability of computational chemistry software development are among the most pressing issues. In this perspective, we discuss software infrastructure needs and investments with an eye to fully utilize exascale resources and provide unique computational tools for next-generation science problems and scientific discoveries.
Tensor algebra operations such as contractions in computational chemistry consume a significant fraction of the computing time on large-scale computing platforms. The widespread use of tensor contractions between large multi-dimensional tensors in describing electronic structure theory has motivated the development of multiple tensor algebra frameworks targeting heterogeneous computing platforms. In this paper, we present Tensor Algebra for Many-body Methods (TAMM), a framework for productive and performance-portable development of scalable computational chemistry methods. TAMM decouples the specification of the computation from the execution of these operations on available high-performance computing systems. With this design choice, the scientific application developers (domain scientists) can focus on the algorithmic requirements using the tensor algebra interface provided by TAMM, whereas high-performance computing developers can direct their attention to various optimizations on the underlying constructs, such as efficient data distribution, optimized scheduling algorithms, and efficient use of intra-node resources (e.g., graphics processing units). The modular structure of TAMM allows it to support different hardware architectures and incorporate new algorithmic advances. We describe the TAMM framework and our approach to the sustainable development of scalable ground- and excited-state electronic structure methods. We present case studies highlighting the ease of use, including the performance and productivity gains compared to other frameworks.
We report the implementation of the real-time equation-of-motion coupled-cluster (RT-EOM-CC) cumulant Green's function method [ J. Chem. Phys. 2020, 152, 174113] within the Tensor Algebra for Many-body Methods (TAMM) infrastructure. TAMM is a massively parallel heterogeneous tensor library designed for utilizing forthcoming exascale computing resources. The two-body electron repulsion matrix elements are Cholesky-decomposed, and we imposed spin-explicit forms of the various operators when evaluating the tensor contractions. Unlike our previous real algebra Tensor Contraction Engine (TCE) implementation, the TAMM implementation supports fully complex algebra. The RT-EOM-CC singles (S) and doubles (D) time-dependent amplitudes are propagated using a first-order Adams-Moulton method. This new implementation shows excellent scalability tested up to 500 GPUs using the Zn-porphyrin molecule with 655 basis functions, with parallel efficiencies above 90% up to 400 GPUs. The TAMM RT-EOM-CCSD was used to study core photoemission spectra in the formaldehyde and ethyl trifluoroacetate (ESCA) molecules. Simulations of the latter involve as many as 71 occupied and 649 virtual orbitals. The relative quasiparticle ionization energies and overall spectral functions agree well with available experimental results.
Newly developed coupled-cluster (CC) methods enable simulations of ionization potentials and spectral functions of molecular systems in a wide range of energy scales ranging from core-binding to valence. This paper discusses the results obtained with the real-time equation-of-motion CC cumulant (RT-EOM-CC) approach and CC Green's function (CCGF) approaches in applications to the water and water dimer molecules. We compare the ionization potentials obtained with these methods for the valence region with the results obtained with the coupled-cluster with singles, doubles, and perturbative triples formulation as a difference of energies for N and N - 1 electron systems. All methods show good agreement with each other. They also agree well with the experiment with errors usually below 0.1 eV for the ionization potentials. We also analyze unique features of the spectral functions, associated with the position of satellite peaks, obtained with the RT-EOM-CC and CCGF methods employing single and double excitations, as a function of the monomer OH bond length and the proton transfer coordinate in the dimer. Finally, we analyze the impact of the basis set effects on the quality of calculated ionization potentials and find that the basis set effects are less pronounced for the augmented-type sets.
The computational power increases over the past decades have greatly enhanced the ability to simulate chemical reactions and understand ever more complex transformations. Tensor contractions are the fundamental computational building block of these simulations. These simulations have often been tied to one platform and restricted in generality by the interface provided to the user. The expanding prevalence of accelerators and researcher demands necessitate a more general approach which is not tied to specific hardware or requires contortion of algorithms to specific hardware platforms. In this paper we present COMET, a domain-specific programming language and compiler infrastructure for tensor contractions targeting heterogeneous accelerators. We present a system of progressive lowering through multiple layers of abstraction and optimization that achieves up to 1.98x speedup for 30 tensor contractions commonly used in computational chemistry and beyond.
Since the advent of the first computers, chemists have been at the forefront of using computers to understand and solve complex chemical problems. As the hardware and software have evolved, so have the theoretical and computational chemistry methods and algorithms. Parallel computers clearly changed the common computing paradigm in the late 1970s and 80s, and the field has again seen a paradigm shift with the advent of graphical processing units. This review explores the challenges and some of the solutions in transforming software from the terascale to the petascale and now to the upcoming exascale computers. While discussing the field in general, NWChem and its redesign, NWChemEx, will be highlighted as one of the early codesign projects to take advantage of massively parallel computers and emerging software standards to enable large scientific challenges to be tackled.
The widespread use of tensor operations in describing electronic structure calculations has motivated the design of software frameworks for productive development of scalable optimized tensor-based electronic structure methods. Whereas prior work focused on Cartesian abstractions for dense tensors, we present an algebra to specify and perform tensor operations on a larger class of block-sparse tensors. We illustrate the use of this framework in expressing real-world computational chemistry calculations beyond the reach of existing frameworks.
Implementing highly concurrent programs can be challenging because programmers can easily introduce unintended nondeterminism, which has the potential to affect the program output. We propose and implement a technique for detecting unintended nondeterminism in applications developed on shared memory systems with dataflow execution model. Such nondeterminism bugs may be caused by missing or incorrect ordering of task dependencies that are used for ensuring certain ordering of tasks. The proposed method is based on the formulation of happens-before relation on tasks executions in a dataflow dependency graph. Its implementation is composed of two main phases; log recording and detection. For recording the necessary information from the execution, the tool instruments the dataflow framework and the applications, on top of the LLVM compiler infrastructure. Later it processes the collected log and reports on the found output nondeterminism in the execution. The tool can integrate well with the development cycle to provide the programmer with a testing framework against possible nondeterminism bugs. To demonstrate its effectiveness, we study a set of benchmark applications written in Atomic DataFlow programming model and report on real nondeterminism bugs in them. (C) 2017 Elsevier B.V. All rights reserved.
Modern geo-replicated data stores provide high availability by relaxing the underlying consistency requirements. Programs layered over such data stores are called weakly consistent programs. Due to the reduced consistency requirements, they exhibit highly nondeterministic behaviors, some of which might violate program invariants. Therefore, implementing correct weakly consistent programs and reasoning about them is challenging. In this paper, we present a systematic scheduling approach that is aware of the underlying consistency model. Our approach dynamically explores all possible program behaviors allowed by the used data store consistency model, and it evaluates program invariants during the exploration. We implement the approach in a prototype model checker for Antidote, which is a causally consistent key-value data store with convergent con ict handling. We evaluate our tool on several benchmarks. The results show that our approach is e effective in detecting buggy behaviors in weakly consistent programs.
As HPC platforms get increasingly complex, the complexity of software optimized for these platforms has also increased. There is a pressing need to ensure correctness of scientific applications to enhance our confidence in the results they produce. In this paper, we focus on checking the functional equivalence of libraries providing a small but important functionality---tensor transposition---used in computational chemistry applications. While several correctness tools have been developed and deployed, there are several practical challenges in using them to check correctness of production HPC software. We present our experiences using two tools---CIVL and CodeThorn---in checking the functional equivalence of two index permutation libraries. We observe that, with some effort, the tools we evaluated can handle kernels from production codes. We present observations that will aid library writers to write code that can be checked with these tools.
As JavaScript has become virtually omnipresent as the language for programming large and complex web applications in the last several years, we have seen an increase in interest in finding data races in client-side JavaScript. While JavaScript execution is single-threaded, there is still enough potential for data races, created largely by the non-determinism of the scheduler. Recently, several academic efforts have explored both static and run-time analysis approaches in an effort to find data races. However, despite this, we have not seen these analysis techniques deployed in practice and we have only seen scarce evidence that developers find and fix bugs related to data races in JavaScript. In this paper we argue for a different formulation of what it means to have a data race in a JavaScript application and distinguish between benign and harmful races, affecting persistent browser or server state. We further argue that while benign races — the subject of the majority of prior work — do exist, harmful races are exceedingly rare in practice (19 harmful vs. 621 benign). Our results shed a new light on the issues of data race prevalence and importance. To find races, we also propose a novel lightweight run-time symbolic exploration algorithm for finding races in traces of run-time execution. Our algorithm eschews schedule exploration in favor of smaller run-time overheads and thus can be used by beta testers or in crowd-sourced testing. In our experiments on 26 sites, we demonstrate that benign races are considerably more common than harmful ones.
We present a dynamic verification technique for a class of concurrent programming models that combine dataflow and shared memory programming. In this class of hybrid concurrency models, programs are built from tasks whose data dependencies are explicitly defined by a programmer and used by the runtime system to coordinate task execution. Differently from pure dataflow, tasks are allowed to have shared state which must be properly protected using synchronization mechanisms, such as locks or transactional memory (TM). While these hybrid models enable programmers to reason about programs, especially with irregular data sharing and communication patterns, at a higher level, they may also give rise to new kinds of bugs as they are unfamiliar to the programmers. We identify and illustrate a novel category of bugs in these hybrid concurrency programming models and provide a technique for randomized exploration of program behaviors in this setting.
Despite JavaScript runtime’s lack of conventional threads, the presence of asynchrony creates a real potential for concurrency errors. These concerns have lead to investigations of race conditions in the Web context. However, focusing on races does not produce actionable error reports that would at the end of the day appeal to developers and cause them to fix possible underlying problems. In this paper, we advocate for the notion of observable races, focusing on concurrency conditions that lead to visually apparent glitches caused by non-determinism within the runtime scheduler on the network. We propose and investigate ways to find observable races via systematically exploring possible network schedules and shepherding the scheduler towards correct executions. We propose crowd-sourcing both to spot when different schedules lead to visually broken sites and also to determine under what environment conditions (OS, browser, network speed) these schedules may in fact happen in practice for some fraction of the users.