The evolution of molecular dynamics (MD) simulations has been intimately linked to that of computing hardware. For decades following the creation of MD, simulations have improved with computing power along the three principal dimensions of accuracy, atom count (spatial scale), and duration (temporal scale). Since the mid-2000s, computer platforms have, however, failed to provide strong scaling for MD, as scale-out central processing unit (CPU) and graphics processing unit (GPU) platforms that provide substantial increases to spatial scale do not lead to proportional increases in temporal scale. Important scientific problems therefore remained inaccessible to direct simulation, prompting the development of increasingly sophisticated algorithms that present significant complexity, accuracy, and efficiency challenges. While bespoke MD-only hardware solutions have provided a path to longer timescales for specific physical systems, their impact on the broader community has been mitigated by their limited adaptability to new methods and potentials. In this work, we show that a novel computing architecture, the Cerebras wafer scale engine, completely alters the scaling path by delivering unprecedentedly high simulation rates up to 1.144 M steps/s for 200 000 atoms whose interactions are described by an embedded atom method potential. This enables direct simulations of the evolution of materials using general-purpose programmable hardware over millisecond timescales, dramatically increasing the space of direct MD simulations that can be carried out. In this paper, we provide an overview of advances in MD over the last 60 years and present our recent result in the context of historical MD performance trends.
Simulation of physical systems is essential in many scientific and engineering domains. Commonly used domain decomposition methods are unable to deliver high simulation rate or high utilization in network computing environments. In particular, Exascale systems deliver only a small fraction their peak performance for these workloads. This paper introduces the novel \algorithmpropernoun{} algorithm, designed to overcome these limitations. We apply this method and show simulations running in excess of 1.6 million time steps per second and simulations achieving 84 PFLOP/s. Our implementation can achieve 90\% of peak performance in both single-node and clustered environments. We illustrate the method by applying the shallow-water equations to model a tsunami following an asteroid impact at 460m-resolution on a planetary scale running on a cluster of 64 Cerebras CS-3 systems.
Molecular dynamics (MD) simulations have transformed our understanding of the nanoscale, driving breakthroughs in materials science, computational chemistry, and several other fields, including biophysics and drug design. Even on exascale supercomputers, however, runtimes are excessive for systems and timescales of scientific interest. Here, we demonstrate strong scaling of MD simulations on the Cerebras Wafer-Scale Engine. By dedicating a processor core for each simulated atom, we demonstrate a 457-fold improvement in timesteps per second versus the Frontier GPU-based Exascale platform, along with a large improvement in timesteps per unit energy. Reducing every year of runtime to less than a day unlocks currently inaccessible timescales of slow microstructure transformation processes that are critical for understanding material behavior and function. Our dataflow algorithm runs Embedded Atom Method (EAM) simulations at rates over 699k timesteps per second for problems with up to 800k atoms. This demonstrated performance is unprecedented for general-purpose processing cores.
Poetry & Strikes examines shifting representations of strike action in the work of six British poets from the 1970s to the present day. It considers how these poets have come to contend with, and contribute to, narratives surrounding industrial disputes. Through these conversations, the book attempts to question the way in which union narratives and legacies are constructed, and to investigate the power dynamics that underpin the presentation of labour histories. The work of these poets helps us to understand how cultural memories have been formed, and makes it possible to see how these legacies may still be rewritten and reframed. The book focuses on two distinct poetic eras: from 1972 to the late 1980s through the work of Barry MacSweeney, Tony Harrison and Sean O’Brien, and the 2010s, through the writings of Helen Mort, Steve Ely and Paul Bentley. The poets bring influences from other poems, from movies and from ‘official’ narratives on to the page, sometimes brazenly, sometimes obscured, but always with a sense that they are writing out of and into a strike history. The book explores these influences and their engagement with strike history. The poets examined see these legacies as artful constructions, foregrounding an awareness of working-class stories and histories as products, products that have been manufactured and arranged, in the same way that the poems themselves are artfully rendered and established on the page.
Arash Abizadeh argues that all coercive enforcement of borders is democratically illegitimate, since foreigners do not participate in the creation of border laws. It is irrelevant whether the border laws are substantively just or unjust, whether the state enforcing them is affluent or poor, and whether the individual being coerced autonomously chooses to cross the border or is forced by desperate circumstances to do so. His argument involves (1) a foundational commitment to individual autonomy; (2) a normative premise that coercion requires democratic legitimation; (3) and an empirical premise that border enforcement laws subject all foreigners to state coercion. In this essay, I contest each of these components. I challenge the empirical premise through examples illustrating the empirical limits to state coercion over foreigners. I contest the normative premise by showing that state coercion requires democratic legitimation only for those involuntarily and indefinitely subject to it. Finally, I challenge the commitment to individual autonomy as foundational to political legitimacy by distinguishing political legitimacy from political authority. I conclude by demonstrating how my critique renders a more plausible account of the normative limits of border coercion, one that coheres more readily with stances advanced by Javier Hidalgo and Abizadeh himself.
In 2015, Spain passed a law expediting citizenship for the descendants of the Sephardic Jews expelled in 1492, but not to the descendants of the Moriscos expelled in 1609. In this essay, I use Spain's 2015 citizenship law as a test case for assessing three normative models for linking citizenship with collective responsibility for the past: reparations for historic injustice; the principle of coercively constituted identities; and remedial responsibility. I argue that the first two models confront intractable philosophical problems that are circumvented by the third model, remedial responsibility, which prioritizes contemporary suffering and looks to the past only to identify agents who must provide a remedy. However, remedial responsibility subordinates the obligation to expedite citizenship to descendants of the Sephardim and the Moriscos in favor of the Saharawis, citizens of the former Spanish colony of Western Sahara who still languish in refugee camps forty years after decolonization.
Fast neutron identification and spectroscopy is of great interest to nuclear physics experiments. Using the neutron elastic scattering, the fast neutron momentum can be measured. Wang and Morris introduced the theoretical concept that the initial fast neutron momentum can be derived from up to three consecutive elastic collisions between the neutron and the target, including the information of two consecutive recoil ion tracks and the vertex position of the third collision or two consecutive elastic collisions with the timing information. Here, we also include the additional possibility of measuring the deposited energies from the recoil ions. In this paper, we simulate the neutron elastic scattering using the Monte Carlo N-Particle Transport Code (MCNP) and study the corresponding neutron detection and tracking efficiency. The corresponding efficiency and the scattering distances are simulated with different target materials, especially natural silicon (92.23%28Si, 4.67%29Si, and 3.1%30Si) and helium-4 (4He). The timing of collision and the recoil ion energy are also investigated, which are important characters for the detector design. We also calculate the ion traveling range for different energies using the software, “The Stopping and Range of Ions in Matter (SRIM)”, showing that the ion track can be most conveniently observed in 4He unless sub-micron spatial resolution can be obtained in silicon.
Do past state actions, such as the American conquest of northern Mexico, the British colonization of South Asia, and the Spanish expulsion of the Sephardim and Moriscos, grant contemporary Mexicans, South Asians, and the descendants of the Sephardim and Moriscos a particular right to immigrate to the United States, the United Kingdom, and Spain respectively? In this paper I examine three theoretical models for addressing this question: retrospective responsibility for historic injustice; the principle of coercively constituted identities; and the theory of remedial responsibility. I argue that remedial responsibility best justifies a particular right to immigrate through responsibility for the past for three reasons. First, it relieves us of the epistemological task of establishing causal responsibility. Second, it lessens the normative task of identifying a theory of unjust harm to establish moral responsibility. Finally, it facilitates the normative task of ranking the claims to immigrate of different individuals.
We present a high-level and accessible Application Programming Interface (API) for the solution of field equations on the Cerebras Systems Wafer-Scale Engine (WSE) with over two orders of magnitude performance gain relative to traditional distributed computing approaches. The domain-specific API is called the WSE Field-equation API (WFA). The WFA outperforms OpenFOAM on NETL's Joule 2.0 supercomputer by over two orders of magnitude in time to solution. While this performance is consistent with hand-optimized assembly codes, the WFA provides an easy-to-use, high-level Python interface that allows users to form and solve field equations effortlessly. We report here the WFA programming methodology and achieved performance on the latest generation of WSE, the CS-2.
ABSTRACTSolving 3-D partial differential equations in a Finite Element model is computationally intensive and requires extremely high memory and communication bandwidth. This paper describes a novel way where the Finite Element mesh points of varying resolution are mapped on a large 2-D homogenous array of processors. Cerebras developed a novel supercomputer that is powered by a 21.5cm by 21.5cm Wafer-Scale Engine (WSE) with 850,000 programmable compute cores. With 2.6 trillion transistors in a 7nm process this is by far the largest chip in the world. It is structured as a regular array of 800 by 1060 identical processing elements, each with its own local fast SRAM memory and direct high bandwidth connection to its neighboring cores. For the 2021 ISPD competition we propose a challenge to optimize placement of computational physics problems to achieve the highest possible performance on the Cerebras supercomputer. The objectives are to maximize performance and accuracy by optimizing the mapping of the problem to cores in the system. This involves partitioning and placement algorithms.
High-energy (>20 keV) X-ray photon detection at high quantum yield, high spatial resolution, and short response time has long been an important area of study in physics. Scintillation is a prevalent method but limited in various ways. Directly detecting high-energy X-ray photons has been a challenge to this day, mainly due to low photon-to-photoelectron conversion efficiencies. Commercially available state-of-the-art Si direct detection products such as the Si charge-coupled device (CCD) are inefficient for >10 keV photons. Here, we present Monte Carlo simulation results and analyses to introduce a highly effective yet simple high-energy X-ray detection concept with significantly enhanced photon-to-electron conversion efficiencies composed of two layers: a top high-Z photon energy attenuation layer (PAL) and a bottom Si detector. We use the principle of photon energy down conversion, where high-energy X-ray photon energies are attenuated down to ≤10 keV via inelastic scattering suitable for efficient photoelectric absorption by Si. Our Monte Carlo simulation results demonstrate that a 10–30× increase in quantum yield can be achieved using PbTe PAL on Si, potentially advancing high-resolution, high-efficiency X-ray detection using PAL-enhanced Si CMOS image sensors.
Neutron radiography through Spectroscopic Imaging by Fast Neutrons (SIFaN) is described. The fast neutron tracking principle [1] has been extended to include neutron-induced fissions in actinide perovskites. The design and performance of a new SIFaN instrument, called SIFaN-perovskite (or SIFaN-P), are studied by a combination of MCNP and semi-analytical models. Hybrid organic-inorganic and actinide perovskites are considered and compared with gas and silicon for neutron detection and tracking. The Los Alamos LANSCE facility provides access to the initial SIFaN-P testing and demonstration.
The advent of smartphone technology has provided us with intelligent devices for communication as well as playing game. Unfortunately, applications that exploit available sensors in the smartphone are mostly designed for people with no physical handicap. This paper presents Mata, a game user interface using eye-tracking to operate and control games running on Android smartphone. This system is designed to enhance user experiences and help motoric impaired peoples in using smartphone for playing games. Development and testing of the Mata system has proven the concepts of eye-tracking and eyegazing usage as unimodal input for game user interface.
In this work, we present a new analysis method applied to revitalize permanent magnet Compton spectrometers used to measure photon energy spectra in the MeV range. The inversion of the measured electron distribution to determine the original photon distribution is achieved via a method of consistent coupled radiation transport and magnetic field mapping of the input photon spectra to the measured electron distribution. The method of linear least squares was used to perform the unfolding of the electron distribution to the initial photon spectra, without any assumptions made regarding the electron distribution. We present an application of this method to data from a nominal 19.4 MeV flash radiographic source (the first axis of the Dual Axis Radiographic Hydro-Test Facility) capable of generating 500 R @ 1 m in ∼60 ns and a medical therapy source (a Scanditronix M22, Microtron) capable of variable energies with nominal endpoints of 6, 10, 15, and 20 MeV and an output of ∼1000-2000 R/min @ 1 m. The results provide agreement between the modeled and unfolded experimentally measured photon spectra as quantified by statistical tests, from 1.5 to 20 MeV. Experimental results are presented as well as a discussion of the novel MCNP6-based simulations and methods for reconstruction of the spectra.
The performance of CPU-based and GPU-based systems is often low for PDE codes, where large, sparse, and often structured systems of linear equations must be solved. Iterative solvers are limited by data movement, both between caches and memory and between nodes. Here we describe the solution of such systems of equations on the Cerebras Systems CS-1, a wafer-scale processor that has the memory bandwidth and communication latency to perform well. We achieve 0.86 PFLOPS on a single wafer-scale system for the solution by BiCGStab of a linear system arising from a 7-point finite difference stencil on a 600 X 595 X 1536 mesh, achieving about one third of the machine's peak performance. We explain the system, its architecture and programming, and its performance on this problem and related problems. We discuss issues of memory capacity and floating point precision. We outline plans to extend this work towards full applications.
This paper introduces a special case of the floorplanning problem for optimizing neural networks to run on a wafer-scale computing engine. From a compute perspective, neural networks can be represented by a deeply layered structure of compute kernels. During the training of a neural network, gradient descent is used to determine the weight factors. Each layer then uses a local weight tensor to transform "activations" and "gradients" that are shared among connected kernels according to the topology of the network. This process is computationally intensive and requires high memory and communication bandwidth. Cerebras has developed a novel computer system designed for this work that is powered by a 21.5cm by 21.5cm wafer-scale processor with 400,000 programmable compute cores. It is structured as a regular array of 633 by 633 processing elements, each with its own local high bandwidth SRAM memory and direct high bandwidth connection to its neighboring cores. In addition to supporting traditional execution models for neural network training and inference, this engine has a unique capability to compile and compute every layer of a complete neural network simultaneously. Mapping a neural network in this fashion onto Cerebras' Wafer-Scale Engine (WSE) is reminiscent of the traditional floorplanning problem in physical design. A kernel ends up as a rectangle of x by y compute elements. These are the flexible blocks that need to be placed to optimize performance. This paper describes an ISPD 2020 challenge to develop algorithms and heuristics that produce compiled neural networks that achieve the highest possible performance on the Cerebras WSE.
We report a Monte Carlo feasibility study on a high-energy (20-50keV) X-ray detection concept using photon energy attenuation with high-Z materials to enhance the efficiency by more than 10 times compared to the state-of-the-art technologies.