Over the last decade block-structured adaptive mesh refinement (SAMR) has found increasing use in large, publicly available codes and frameworks. SAMR frameworks have evolved along different paths. Some have stayed focused on specific domain areas, others have pursued a more general functionality, providing the building blocks for a larger variety of applications. In this survey paper we examine a representative set of SAMR packages and SAMR-based codes that have been in existence for half a decade or more, have a reasonably sized and active user base outside of their home institutions, and are publicly available. The set consists of a mix of SAMR packages and application codes that cover a broad range of scientific domains. We look at their high-level frameworks, their design trade-offs and their approach to dealing with the advent of radical changes in hardware architecture. The codes included in this survey are BoxLib, Cactus, Chombo, Enzo, FLASH, and Uintah.
The design of hardware for next-generation exascale computing systems will require a deep understanding of how software optimizations impact hardware design trade-offs. In order to characterize how co-tuning hardware and software parameters affects the performance of combustion simulation codes, we created ExaSAT, a compiler-driven static analysis and performance modeling framework. Our framework can evaluate hundreds of hardware/ software configurations in seconds, providing an essential speed advantage over simulators and dynamic analysis techniques during the co-design process. Our analytic performance model shows that advanced code transformations, such as cache blocking and loop fusion, can have a significant impact on choices for cache and memory architecture. Our modeling helped us identify tuned configurations that achieve a 90% reduction in memory traffic, which could significantly improve performance and reduce energy consumption. These techniques will also be useful for the development of advanced programming models and runtimes, which must reason about these optimizations to deliver better performance and energy efficiency.
We describe a second-order accurate sequential algorithm for solving two-phase multicomponent flow in porous media. The algorithm incorporates an unsplit second-order Godunov scheme that provides accurate resolution of sharp fronts. The method is implemented within a block structured adaptive mesh refinement (AMR) framework that allows grids to dynamically adapt to features of the flow and enables efficient parallelization of the algorithm. We demonstrate the second-order convergence rate of the algorithm and the accuracy of the AMR solutions compared to uniform fine-grid solutions. The algorithm is then used to simulate the leakage of gas from a Liquified Petroleum Gas (LPG) storage cavern, demonstrating its capability to capture complex behavior of the resulting flow. We further examine differences resulting from using different relative permeability functions.
In this paper we investigate the performance of the Porous Media with Adaptive Mesh Refinment (PMAMR) code which was developed in the Center for Computational Science and Engineering at Lawrence Berkeley National Laboratory. This code is being used to model carbon sequestration and contaminant transport as part of the Advanced Simulation Capability for Environmental Management (ASCEM) project. The goal of the ASCEM project is to better understand and quantify flow and contaminant transport behavior in complex geological systems. It will also address the long-term performance of engineered components including cementitious materials in nuclear waste disposal facilities, in order to reduce uncertainties and risks associated with DOE EM’s environmental cleanup and closure activities. For this paper we have chosen a specific, computationally-intensive problem involving uranium contamination at a site in Georgia. Currently we are interested in how long it will take to simulate 25 years of flow, because this roughly corresponds to the time between experimental measurements; for example, see the measurements taken in Figure 1. Eventually, we would like to be able to simulate to 100 years or more, but this is likely to require the development of new algorithms beyond simple code tuning. The PMAMR code is built upon the BoxLib framework and was previously parallelized using MPI [4]. In this work we used tools on the XE6 platform to identify performance bottlenecks and show how we can simulate 25 years of data faster by adopting a hybrid MPI/OpenMP programming model. Overall our hybrid approach yields a net speedup of 2.6× compared to the MPI only version, with substantially reduced memory usage. We note that while the use of OpenMP and MPI leads to considerable speed-ups, the calculations would still take on the order of a year to complete. Thus increasing the performance of our code with a hybrid programming model on Hopper is a first step towards the goal of simulating 25 years of flow.
This paper is the first in a series that presents a combined computational and experimental study to investigate and characterize the structure of premixed turbulent low swirl laboratory flames. The simulations discussed here are based on an adaptive solution of the low Mach number equations for turbulent reacting flow, and incorporate detailed models for transport and thermo-chemistry. Experimental diagnostics of the laboratory flame include PIV and OH-PLIF imaging, and are used to quantify the flow field, mean flame location, and local flame wrinkling characteristics. We present a framework for relating the simulation results to the flame measurements, and then use the simulation data to further probe the time-dependent, 3D structure of the flames as they interact with the turbulent flow. The present study is limited to lean methane–air flames over a range of flow conditions, and demonstrates that in the regime studied, local flame profiles are structurally very similar to the flat, unstrained steady (“laminar”) flame. The analysis here will serve as a framework for discussing a broader set of premixed flames in this same configuration. Papers II and III will discuss corresponding analysis for pure hydrogen–air and hydrogen–methane mixed fuels, respectively.
Performing high-resolution, high-fidelity, three-dimensional simulations of Type Ia supernovae (SNe Ia) requires not only algorithms that accurately represent the correct physics, but also codes that effectively harness the resources of the most powerful supercomputers. We are developing a suite of codes that provide the capability to perform end-to-end simulations of SNe Ia, from the early convective phase leading up to ignition to the explosion phase in which deflagration/detonation waves explode the star to the computation of the light curves resulting from the explosion. In this paper we discuss these codes with an emphasis on the techniques needed to scale them to petascale architectures. We also demonstrate our ability to map data from a low Mach number formulation to a compressible solver.
Simulations are routinely used to study the process of carbon dioxide (CO2) sequestration in saline aquifers. In this paper, we describe the modeling and simulation of the dissolution–diffusion–convection process based on a total velocity splitting formulation for a variable-density incompressible single-phase model. A second-order accurate sequential algorithm, implemented within a block-structured adaptive mesh refinement (AMR) framework, is used to perform high-resolution studies of the process. We study both the short-term and long-term behaviors of the process. It is found that the onset time of convection follows closely the prediction of linear stability analysis. In addition, the CO2 flux at the top boundary, which gives the rate at which CO2 gas dissolves into a negatively buoyant aqueous phase, will reach a stabilized state at the space and time scales we are interested in. This flux is found to be proportional to permeability, and independent of porosity and effective diffusivity, indicative of a convection-dominated flow. A 3D simulation further shows that the added degrees of freedom shorten the onset time and increase the magnitude of the stabilized CO2 flux by about 25%. Finally, our results are found to be comparable to results obtained from TOUGH2-MP.
Simulations are routinely used to study the process of carbon dioxide (CO 2 ) sequestration in saline aquifers. In this paper, we look at some numerical aspects of the accurate modeling and simulation of the dissolution-diffusion-convection process. We perform convergence studies with respect to solver tolerances, grid resolutions, fluctuation strength, and domain size. We show that stringent tolerances and grid resolutions are needed to accurately predict onset time. Domain size must be sufficiently large to contain at least 2 extended fingers to accurately predict the long-term stabilized mass flux of CO 2 ; otherwise, finite domain effects will adversely change the flow behavior of the system we are modeling.
In this paper, we present a second-order accurate adaptive algorithm for solving multi-phase, incompressible flow in porous media. We assume a multi-phase form of Darcy's law with relative permeabilities given as a function of the phase saturation. The remaining equations express conservation of mass for the fluid constituents. In this setting, the total velocity, defined to be the sum of the phase velocities, is divergence free. The basic integration method is based on a total-velocity splitting approach in which we solve a second-order elliptic pressure equation to obtain a total velocity. This total velocity is then used to recast component conservation equations as nonlinear hyperbolic equations. Our approach to adaptive refinement uses a nested hierarchy of logically rectangular grids with simultaneous refinement of the grids in both space and time. The integration algorithm on the grid hierarchy is a recursive procedure in which coarse grids are advanced in time, fine grids are advanced multiple steps to reach the same time as the coarse grids and the data at different levels are then synchronized. The single-grid algorithm is described briefly, but the emphasis here is on the time-stepping procedure for the adaptive hierarchy. Numerical examples are presented to demonstrate the algorithm's accuracy and convergence properties and to illustrate the behaviour of the method.
We present numerical simulations of lean hydrogen flames interacting with turbulence. The simulations are performed in an idealized setting using an adaptive low Mach number model with a numerical feedback control algorithm to stabilize the flame. At the conditions considered here, hydrogen flames are thermodiffusively unstable, and burn in cellular structures. For that reason, we consider two levels of turbulence intensity and a case without turbulence whose dynamics is driven by the natural flame instability. An overview of the flame structure shows that the burning in the cellular structures is quite intense, with the burning patches separated by regions in which the flame is effectively extinguished. We explore the geometry of the flame surface in detail, quantifying the mean and Gaussian curvature distributions and the distribution of the cell sizes. We next characterize the local flame speed to quantify the effect of flame intensification on local propagation speed. We then introduce several diagnostics aimed at quantifying both the level of intensification and diffusive mechanisms that lead to the intensification.
We present a second-order accurate adaptive algorithm for solving reactive transport flow in geochemical systems. A Strang-splitting approach is used to numerically decouple the transport component and the reaction component of the problem. The transport component is solved using a second-order accurate IMPES-like algorithm that exhibits excellent control of numerical dispersion. The numerical scheme for the reaction component uses reaction network details from the chemistry module of TOUGHREACT, COREREACT, and VODE, a high-order ODE integrator, to maintain second-order accuracy of the algorithm. The algorithm is implemented within an adaptive refinement framework that uses a nested hierarchy of logically rectangular grids with simultaneous refinement of the grids in both space and time. The integration algorithm on the grid hierarchy is a recursive procedure in which coarse grids are advanced in time, fine grids are advanced multiple steps to reach the same time as the coarse grids, and the data at different levels are then synchronized. Numerical examples are presented to demonstrate the algorithm's accuracy and convergence properties. We conclude with a simulation of a reactive salt dome problem.
The theory of turbulent premixed flames is based on a characterization of the flame as a discontinuous surface propagating through the fluid. The displacement speed, defined as the local speed of the flame front normal to itself, relative to the unburned fluid, provides one characterization of the burning velocity. In this paper, we introduce a geometric approach to computing displacement speed and discuss the efficacy of the displacement speed for characterizing a turbulent flame.
After a decade where HEC (high-end computing) capability was dominated by the rapid pace of improvements to CPU clock frequency, the performance of next-generation supercomputers is increasingly differentiated by varying interconnect designs and levels of integration. Understanding the tradeoffs of these system designs, in the context of high-end numerical simulations, is a key step towards making effective petascale computing a reality. This work represents one of the most comprehensive performance evaluation studies to date on modern NEC systems, including the IBM Power5, AMD Opteron, IBM BG/L, and Cray X1E. A novel aspect of our study is the emphasis on full applications, with real input data at the scale desired by computational scientists in their unique domain. We examine six candidate ultra-scale applications, representing a broad range of algorithms and computational structures. Our work includes the highest concurrency experiments to date on five of our six applications, including 32K processor scalability for two of our codes and describe several successful optimizations strategies on BG/L, as well as improved X1E vectorization. Overall results indicate that our evaluated codes have the potential to effectively utilize petascale resources; however, several applications would require reengineering to incorporate the additional levels of parallelism necessary to achieve the vast concurrency of upcoming ultra-scale systems.
After a decade where HEC (high-end computing) capability was dominated by the rapid pace of improvements to CPU clock frequency, the performance of next-generation supercomputers is increasingly differentiated by varying interconnect designs and levels of integration. Understanding the trade-offs of these system designs, in the context of high-end numerical simulations, is a key step towards making effective petascale computing a reality. This work represents one of the most comprehensive performance evaluation studies to date on modern HEC systems, including the IBM Power5, AMD Opteron, IBM BG/L, and Cray X1E. A novel aspect of our study is the emphasis on full applications, with real input data at the scale desired by computational scientists in their unique domain. We examine five candidate ultra-scale applications, representing a broad range of algorithms and computational structures. Our work includes the highest concurrency experiments to date on five of our six applications, including 32K processor scalability for two of our codes and describes several successful optimization strategies on BG/L, as well as improved X1E vectorization. Overall results indicate that our evaluated codes have the potential to effectively utilize petascale resources; however, several applications will require reengineering to incorporate the additional levels of parallelism necessary to utilize the vast concurrency of upcoming ultra-scale systems.
In this paper, we discuss some of the issues in obtaining high performance for block-structured adaptive mesh refinement software for partial differential equations. We show examples in which AMR scales to thousands of processors. We also discuss a number of metrics for performance and scalability that can provide a basis for understanding the advantages and disadvantages of this approach.
As we enter the era of peta-scale computing, system architects must plan for machines composed of tens or even hundreds of thousands of processors. Although fully connected networks such as fat-tree configurations currently dominate HPC interconnect designs, such approaches are inadequate for ultra-scale concurrencies due to the superlinear growth of component costs. Traditional low-degree interconnect topologies, such as 3D tori, have reemerged as a competitive solution due to the linear scaling of system components relative to the node count; however, such networks are poorly suited for the requirements of many scientific applications at extreme concurrencies. To address these limitations, we propose HFAST, a hybrid switch architecture that uses circuit switches to dynamically reconfigure lower-degree interconnects to suit the topological requirements of a given scientific application. This work presents several new research contributions. We develop an optimization strategy for HFAST mappings and demonstrate that efficiency gains can be attained across a broad range of static numerical computations. Additionally, we conduct an extensive analysis of the communication characteristics of a dynamically adapting mesh calculation and show that the HFAST approach can achieve significant advantages, even when compared with traditional fat-tree configurations. Overall results point to the promising potential of utilizing hybrid reconfigurable networks to interconnect future peta-scale architectures, for both static and dynamically adapting applications.
The speed of propagation of a premixed turbulent flame correlates with the intensity of the turbulence encountered by the flame. One consequence of this property is that premixed flames in both laboratory experiments and practical combustors require some type of stabilization mechanism to prevent blow-off and flashback. The stabilization devices often introduce a level of geometric complexity that is prohibitive for detailed computational studies of turbulent flame dynamics. Furthermore, the stabilization introduces additional fluid mechanical complexity into the overall combustion process that can complicate the analysis of fundamental flame properties. To circumvent these difficulties we introduce a feedback control algorithm that allows us to computationally stabilize a turbulent premixed flame in a simple geometric configuration. For the simulations, we specify turbulent inflow conditions and dynamically adjust the integrated fueling rate to control the mean location of the flame in the domain. We outline the numerical procedure, and illustrate the behavior of the control algorithm on methane flames at various equivalence ratios in two dimensions. The simulation data are used to study the local variation in the speed of propagation due to flame surface curvature.
Turbulent methane flames exhibit a change in Markstein number as the equivalence ratio changes. In this paper, we demonstrate how these changes in Markstein number are related to shifts in the chemistry and transport. We focus on the analysis of simulations of two-dimensional premixed turbulent methane flames at equivalence ratios =0.55 and =1.00 computed using the GRI-Mech 3.0 mechanism. The simulations were performed using a low Mach number adaptive mesh refinement algorithm coupled to an automatic feedback control algorithm that stabilizes the flame on the computational grid. The flames are characterized in terms of curvature, strain and a local fuel-consumption- based flame speed. Joint probability density functions show correlations of the local flame speed with curvature and a lack of correlation with the tangential strain rate. We then introduce a pathline diagnostic that follows parcels of fluid through the flame, enabling us to quantify the associated reaction and diusive transport processes. Using this diagnostic, we examine the dierences in the chemical and transport properties of the two flames, and isolate those responsible for the shift in the correlation of local flame speed with flame surface curvature.