AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.
Anecdotal evidence suggests that Research Software Engineers (RSEs) and Software Engineering Researchers (SERs) often use different terminologies for similar concepts, creating communication challenges. To better understand these divergences, we have started investigating how SE fundamentals from the SER community are interpreted within the RSE community, identifying aligned concepts, knowledge gaps, and areas for potential adaptation. Our preliminary findings reveal opportunities for mutual learning and collaboration, and our systematic methodology for terminology mapping provides a foundation for a crowd-sourced extension and validation in the future.
This report summarizes insights from the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science, which convened more than 40 experts from national laboratories, academia, industry, and community organizations to chart a path toward more powerful, sustainable, and collaborative scientific software ecosystems. To address urgent challenges at the intersection of high-performance computing (HPC), AI, and scientific software, participants envisioned agile, robust ecosystems built through socio-technical co-design–the intentional integration of social and technical components as interdependent parts of a unified strategy. This approach combines advances in AI, HPC, and software with new models for cross-disciplinary collaboration, training, and workforce development. Key recommendations include building modular, trustworthy AI-enabled scientific software systems; enabling scientific teams to integrate AI systems into their workflows while preserving human creativity, trust, and scientific rigor; and creating innovative training pipelines that keep pace with rapid technological change. Pilot projects were identified as near-term catalysts, with initial priorities focused on hybrid AI/HPC infrastructure, cross-disciplinary collaboration and pedagogy, responsible AI guidelines, and prototyping of public-private partnerships. This report presents a vision of next-generation ecosystems for scientific computing where AI, software, hardware, and human expertise are interwoven to drive discovery, expand access, strengthen the workforce, and accelerate scientific progress.
High-quality research software is a cornerstone of modern scientific progress, enabling researchers to analyze complex data, simulate phenomena, and share reproducible results. However, creating such software requires adherence to best practices that ensure robustness, usability, and sustainability. This paper presents ten guidelines for producing high-quality research software, covering every stage of the development lifecycle. These guidelines emphasize the importance of planning, writing clean and readable code, using version control, and implementing thorough testing strategies. Additionally, they address key principles such as modular design, reproducibility, performance optimization, and long-term maintenance. The paper also highlights the role of documentation and community engagement in enhancing software usability and impact. By following these guidelines, researchers can create software that advances their scientific objectives and contributes to a broader ecosystem of reliable and reusable research tools. This work serves as a practical resource for researchers and developers aiming to elevate the quality and impact of their research software.
In the HPC area, both hardware and software move quickly. Often new hardware is developed and deployed, the corresponding software stack, including compilers and other tools, are under active development while leading edge software developers are working to port and tune their applications, all at the same time. While the software ecosystem is in flux, one of the key challenges for users is obtaining insight into the state of implementation of key features in the programming languages and models their applications are using - whether they have been implemented, and whether the implementation conforms to the specification, especially for newly implemented features (less tested by widespread use). OpenMP is one of the most prominent shared memory programming models used for on-node programming in HPC. With the shift towards accelerators (such as GPUs and FPGAs) and heterogeneous programming OpenMP features are getting more complex. It is natural to ask whether generative AI approaches, and large language models (LLMs) in particular, can help in producing validation and verification test suites to allow users better and faster insights into the availability and correctness of OpenMP features of interest. In this work, we explore the use of ChatGPT-4 to generate a suite of tests for OpenMP features. We have chosen a set of directives and clauses, a total of 78 combinations, which first appeared in OpenMP 3.0 (released in May 2008) but are also relevant for accelerators. We prompted ChatGPT to generate tests in the C and Fortran languages, for both host (CPU) and device (accelerator). On the Summit super-computer using the GNU implementation, we found that, of the 78 generated tests 67 C tests and 43 Fortran tests compiled successfully and fewer than those executed to completion. On further analysis we show that not all generated tests are valid. We document the process, results, and provide detailed analysis regarding the quality of tests generated. With the aim of providing input to a production quality validation and verification suite, we manually implement the corrections required to make the tests valid according to the current OpenMP specification. We quantify this effort as small, medium, or large, and record the lines of code changed to correct the invalid tests. With the corrected tests we validate recent implementations from HPE, AMD, and GNU on the Frontier supercomputer. Our experiment and subsequent analysis show that although LLMs are capable of producing HPC specific codes, they are limited by their understanding of the deeper semantics and restrictions of programming models such as OpenMP. Unsurprisingly more commonly used features have better support, while some OpenMP 3.0 directives such as sections and tasking are not universally supported on accelerators. We demonstrate that successful compilation and execution to completion are inadequate metrics for evaluating generated code and that, at this time, commodity LLMs require expert intervention for code verification. This points to gaps in the training data that is currently available for HPC. We demonstrate that with "small" effort 37% of generated invalid C tests and 63% of generated invalid Fortran tests could be corrected. This improves productivity of test generation as we circumvent writing from scratch and the common programming errors associated with it.
The Open MPI for Exascale (OMPI-X) project was one of two in the Exascale Computing Project (ECP) focused on advancing the MPI ecosystem. The OMPI-X team worked with other MPI Forum members to champion several important features for inclusion in the MPI 4.0, 4.1, and upcoming 5.0 MPI standard versions, in support of the needs of exascale applications and systems. The team also worked with the larger Open MPI community to bring implementations of these new features and other enhancements into Open MPI, one of the leading open-source implementations of the MPI interface. This paper describes the motivation for the work of the OMPI-X project in the context of exascale computing needs, the nature of the resulting new capabilities in the MPI standard, and how they were implemented in the Open MPI library. Features include improved support for “MPI + X” programming models through partitioned communications and support for user-level threading, sessions, fault tolerance through the user-level fault mitigation (ULFM) and Reinit models, and other features. We also discuss enhancements to Open MPI providing improved performance and scalability for existing features, such as collective operations, one-sided operations, support for the Slingshot-11 interconnect of the initial exascale systems, and how the OMPI-X team worked to improve quality assurance for the Open MPI library, particularly on platforms of interest to the Department of Energy community.
Integrated modeling of plasma-surface interactions provides a comprehensive and self-consistent description of the system, moving the field closer to developing predictive and design capabilities for plasma facing components. One such workflow, including descriptions for the scrape-off-layer plasma, ion-surface interactions and the sub-surface evolution, was previously used to address steady-state scenarios and has recently been extended to incorporate time-dependence and two-way information flow. The new model can address dynamic recycling in transient scenarios, such as the application presented in this paper: the evolution of W samples pre-damaged by helium and exposed to ELMy H-mode plasmas in the DIII-D DiMES. A first set of simulations explored the effect of ELM frequency. This study was discussed in detail in this conference’s proceedings and is summarized here. The 2nd set of simulations, which is the focus of this paper, explores the effect of code-coupling frequency. These simulations include initial SOLPS solutions converged to the inter-ELM state, ion impact energy ( E in ) and angles ( A in ) calculated by hPIC2, and an improved heat transfer description in Xolotl. The model predicts increases in particle fluxes and decreases in heat fluxes by 10%–20% with the coupling time-step. Compared with the first set of simulations, the less shallow impact angle leads to smaller reflection rates and significant D implantation. The higher fraction of implanted flux (and deeper), in particular during ELMs, increases the accumulated D content in the W near-surface region. Future expansion of the workflow includes coupling to hPIC2 and GITR to ensure accurate descriptions of E in and A in , and W impurity transport.
The Cray HPE Slingshot 11 network is used on the new exascale systems arriving at the U.S. Department of Energy (DoE) laboratories (e.g., Frontier, Aurora, Perlmutter). As such, the support of this network is an important capability to meet the needs of exascale applications. This article highlights recent work to develop supporting infrastructure to enable Open MPI to efficiently support these new platforms. A key component of this effort involves development of a new Open Fabrics Interface (OFI) provider, LinkX. We discuss the design and development of enhancements that take advantage of the new Slingshot 11 network and AMD GPUs. We include performance data from tests on the Frontier supercomputer using synthetic communication benchmarks, and the vendor provided MPI as a baseline for comparison. The tests demonstrate full functionality of Open MPI on the system and initial results show favorable performance when compared to the highly tuned vendor implementation.
The development of scientific software—a cornerstone of long-term collaboration and scientific progress—parallels the development of other types of software but still poses distinct challenges, especially in high-performance computing. Although web searches yield numerous resources on software engineering, there is still a scarcity specifically for scientific software development. This article introduces the Better Scientific Software site (https://bssw.io), a platform that hosts a community of researchers, developers, and practitioners who share their experiences and insights on scientific software development. Since 2017, this collaborative hub has gained traction within the scientific computing community, attracting a growing number of readers and contributors eager to share ideas and elevate their software development practices. In sharing the BSSw.io site’s story, we hope to encourage further growth of the BSSw.io community through both readership and contributors, with a long-term goal of fostering culture change by increasing emphasis on best practices in scientific software.
Computational and data-enabled science and engineering are revolutionizing advances throughout science and society, at all scales of computing. For example, teams in the U.S. DOE Exascale Computing Project have been tackling new frontiers in modeling, simulation, and analysis by exploiting unprecedented exascale computing capabilities-building an advanced software ecosystem that supports next-generation applications and addresses disruptive changes in computer architectures. However, concerns are growing about the productivity of the developers of scientific software, its sustainability, and the trustworthiness of the results that it produces. Members of the IDEAS project serve as catalysts to address these challenges through fostering software communities, incubating and curating methodologies and resources, and disseminating knowledge to advance developer productivity and software sustainability. This paper discusses how these synergistic activities are advancing scientific discovery-mitigating technical risks by building a firmer foundation for reproducible, sustainable science at all scales of computing, from laptops to clusters to exascale and beyond.
As new compute systems are developed, there is still a need to compile and execute codes authored in Fortran on these leading edge systems. In order to achieve this, development of compilers that support the latest hardware is continuously under development. Though the specification of Fortran is extensive, it is helpful to compiler authors to be able to prioritize the development of key features in order to get certain codes deemed important, e.g., applications of interest to leadership computing facilities, executable on leading edge compute systems. Identifying key features though is largely done through querying software experts or users of the Fortran applications of interest, who then manually report what features are and are not present. This exercise can both time consuming and error prone. To automate this process, we present a compiler plugin to Flang, the Fortran frontend for LLVM. This plugin is a tool that operates on the parse tree representation generated by Flang and detects key features based on walking parse tree nodes that correspond to features of interest. We show the result of our tool on four applications, three of which were manually profiled by software experts. We show the discrepancies between our tool and the manual characterization of the three applications, as well as generate a characterization for an application not yet profiled. We intend to open-source our tool in order to invite the community to benefit from the tool and make contributions for other features.
for predicting reduction potentials and and calculated H adsorption energy as activity descriptor; and LANL Institutional Computing (IC) resources are essential for successful project execution given the size of structures involved in modeling molecule-support interaction.
Several configurations for the core and pedestal plasma are examined for a predefined tokamak design by implementing multiple heating/current drive (H/CD) sources to achieve an optimum configuration of high fusion power in a noninductive operation while maintaining an ideally magnetohydrodynamic (MHD) stable core plasma using the IPS-FASTRAN framework. IPS-FASTRAN is a component-based lightweight coupled simulation framework that is used to simulate magnetically confined plasma by integrating a set of high-fidelity codes to construct the plasma equilibrium (EFIT, TOQ, and CHEASE), calculate the turbulent heat and particle transport fluxes (TGLF), model various H/CD systems (TORIC, TORAY, GENRAY, and NUBEAM), model the pedestal pressure and width (EPED), and estimate the ideal MHD stability (DCON). The TGLF core transport model and EPED pedestal model are used to self-consistently predict plasma profiles consistent with ideal MHD stability and H/CD (and bootstrap) current sources. In order to evaluate the achievable and sustainable plasma beta, varying configurations are produced ranging from the no-wall stability to with-wall stability regimes, simultaneously subject to the self-consistent TGLF, EPED, and H/CD source profile predictions that optimize configuration performance. The pedestal density, plasma current, and total injected power are scanned to explore their impact on the target plasma configuration, fusion power, and confinement quality. A set of fully noninductive scenarios are achieved by employing ion-cyclotron, neutral beam injection, helicon, and lower-hybrid H/CDs to provide a broad profile for the total current drive in the core region for a predefined tokamak design. These noninductive scenarios are characterized by high fusion gain (Q similar to 4) and power (P-fus similar to 600 MW), optimum confinement quality (H-98 similar to 1.1), and high bootstrap current fraction (f(BS) similar to 0.7) for Greenwald fraction below unity. The broad current profile configurations identified are stable to low-n kink modes either because the normalized pressure beta(N) is below the no-wall limit or a wall is present.
Increasingly powerful and affordable computing has revolutionized scientific and scholarly discovery across a broad range of fields. Computing relies on software, which has been rapidly growing in scope, diversity, and complexity. At the same time, the methods, processes, and tools used to produce and utilize this essential software are often ad hoc, and the study and improvement of them are often done without the benefit of direct funding or prioritization. Consequently, concerns are growing about the productivity of the developers and users of scientific software, its sustainability, and the trustworthiness of the results that it produces. Increased investment, especially in the characterization and improvement of how scientific software is developed and used, is important for sustaining and improving the impact of software as the scope and complexity of scientific efforts expand. Without this investment, we face the risk of diminishing returns on our software investments because the demands for increased functionality, usability, reliability, and more will not be sufficiently met. The US Department of Energy Office of Science (DOE/SC) is at the forefront of modern software-enabled scientific discovery across numerous areas of computational, experimental, and observational science, including major investments in national user facilities that support these activities. For many years, DOE/SC software investments have provided tremendous value to the scientific community. We want to continue and further improve the value of DOE/SC software efforts by using a scientific approach to understanding and improving how scientific software is developed and used. In December 2021, the DOE/SC Office of Advanced Scientific Computing Research (ASCR) convened a workshop on basic research needs for the Science of Scientific-Software Development and Use (SSSDU). Through keynote presentations, lightning talks, and breakout groups, which built on insights from 124 pre-workshop position papers, participants discussed the current practice of software development, maintenance, evolution, and use, and considered how the scientific method could be used to examine these practices and develop more evidence-based approaches to enhance the impact of software and computing on all areas of science. Workshop participants identified three priority research directions (PRDs) and three important crosscutting themes that center on the following overarching insight: Software has become an essential part of modern science, impacting discoveries, policy, and technological development. To maintain and improve confidence in science delivered via software, we must improve the processes and tools that help us create and use software, and this enhancement requires a deep understanding of the diverse array of teams and individuals doing the work.
As the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project.
In the search for a sustainable approach for software ecosystems that supports experimental and observational science (EOS) across Oak Ridge National Laboratory (ORNL), we conducted a survey to understand the current and future landscape of EOS software and data. This paper describes the survey design we used to identify significant areas of interest, gaps, and potential opportunities, followed by a discussion on the obtained responses. The survey formulates questions about project demographics, technical approach, and skills required for the present and the next five years. The study was conducted among 38 ORNL participants between June and July of 2021 and followed the required guidelines for human subjects training. We plan to use the collected information to help guide a vision for sustainable, community-based, and reusable scientific software ecosystems that need to adapt effectively to: (i) the evolving landscape of heterogeneous hardware in the next generation of instruments and computing (e.g. edge, distributed, accelerators), and (ii) data management requirements for data-driven science using artificial intelligence.
As the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project.
The Better Scientific Software Fellowship (BSSwF) was launched in 2018 to foster and promote practices, processes, and tools to improve developer productivity and software sustainability of scientific codes. The BSSwF’s vision is to grow the community with practitioners, leaders, mentors, and consultants to increase the visibility of scientific software. Over the last five years, many fellowship recipients and honorable mentions have identified as research software engineers (RSEs). Case studies from several of the program’s participants illustrate the diverse ways the BSSwF has benefited both the RSE and scientific communities. In an environment where the contributions of RSEs are too often undervalued, we believe that programs such as the BSSwF can help recognize and encourage community members to step outside of their regular commitments and expand on their work, collaborations, and ideas for a larger audience.
We present a data-driven strategy for effective construction of a surrogate model in high-dimensional parameter space for the ion energy-angle distribution (IEAD) output of hPIC simulations of plasma-surface interactions. The methodology is based on a bin-by-bin least-squares fitting of the IEAD in the parameter space. The fitting is performed in a transformed coordinate system to normalize the IEAD, and it employs sparse grids for sampling the parameter space to overcome sampling challenges in high dimensions. The surrogate model is significantly cheaper computationally than direct hPIC simulations yet maintains high fidelity to them, providing a fast emulator for hPIC simulations. Sensitivity analysis based on the surrogate model is utilized to characterize the dependence of the ion impact angle and energy moments on the physical parameters.1
Boyana Norris合作论文数Mathematics and Computer Science Division;Argonne National Laboratory10
Katherine Riley合作论文数Argonne National Laboratory7