Reproducibility remains a critical challenge in high-performance computing (HPC), where increasing system complexity and evolving software stacks often impede the validation of published results. This work, conducted as part of the Student Cluster Competition (SCC) 2024 Reproducibility Challenge at SC24 in Atlanta, targets this issue by reproducing key findings from Data Flow Lifecycles for Optimizing Workflow Coordination. The target paper introduces Data Flow Lifecycles (DFL), a methodology for optimizing data-intensive scientific workflows. We replicate the original experiments on SCC24 competition hardware with AMD EPYC 9634 CPU, utilizing RAM-disk and SSD-based staging. Our results successfully reproduce the original qualitative trends and demonstrate a $2.0\times$ performance speedup with these optimizations. However, quantitative results diverge by 50–62.5%, primarily due to differences in cluster architecture and system capabilities. We analyze these sources of variation and discuss implications for reproducibility in portable, power-constrained HPC environments. Our study offers practical insights for improving transparency and reproducibility in data-intensive scientific computing.
The second annual NSF, OAC CSSI, CyberTraining and related programs PI meeting was held August 12 to 13 in Charlotte, NC, with participation from PIs or representatives of all major awards. Keynotes, panels, breakouts, and poster sessions allowed PIs to engage with each other, NSF staff, and invited experts. The 286 attendees represented 292 awards across CSSI, CyberTraining, OAC Core, CIP, SCIPE CDSE, and related programs, and presented over 250 posters. This report documents the meetings structure, findings, and recommendations, offering a snapshot of current community perspectives on cyberinfrastructure. A key takeaway is a vibrant, engaged community advancing science through CI. AI-driven research modalities complement established HPC and data centric tools. Workforce development efforts align well with the CSSI community.
DaCe is a framework for Python that claims to provide massive speedups with C-like speeds compared to already existing high-performance Python frameworks (e.g. Numba or Pythran). In this work, we take a closer look at reproducing the NPBench work. We use performance results to confirm that NPBench achieves higher performance than NumPy in a variety of benchmarks and provide reasons as to why DaCe is not truly as portable as it claims to be, but with a small adjustment it can run anywhere.
The needs of cyberinfrastructure (CI) Users are different from those of CI Contributors. Typically, much of the training in advanced CI addresses developer topics such as MPI, OpenMP, CUDA and application profiling, leaving a gap in training for these users. To remedy this situation, we developed a new program: COMPrehensive Learning for end-users to Effectively utilize CyberinfraStructure (COMPLECS). COMPLECS focuses exclusively on helping CI Users acquire the skills and knowledge they need to efficiently accomplish their compute- and data-intensive research, covering topics such as parallel computing concepts, data management, batch computing, cybersecurity, HPC hardware overview, and high throughput computing.
In recent years, Artificial Intelligence (AI) has reshaped various facets of our day-to-day lives. To evaluate both hardware and software deployment of Machine Learning (ML) applications, it is necessary to measure system-wide performance and accuracy. MLPerf, a Machine Learning Benchmark from MLCommons, provides a standard for measuring Machine Learning Performance on hardware. To our knowledge, there is little known public effort on porting and optimizing these inference benchmark models on AMD GPUs. This paper focuses on the effort to port MLPerf BERT, one of the MLPerf Inference Benchmarks, to run on AMD Instinct MI210 and MI250 Accelerators. We describe the challenges encountered, solutions applied, and the successful porting of the benchmark resulting in submission of the code to the MLPerf community codebase. This paper derives from unpublished work completed as part of the 2023 Student Cluster Competition (SCC23). We present the first unofficial results for running the MLPerf BERT Inference Benchmark using our optimization strategies on AMD accelerators, specifically the MI250 and MI210. Additionally, as a result of this effort, the CM Automation framework now supports AMD ROCm for PyTorch and ONNX Runtime platforms.
To improve the sharing and discovery of CyberTraining materials, the HPC-ED Pilot project team is building a platform for the community to better share and find training materials through a federated catalog. The platform, currently in early test mode, is focused on a flexible platform, informative metadata, and community participation. By creating a framework for identifying, sharing, and including content broadly, HPC-ED will: allow providers of training materials to reach new groups of learners; extend the breadth and depth of training materials; and enable local sites to add or extend local portals.
Today, machine learning is being deployed for use in practice across all fields of science, engineering, and medicine. Although there are many educational resources on machine learning, they focus on executing workflows at modest scale. Our training program, Cyberinfrastructure-Enabled Machine Learning (CIML), focuses on the core competencies for at-scale ML workflows, and integrates topics from high-performance computing, data management, advanced cyberinfrastructure, reproducible computing, calable machine learning, and deep learning. The CIML project hosts an annual workshop that brings together researchers and practitioners from all fields, where the training program focuses on teaching participants the basics of high-performance computing (HPC) and ML at scale. Adopting "Findable, Accessible, Interoperable, and Reusable" (FAIR) practices, all CIML training material is made freely available online via GitHub. We describe the CIML project and its training program, report on its impact to date, and discuss our future plans for the project.
The Artificial Intelligence (AI) institute for Intelligent Cyberinfrastructure with Computational Learning in the Environment (ICICLE) is funded by the NSF to build the next generation of Cyberinfrastructure to render AI more accessible to everyone and drive its further democratization in the larger society. We describe our efforts to develop Jupyter Notebooks and Python command line clients that would access these ICICLE resources and services using ICICLE authentication mechanisms. To connect our clients, we used Tapis, which is a framework that supports computational research to enable scientists to access, utilize, and manage multi-institution resources and services. We used Neo4j to organize data into a knowledge graph (KG). We then hosted the KG on a Tapis Pod, which offers persistent data storage with a template made specifically for Neo4j KGs. In order to demonstrate the capabilities of our software, we developed several clients: Jupyter notebooks authentication, Neural Networks (NN) notebook, and command line applications that provide a convenient frontend to the Tapis API. In addition, we developed a data processing notebook that can manipulate KGs on the Tapis servers, including creations of a KG, data upload and modification. In this report we present the software architecture, design and approach, the successfulness of our client software, and future work.
This document describes a two-day meeting held for the Principal Investigators (PIs) of NSF CyberTraining grants. The report covers invited talks, panels, and six breakout sessions. The meeting involved over 80 PIs and NSF program managers (PMs). The lessons recorded in detail in the report are a wealth of information that could help current and future PIs, as well as NSF PMs, understand the future directions suggested by the PI community. The meeting was held simultaneously with that of the PIs of the NSF Cyberinfrastructure for Sustained Scientific Innovation (CSSI) program. This co-location led to two joint sessions: one with NSF speakers and the other on broader impact. Further, the joint poster and refreshment sessions benefited from the interactions between CSSI and CyberTraining PIs.
We describe the design motivation, architecture, deployment, and early operations of Expanse, a 5 Petaflop, heterogenous HPC system that entered production as an NSF-funded resource in December 2020 and will be operated on behalf of the national community for five years. Expanse will serve a broad range of computational science and engineering through a combination of standard batch-oriented services, and by extending the system to the broader CI ecosystem through science gateways, public cloud integration, support for high throughput computing, and composable systems. Expanse was procured, deployed, and put into production entirely during the COVID-19 pandemic, adhering to stringent public health guidelines throughout. Nevertheless, the planned production date of October 1, 2020 slipped by only two months, thanks to thorough planning, a dedicated team of technical and administrative experts, collaborative vendor partnerships, and a commitment to getting an important national computing resource to the community at a time of great need.
A User Portal is being developed for NSF-funded Expanse supercomputer. The Expanse portal is based on the NSF-funded Open OnDemand HPC portal platform which has gained widespread adoption at HPC centers. The portal will provide a gateway for launching interactive applications such as MATLAB, RStudio, and an integrated web-based environment for file management and job submission. This paper discusses the early experience in deploying the portal and the customizations that were made to accommodate the requirements of the Expanse user community.
The General Curvilinear Coastal Ocean Model (GCCOM) is a 3D curvilinear, structured-mesh, non-hydrostatic, large-eddy simulation model that is capable of running oceanic simulations. GCCOM is an inherently computationally expensive model: it uses an elliptic solver for the dynamic pressure; meter-scale simulations requiring memory footprints on the order of 10 12 cells and terabytes of output data. As a solution for parallel optimization, the Fortran-interfaced Portable–Extensible Toolkit for Scientific Computation (PETSc) library was chosen as a framework to help reduce the complexity of managing the 3D geometry, to improve parallel algorithm design, and to provide a parallelized linear system solver and preconditioner. GCCOM discretizations are based on an Arakawa-C staggered grid, and PETSc DMDA (Data Management for Distributed Arrays) objects were used to provide communication and domain ownership management of the resultant multi-dimensional arrays, while the fully curvilinear Laplacian system for pressure is solved by the PETSc linear solver routines. In this paper, the framework design and architecture are described in detail, and results are presented that demonstrate the multiscale capabilities of the model and the parallel framework to 240 cores over domains of order 10 7 total cells per variable, and the correctness and performance of the multiphysics aspects of the model for a baseline experiment stratified seamount.
At SC16, the SCC teams participated in a new application area: the Reproducibility Challenge. In this paper we report on our efforts to reproduce results presented in a paper titled "A Parallel Connectivity Algorithm for de Bruijn Graphs in Metagenomic Applications," which shows that the parallel graph-based algorithm developed scales to over a thousand cores, and runs faster than traditional Breadth First Search algorithms. In general, using the smaller competition test data sets on over 128 processors, we were able to reproduce some, but not all, of the reported results: we were unable to run the D1 data set on 128 cores and 2GB/core memory; our results did show similar timing trends for the different algorithm variations; we were able to observe the trend of communication dominating the computation time; and the AP and AP_LB versions of our runs on smaller datasets only show a small time improvement in our graphs, which is similar but not exactly what was described within the paper. We believe that cluster architecture, required memory, network tuning, and number of processors available impacted our ability to exactly reproduce the results of the paper. (C) 2017 Elsevier B.V. All rights reserved.
This paper presents the integration of a data assimilation framework and a nonhydrostatic coastal ocean model for the study of turbulent processes and state variables. Interfacing our General Curvilinear Coastal Ocean Model (GCCOM) with NCAR's Data Assimilation Research Testbed (DART), enable the integration of very high-resolution observations into the system. These results included observation system simulation experiments (OSSEs) for test cases using a very steep seamount, by using different observation error variances. Our results demonstrated that the DART-GCCOM model can assimilate high-resolution observations using as few as 30 ensemble members.
The Regional Ocean Modeling System (ROMS) is a hydrostatic free-surface ocean model ideally suited to simulate mesoscale to basin-scale (10 km - 10000 km) ocean processes. The General Curvilinear Coastal Ocean Model (GCCOM) is a nonhy-drostatic large eddy simulation (LES) model designed specifically for high-resolution (meters) simulations. In this research, a hybrid model is developed that nests a fine-grid GCCOM model within a coarse-grid ROMS. The nested GCCOM-ROMS model is tested in an idealized flow over a seamount.
The SDSU online Chemical Equilibrium Services perform numerical heat transfer and fluid flow computations, using the Flame3D simulator, for thousands of researchers, educators, and students. The computation is broken down into a grid of 2D or 3D control volumes, each of which runs for a few seconds, has small memory requirements (100 Bytes), is independent of its neighbors, and is submitted individually to a Web Service. The embarrassingly parallel simulation requires several hours to compute a few thousand control volumes, for 10’s of thousands of iterations on a desktop. To improve the computational performance, a multi-task computing (MTC) approach was adopted. For this, a simple job distribution Web service framework (JODIS) was designed that distributes application workloads across hetergenous computing systems. JODIS has been demonstrated to run millions of Flame3D tasks simultaneously on a variety of resources and queuing systems. In this paper we report on the impact of JODIS on Flame3D computations, along with our experiences gained and challenges encountered when using heterogeneous computing environments, including the TeraGrid. Using JODIS, we have demonstrated a significant increase in the resolution of Flame3D (from 10 to more than 10 control volumes) and significant reduction in run times (by a factor of over 40 for a large test case of 128 processors and 10 tasks). In general, we conclude that the MTC approach can significantly improve Flame3D computational performance, but that changes need to be made to queuing/job submission systems in order to facilitate the rapid cycles needed for jobs similar to the Flame3D tasks.
The Unified Curvilinear Ocean Atmospheric Model (UCOAM) is a Large Eddie Simulation (LES) CFD model capable of running both ocean and atmospheric simulations. It is the only environmental model in existence today using a full, 3D curvilinear coordinate system, which results in increased accuracy and resolution. UCOAM is a petascale model: it is capable of resolving sub-km scale fluctuations requires large arrays (1010 elements, the curvilinear system requires large number of arrays (~100); communication occurs along all 3 axes, and full simulations will generate TBytes of data. Consequently, this model requires parallelization. To facilitate UCOAM computations and data management, we have developed a new parallel framework capable of distributing the computations across arbitrary 3D processor arrangements, manages the complexity of the staggered grid variables, and performs communications along all axes, including diagonal and tridiagonal neighbors. To facilitate computations, we have developed computational environment (CE) based on the Cyber infrastructure Web Application Framework (Cyber Web) which supports the development of web services and portals. In this paper we discuss the design and architecture of the parallel framework and supporting CE infrastructure, as well as challenges associated with parallelizing this novel model. We include the first initial parallel results for a small (105 nodes) 1 meter resolution seamount test case that shows scaling of the parallel model.
The General Curvilinear Ocean Model (GCOM) differs significantly from the traditional approach, where the use of Cartesian coordinates forces the model to simulate terrain as a series of steps. GCOM utilizes a full three-dimensional curvilinear transformation, which has been shown to have greater accuracy than similar models and to achieve results more efficiently. The GCOM model has been validated for several types of water bodies, different coastlines and bottom shapes, including the Alarcon Seamount, Southern California Coastal Region, the Valencia Lake in Venezuela, and more recently the Monterey Bay. In this paper, enhancements to the GCOM model and an overview of the computational environment (GCOM-CE) are presented. Model improvements include migration from F77 to F90; approach to a component design; and initial steps towards parallelization of the model. Through the use of the component design, new models are being incorporated including biogeochemical, pollution, and sediment transport. The computational environment is designed to allow various client interactions via secure Web applications (portal, Web services, and Web 2.0 gadgets). Features include building jobs, managing and interacting with long running jobs; managing input and output files; quick visualization of results; publishing of Web services to be used by other systems such as larger climate models. The CE is based mainly on Python tools including a grid-enabled Pylons Web application Framework for Web services, pyWSRF (python-Web Services-Resource Framework), pyGlobus based web services, SciPy, and Google code tools.
The combination of a heating element and a control unit therefor and methods of making and of operating the same are provided, the heating element normally being adapted to be operated by the continuous full wave pulses of a certain high voltage alternating current source and the control unit being operatively interconnected to the heating element for operatively interconnecting the heating element to a high voltage alternating current source, the control unit having an arrangement for operating the heating element with a certain repeating pattern of skipped full half-wave pulses of the source when the source has a higher voltage than the certain source.
Geoffrey Fox合作论文数Department of Physics, College of Arts and Sciences, Indiana University;Department of Intelligent Systems Engineering, Indiana University;Community Grid Laboratory, Indiana University;Digital Science Center of Pervasive Technology Institute;School of Engineering and Applied Science, University of Virginia7
G. Von Laszewski合作论文数Indiana University2