The Fast Fourier Transform (FFT) is required for chemistry, weather, defense, and signal processing for seismic exploration and radio astronomy. It is communication-bound, making supercomputers thousands of times slower at FFTs then at dense linear algebra. The key to accelerating FFTs is to minimize bits per datum without sacrificing accuracy. The 16-bit fixed point and IEEE float type lack sufficient accuracy for 1024- and 4096-point FFTs of data from analog-to-digital converters. We show that the 16-bit posit, with higher accuracy and larger dynamic range, can perform FFTs so accurately that a forward-inverse FFT restores the original signal perfectly. “Reversible” FFTs with posits are lossless, eliminating the need for 32-bit or higher precision. Similarly, 32-bit posit FFTs can replace 64-bit float FFTs for many HPC tasks. Speed, energy efficiency, and storage costs can thus be improved by 2 × for a broad range of HPC workloads.
Scientists from different communities, especially large-scale experimental facilities, have unique data and compute requirements that cannot be easily realised on a general-purpose supercomputing environment without significant modifications to either the domain requirements or supercomputing operational configurations. One of the crucial customisations required is the orchestration of on-demand and auto-scale workflows on a largely batch-driven supercomputing system. On-demand, auto-scale and resiliency are however features that are typically associated with cloud technologies. In order to empower a wide variety of users, delivery models such as X-as-a-Service have been introduced. The experimental facilities' workflows from PSI that will leverage on Data and Compute as a Service in the supercomputing ecosystem at CSCS will be shared. Three use cases were implemented to demonstrate the feasibility and benefits of applying a clouddriven approach to supercomputing ecosystems.
Large scale experimental facilities such as the Swiss Light Source and the free-electron X-ray laser SwissFEL at the Paul Scherrer Institute, and the particle accelerators and detectors at CERN are experiencing unprecedented data generation growth rates. Consequently, management, processing and storage requirements of data are increasing rapidly. Historically, online and on-demand processing of data generated by the instruments used to be tightly-coupled with a dedicated, domains-specific, site-local IT infrastructure. Cost and performance scaling of these facilities not only pose technical but also planning and scheduling challenges. Supercomputing ecosystems optimize cost and scaling for computing and storage resources but typically exploit a shared batch access model, which is optimized for high utilization of compute resources. In comparison, in public clouds, on-demand service delivery models address the concept of elasticity while maintaining isolation with performance trade-offs. Furthermore, these on-demand access models allow for different degrees of privileges to users for managing IT infrastructure services, in contrast with shared, bare-metal supercomputing ecosystems. This paper outlines an approach for enabling interactive, on-demand supercomputing for experimental data-driven workflows, which are characterised by a managed but bursty data and computing requirements. We present a delegated batch reservation model, controlled by the customer and provisioned by the supercomputing site, that allows scientists at the experimental facility to couple generation of data to the allocation of compute, data and network resources at the supercomputing centre. Scientists are then able to manage resources both at the experimental and supercomputing facilities interactively for managing their scientific workflows. Prototype implementation demonstrates that this rather simple co-designed extension to a supercomputing classic batch scheduling system with a controlled degree of privilege can be easily incorporated to the experimental facilities existing IT resource management and scheduling pipelines.
Forecasting severe hydro-meteorological events under time constraints requires reliable and robust models, which facilitate energy-aware allocations on high performance computing infrastructures. We present an innovative approach to quantify the performance of six heuristics in selecting optimal allocations to distributed HPC resources for an ensemble of meteorological forecasts for a flash-flood producing storm in Genoa (Liguria, Italy) in October 2014. The computing environments are expected to be dynamic and heterogeneous in nature with varying availability, performance and energy-to-solution. The results of the allocations are assessed and compared. The aim is to provide a robust, reliable and energy-aware resource allocation for ensembles of forecasts for time-critical decision support.
Urgent computing requires computations to commence in short order and complete within a stipulated deadline to support mitigation activities in preparation, response and recovery from an event that requires immediate attention. Missing an urgent deadline can lead to dire consequences where avoidable human and financial losses ensued. Timely allocation of resources to meet the deadline is crucial. Robustness is of great importance to ensure that small perturbations on the computing systems do not affect the makespan of the allocations so that the deadline can be met. This work focuses on developing a general mathematical makespan model for urgent computing to enable a robust allocation of ensemble forecasts while minimising the makespan. Three patterns of different resource allocation will be investigated to illustrate the model. The result will aid in satisfying the most crucial requirement, the time criterion, of urgent computing.
Numerical simulations of urgent events, e.g. tsunamis and storm, must be completed within a stipulated deadline. The simulation results are needed by relevant authorities in making timely educated decisions to mitigate financial losses, manage affected areas and reduce casualties. The existing definition of urgent computing is too usage context specific and thus restricts the identification of urgent use cases and the general application of urgent computing. Related paradigms like real-time computing, and crisis and disaster management computing further complicate matters. We thus aim to extend and refine the existing definition to provide a comprehensive general version to clarify the differences. The requirements of urgent computing, the urgent system and its characteristics, deadline and cost will be elaborated. This general definition will aid in the identification of the unique challenges of urgent computing and in turn simulates innovative multi-disciplinary solutions to address these challenges.
The VERCE project has pioneered an e-Infrastructure to support researchers using established simulation codes on high-performance computers in conjunction with multiple sources of observational data. This is accessed and organised via the VERCE science gateway that makes it convenient for seismologists to use these resources from any location via the Internet. Their data handling is made flexible and scalable by two Python libraries, ObsPy and dispel4py and by data services delivered by ORFEUS and EUDAT. Provenance driven tools enable rapid exploration of results and of the relationships between data, which accelerates understanding and method improvement. These powerful facilities are integrated and draw on many other e-Infrastructures. This paper presents the motivation for building such systems, it reviews how solid-Earth scientists can make significant research progress using them and explains the architecture and mechanisms that make their construction and operation achievable. We conclude with a summary of the achievements to date and identify the crucial steps needed to extend the capabilities for seismologists, for solid-Earth scientists and for similar disciplines. Keywords-Science Gateway, HPC, Data-Intensive, Data Science, Metadata and Storage, solid-Earth Sciences, Virtual Research Environment, e-Infrastructure.
The VERCE project has pioneered an e-Infrastructure to support researchers using established simulation codes on high-performance computers in conjunction with multiple sources of observational data. This is accessed and organised via the VERCE science gateway that makes it convenient for seismologists to use these resources from any location via the Internet. Their data handling is made flexible and scalable by two Python libraries, ObsPy and dispel4py and by data services delivered by ORFEUS and EUDAT. Provenance driven tools enable rapid exploration of results and of the relationships between data, which accelerates understanding and method improvement. These powerful facilities are integrated and draw on many other e-Infrastructures. This paper presents the motivation for building such systems, it reviews how solid-Earth scientists can make significant research progress using them and explains the architecture and mechanisms that make their construction and operation achievable. We conclude with a summary of the achievements to date and identify the crucial steps needed to extend the capabilities for seismologists, for solid-Earth scientists and for similar disciplines.
Urgent computing requires computations to commence in short order and complete within a stipulated deadline so as to support mitigation activities in preparation, response and recovery from an event that requires immediate attention. As such, acquiring computation resources swiftly is crucial. Preemptive scheduling, terminating an existing job(s) to make way for an urgent job, is one of the most common approach considered. However, public resource providers are typically faced with policy restrictions that forbid them from allowing preemption. The interruption of existing jobs is believed to have a significant consequence, i.e. cost, to the users and resource providers. This case study on a public HPC resource, SuperMUC, hosted at Leibniz Supercomputing Centre aims to study the cost of preemption. Two cost models, least cost and least disruptive, will be used. With this, we want to demonstrate that the cost of preemption is in fact much lower in comparison to the loss mitigation that can be achieved by allowing an urgent computation. The ultimate aim is to provide evidence to convince policy makers on the feasibility and benefits of supporting urgent computing on public resources.
In spite of the advances in today's technologies, most disasters are hard to predict and prevent. Disaster management to reduce damages and losses is particularly important. Computations to predict the effect of an impending or prevailing disaster, i.e. urgent computing, can support civil protection services to make informed decisions that can improve the efficiency of mitigation activities and reduce the causalities and losses. Accessibility to underlying distributed resource sets should be possible from ubiquitous end user devices, especially in the chaotic environment that entails a disaster. The inherent unpredictability of disasters can render any best made plans to prepare resources in advance futile. It is thus essential to acquire the ability to swiftly organise a resource for urgent computing, commonly while facing uncertainty in computation requirements and dynamism of computing environments on heterogeneous distributed resources. These challenges have to be conquered for urgent computing to be successfully realised. A task-based ubiquitous approach established on top of a three layers architecture is recommended. This architecture allows a separation of concerns between the accessibility needs from the resource/environment and use case specific conditions. Heterogeneity of distributed resources and environments is managed by the task-based setup with a set of subtask functions. Four urgent managers are also introduced to administer the urgent computing requirements. The ultimate aim is to provide an urgent computer system framework for supporting time-critical computations for a wide array of use cases on heterogeneous distributed resources.
Advanced application environments for seismic analysis help geoscientists to execute complex simulations to predict the behaviour of a geophysical system and potential surface observations. At the same time data collected from seismic stations must be processed comparing recorded signals with predictions. The EU-funded project VERCE ( http://verce.eu/ ) aims to enable specific seismological use-cases and, on the basis of requirements elicited from the seismology community, provide a service-oriented infrastructure to deal with such challenges. In this paper we present VERCE’s architecture, in particular relating to forward and inverse modelling of Earth models and how the, largely file-based, HPC model can be combined with data streaming operations to enhance the scalability of experiments. We posit that the integration of services and HPC resources in an open, collaborative environment is an essential medium for the advancement of sciences of critical importance, such as seismology.
Urgent computing enables responsible authorities to make educated decisions by supporting the computations of simulated predictions of time critical events. Unfortunately, most domains of science cannot afford dedicated resources for their urgent computing problems. As a solution, exploiting existing e-Infrastructures is invaluable for many problems if the wide array of available resources in today's e-Infrastructures can be utilised. In this paper, we focus on rarely occurring events that are best suited for urgent computations on existing HPC, Grid and Cloud e-Infrastructures. Since e-Infrastructures are meant to serve more than just one community of users, they have inherent characteristics that have to be modified or adapted in order to enable them effectively for urgent computing. We hope to demonstrate that there are many existing and on-going developments that can be leveraged to prepare existing e-Infrastructures for urgent computing.