This paper focuses on recent extensions to the CyberSKA system to enable multiple CyberSKA portal systems to have a common view of data distributed between the different sites. The extensions also enable remote visualisations to be started at the data centre holding a file. Fast replication of data between sites can be performed if latency is an issue when remotely visualising from the file's original location or to make the file downloadable or accessible from multiple locations. Other enhancements include the integration of Globus Online for importing files and improved security.
The African Data Intensive Research Cloud project aims to establish resources to support data intensive radio astronomy research among collaborating partners in South Africa and African Square Kilometre Array telescope partner countries. Infrastructure as a Service cloud instances using the OpenStack middleware have been deployed at three sites in South Africa as a proof of concept for the future larger scale deployments. By supporting common software, knowledge can be shared amongst the partners leading to faster resolution of problems, thus providing better support to the research community. The project will deploy a mix of scales of systems, with users bursting to the larger facilities as needed. Software to support data distribution, analysis and visualisation of astronomy data through scientific gateways is being deployed. Providing training of computing support teams, astronomers and data scientists at the partner organisations is key to the success of the project.
The Square Kilometer Array (SKA) telescope, which is planned for construction starting in 2019, is currently in its pre-construction design phase. It is planned that high level science data products will be created by the SKA Science Data Processor (SDP) computing systems located in both Perth and Cape Town with these then stored and made available for the research community to use. Due to the nature of the core SDP systems including base computational load, energy consumption constraints and production 24-7 operational requirements, user analysis of generated data products will have to occur on other systems. The SDP DELIV element work package is developing requirements and an architecture for how data could be accessed by scientists and how data should be moved efficiently to globally distributed resources for analysis. This paper describes the current status of work by the SDP DELIV team.
Ancillary services are the mechanisms power grids use to address short-term variability in supply and demand as well as the impact of power plant or transmission line failures. Organizations providing such services can earn revenue or at least reduce their energy costs. This paper explores options for large data centres to reduce costs in this way. Simulation results are presented for a system that models the processing of a workload and the resulting energy use, focusing on the impact of providing specific types of ancillary services. Trace data recording the workload from three supercomputing facilities along with pricing information from a US-based electrical grid are used. Results presented show energy costs reduced by up to 12% with only a small impact on the quality of service provided to users of the data centre. Further reductions in energy costs are shown for data centres willing to cede more control over short-term energy consumption. The potential of ancillary services markets are compared against following-the-renewables approaches, using wind farm data to argue why providing ancillary services may be more effective than optimizing operations based on the production of local renewables. (C) 2013 Elsevier Inc. All rights reserved.
To migrate virtual machines requires not only time but also energy. If pursuing a follow-the-renewables computing paradigm, migrating virtual machines between data centres on the basis of clean energy availability, understanding not only the time requirements of migration but also its energy consumption is key. Existing research has paid little attention to the energy required to perform virtual machine migrations however. In this paper we present the results of an experimental study focused on the power and energy consumption of various workloads throughout the migration process. We also discuss challenges encountered during the study restricting our ability to migrate certain types of workloads. Based on our analysis we draw conclusions about the types of workloads suitable for migration and the manner in which each can be most effectively migrated.
In this article we present a protocol which has been developed for the purposes of providing guidance for estimating the emission reductions that could result from the provision or sourcing of low or zero carbon information and communication technology (ICT) services. This is an increasingly important topic not only because ICT has growing environmental impacts, but also due the technical complexities which underlie the delivery of ICT as a service, especially in respect of the growing use of cloud computing and the provision of ICT services over the internet. The protocol can be used both for creating emission reductions for carbon trading, and the quantification and reporting of related low or zero carbon ICT initiatives within corporate sustainability reports. (C) 2012 Published by Elsevier Inc.
Over the last decade, the grid has emerged as a paradigm of distributed and collaborative computing focusing on the sharing of computational and storage resources spanning across geographical and organizational domains. Greater access to high-end computational facilities provides researchers from a broad spectrum of domains an inexpensive option of carrying out sophisticated computational experiments. However, the inherent dynamics and heterogeneity of grid environments make the execution of resource and compute intensive applications a challenging task. Increasing fault tolerance by checkpointing and migrating jobs between resources requires significant expertise and intervention from users. Automation of such tasks can allow them to focus more on the scientific results and less on the technical details. This thesis addresses the issues associated with management of execution of long running applications in grid environments. It presents a generic framework for automating execution of such applications. The framework is driven by a set of information models that capture knowledge about the resources and the applications. Crucial to the functioning of the framework is information on two application characteristics: the configurability, and the memory usage behaviour. Separate models are presented to encode knowledge of both of these characteristics. Use of a common representation of knowledge abstracts the heterogeneity of both the resources and the applications and makes the framework functional without the need to be tailored to any specific application. Two important issues that need to be considered in managing job execution are the amount of memory required by the job and the wait time the job may experience on a specific resource. The framework presented in this thesis is equipped with mechanisms to address both of these issues. It is able to make estimations about the wait time for jobs with different resource requirements. A learning system has been designed as part of the framework to characterize the memory usage behaviour of application instances. The system facilitates execution management operations by providing accurate estimation of job's memory usage.
Welcome to this special issue of SIMULATION devoted to the principles of advanced and distributed simulation (PADS). The objective of this special issue is to provide the reader with recent results on hot topics in the PADS research area. All of the papers appearing in this special issue have been selected using a peer review process and are revised/extended versions of papers that recently appeared in the workshop on principles of advanced and distributed simulation. The set of topics addressed by the selected papers is wide with results ranging from simulation support for specific challenging application problems, to general-purpose solutions. Also, different perspectives are provided, with some of the papers presenting a methodological perspective and others being more oriented to architectural and implementation aspects. The paper ‘Discrete event modeling and massively parallel execution of epidemic outbreak phenomena’, by Kalyan S Perumalla and Sudip K Seal, addresses the issue of efficiently simulating complex phenomena such as epidemiological outbreaks. The solution presented by the authors includes a reaction–diffusion simulation model for epidemic propagation and an implementation that utilizes reverse computation techniques that enables efficient optimistic parallel discrete event simulation. Results demonstrating the scalability of the system on thousands of processor cores are provided for an IBM Blue Gene/P system and a Cray XT5 system. The paper ‘Ad hoc distributed simulation methodology for open queuing networks’, by Ya-Lin Huang, Christos Alexopoulos, Michael Hunter and Richard M Fujimoto, provides innovative contributions in the context of the recently developed ad-hoc distributed simulation methodology. This is based on the integration (embedding) of a distributed simulation system into an operational system (such as a sensor network) in order to support predictions of the future state trajectory for the system. In particular, the paper shows how the ad-hoc methodology can be applied in all the scenarios where the embedded simulation system can be modeled as an open queuing network. Experimental results reported in the paper also show how the approximations used in the approach described (e.g. in relation to the flow of units along the links between different simulators within the ad-hoc distributed simulation system) enable predictions that are close to those achievable via a sequential simulation. The paper ‘Multicore acceleration of Discrete Event System Specification systems’, by Qi Liu and Gabriel Wainer, addresses the important issue of running simulations efficiently on heterogeneous multi-core platforms. The authors discuss how these types of architectures require reconsidering the design and implementation of high-performance simulation kernels. The paper also provides a specific solution tailored to the IBM Cell processor, based on the Discrete Event System Specification methodology. A discussion on the applicability of the proposed approach in the context of different multi-core architectures is provided, which enlarges the applicability of the approach presented. The paper ‘Towards flexible exascale stream processing system simulation’, by Alfred J Park, Cheng-Hong Li, Ravi Nair, Nobuyuki Ohba, Uzi Shvadron, Ayal Zaks and Eugen Schenfeld, is targeted at simulating stream processing applications. Parallel/distributed simulation methodologies have been used as the reference by the authors in order to build a simulation environment called Flow, which allows for parallelization of the simulation model execution with minimum overhead for scenarios modeling acyclic stream application graphs. The system has been evaluated on a cluster computing platform with 512 processor cores. Finally, the paper ’A framework for simulation-based optimization of business process models’, by Farzad Kamrani, Rassul Ayani and Farshad Moradi, exploits simulation to address the assignment problem. The innovation from this paper is that the tasks to be assigned to agents are part of a business process model, hence they are characterized by interdependencies related to proper rules and constraints. The authors consider different business process scenarios, where the assignment is either independent or dependent, with the meaning that the business process flow is not, or is, affected by the actual assignment. The solution presented has applications in multiple real-world contexts. Also, the results of an experimental study on medium-size business processes show how the cost of simulating the processes is low enough to be performed on entry-level desktop computers. We thank all of the authors who have contributed to this special issue. We also thank the editor in chief Professor Levent Yilmaz and the special issues editor Professor Gabriel A Wainer for giving us the opportunity to edit the special issue. We hope all of you will enjoy the proposed contents.
Ancillary services are the mechanisms power grids use to address short-term variability in supply and demand as well as the impact of power plant or transmission line failures. Organizations providing such services can earn revenue, or at least reduce their energy costs. This paper explores options for large data centres to reduce costs in this way. Simulation results are presented for a system that models the processing of a workload and the resulting energy use, focusing on the impact of providing specific types of ancillary services. Trace data recording the workload from three supercomputing facilities along with pricing information from a US-based electrical grid are used. Results presented show energy costs reduced by up to 12% with only a small impact on the quality of service provided to users of the data centre. Further reductions in energy costs are shown for data centres willing to cede more control over short-term energy consumption.
The next generation of telescopes, such as the Square Kilometre Array (SKA), will generate orders of magnitude more data than previous instruments, far in excess of current storage and networking system handling abilities. To address this problem, we propose an architecture where data is distributed over several archive sites, each holding only a portion of the overall data, that provides efficient and transparent access to the archive as a whole. This paper describes that architecture in detail and the design and implementation of a prototype system,based on the Integrated Rule-Oriented Data System (iRODS) software.
Job submission in high performance computing workloads exhibits a diurnal pattern similar to electrical prices. While high-priority jobs may need immediate access to resources, by altering the cluster scheduler to delay the execution of lower-priority jobs when power prices are high, significant cost savings can be achieved. Reduction of power demands by consumers such as data centres when energy availability is low, as signaled by high prices, can also help to simplify challenges faced in reducing the carbon footprint of the electrical grid. In this paper we discuss patterns in electrical pricing and also look at some challenges in integrating more volatile, but environmentally friendly renewable energy sources into the electrical grid. Simulation results are also presented showing that high-priority jobs can still receive rapid service while achieving 25-50% electricity cost savings for lower priority jobs.
Data centres containing high-performance computing (HPC) clusters may be able to coordinate with the operation of wind farms for mutual benefit. Large data centres consume megawatts of power, typically accounting for a majority of life cycle carbon emissions and a significant portion of the total cost of ownership.We ran simulations to explore the potential for data centres to adapt to dynamic electrical prices, variation in carbon intensity within an electrical grid, or the availability of local renewables. Using workloads from the Parallel Workloads Archive alongside real-world pricing data, we demonstrate potential savings on the cost of electricity ranging typically between 10-50%. Adaptation to the variation in the electrical grid carbon intensity was not as successful, but adaptation to the availability of local renewables showed potential to significantly increase their use. In one example the fraction of power obtained from a local wind installation increased by 10-80%.
Accurate estimation of the amount of memory required is an important step in discovering and selecting configurations for computational jobs in a grid environment. In order to mitigate the adverse impact of over and under estimation of memory requirements, job submitters need to carry out benchmarking in order to understand the memory usage behaviour of specific applications. An automated mechanism for learning the memory usage pattern can allow them to focus more on the results of the jobs. This paper presents mechanisms that automate the process of understanding the memory usage behaviour of jobs. The mechanism uses the provenance data from previously completed jobs. The proposed mechanism enables populating an information model describing the memory usage patterns of applications. The paper also presents techniques for making predictions on the memory usage while the development of the knowledge base is still being carried out. The utility of the mechanism is demonstrated by showing its increasing ability to predict the memory usage for incoming jobs. Also highlighted are the factors that effect the pace in which the system acquires knowledge about the memory usage behaviour.
Managing the execution of scientific applications in a heterogeneous grid computing environment can be a daunting task, particularly for long running jobs. Increasing fault tolerance by checkpointing and migrating jobs between resources requires expertise and time of the scientist. Automation of such tasks can allow the scientist to focus more on the scientific results and less on the technical details. In this paper a generic framework for managing and automating the execution of jobs is presented. It uses of a variety of information models describing systems, policies, and application details/requirements to make suitable decisions on where and how to run, checkpoint, migrate and reconfigure jobs as needed. To demonstrate the utility of the framework, it is used as part of a simulation study to assess the impact availability of application memory usage information has on meeting the QoS objectives of job submitters and on overall utilization of resources. The study shows that with greater availability of memory usage information, the execution management framework is able to better meet user objectives and improve utilization of resources, particularly
Scientific data continues to grow in volume making the tasks of managing, accessing and sharing such data more challenging. Providing data to scientists via scientific gateways or collaborative portals can aid scientists in achieving these tasks. This paper presents a general data management system that has been built on top of Elgg, an open source social networking platform. The tool enables scientists to upload, browse, view and share a wide variety of scientific data, as well as define and evolve meta data standards in a collaborative manner. The data management system is currently being used as part of GeoChronos, a scientific gateway for Earth observation scientists, for creating and sharing collections of spectral and satellite data.
To manage jobs in multi-institutional grid environments, an automation tool needs to know not only the characteristics of resources, but also whether a job’s credentials will be mapped to accounts on them. Credentials may be mapped to an existing dedicated or shared account on a resource, or a new account may be created. Existing information models provide little account policy information, even though the development of virtual organization and account management tools means that account policies may be increasingly dynamic. Without automation tools being able to understand account policies, projects are unable to take full advantage of modern virtual organization and account management systems. Using advertised account policies, automation tools could consider whether the account creation, access, expiry, and cleanup policies of a service provider make it a good candidate for running particular jobs. Additionally, account renewals could be managed automatically using information in an expiry policy model.
Accessing, running and sharing applications and data presents researchers with many challenges. Cloud computing and social networking technologies have the potential to simplify or eliminate many of these challenges. Cloud computing technologies can provide scientists with transparent and on-demand access to applications served over the Internet in a dynamic and scalable manner. Social networking technologies provide a means for easily sharing applications and data. In this paper we present an on-line/on-demand interactive application service. The service is built on a cloud computing infrastructure that dynamically provisions virtualized application servers based on user demand. An open source social networking platform is leveraged to establish a portal front end that enables applications and results to be easily shared between researchers. Furthermore, the service works with existing/legacy applications without requiring any modifications.
Francesco Quaglia合作论文数Universita di Roma "La Sapienza"3
Carl Tropper合作论文数McGill University2