To migrate virtual machines requires not only time but also energy. If pursuing a follow-the-renewables computing paradigm, migrating virtual machines between data centres on the basis of clean energy availability, understanding not only the time requirements of migration but also its energy consumption is key. Existing research has paid little attention to the energy required to perform virtual machine migrations however. In this paper we present the results of an experimental study focused on the power and energy consumption of various workloads throughout the migration process. We also discuss challenges encountered during the study restricting our ability to migrate certain types of workloads. Based on our analysis we draw conclusions about the types of workloads suitable for migration and the manner in which each can be most effectively migrated.
In this article we present a protocol which has been developed for the purposes of providing guidance for estimating the emission reductions that could result from the provision or sourcing of low or zero carbon information and communication technology (ICT) services. This is an increasingly important topic not only because ICT has growing environmental impacts, but also due the technical complexities which underlie the delivery of ICT as a service, especially in respect of the growing use of cloud computing and the provision of ICT services over the internet. The protocol can be used both for creating emission reductions for carbon trading, and the quantification and reporting of related low or zero carbon ICT initiatives within corporate sustainability reports. (C) 2012 Published by Elsevier Inc.
We describe the design and implementation of a an interactive application that is capable of visualizing large remote data sets over the Internet. The application is accessible using a standard Internet browser, allowing users the ability to run visualizations with no need to download any data or to install any additional software. The application has two parts: a graphical user interface client that runs inside the user's browser, and a server code that performs all CPU and data intensive tasks, including rendering.
The next generation of telescopes, such as the Square Kilometre Array (SKA), will generate orders of magnitude more data than previous instruments, far in excess of current storage and networking system handling abilities. To address this problem, we propose an architecture where data is distributed over several archive sites, each holding only a portion of the overall data, that provides efficient and transparent access to the archive as a whole. This paper describes that architecture in detail and the design and implementation of a prototype system,based on the Integrated Rule-Oriented Data System (iRODS) software.
Managing the growing volume of data being output by radio telescopes is a significant challenge faced by radio astronomers today. This challenge will only be further compounded with future telescopes such as the Square Kilometre Array (SKA), which will be the world's largest radio telescope when completed and produce data at unprecedented rates. This paper introduces the CyberSKA collaborative portal which is aimed at addressing the current and future needs of data-intensive radio astronomy. A wide variety of tools and services that have been developed and integrated with the CyberSKA portal, including a distributed data management system, a data access tool, remote visualization tools and a third party application interface are described. Current international usage of CyberSKA focusing on several different SKA Pathfinder survey projects and how they make use of the portal are also highlighted.
Job submission in high performance computing workloads exhibits a diurnal pattern similar to electrical prices. While high-priority jobs may need immediate access to resources, by altering the cluster scheduler to delay the execution of lower-priority jobs when power prices are high, significant cost savings can be achieved. Reduction of power demands by consumers such as data centres when energy availability is low, as signaled by high prices, can also help to simplify challenges faced in reducing the carbon footprint of the electrical grid. In this paper we discuss patterns in electrical pricing and also look at some challenges in integrating more volatile, but environmentally friendly renewable energy sources into the electrical grid. Simulation results are also presented showing that high-priority jobs can still receive rapid service while achieving 25-50% electricity cost savings for lower priority jobs.
Accurate estimation of the amount of memory required is an important step in discovering and selecting configurations for computational jobs in a grid environment. In order to mitigate the adverse impact of over and under estimation of memory requirements, job submitters need to carry out benchmarking in order to understand the memory usage behaviour of specific applications. An automated mechanism for learning the memory usage pattern can allow them to focus more on the results of the jobs. This paper presents mechanisms that automate the process of understanding the memory usage behaviour of jobs. The mechanism uses the provenance data from previously completed jobs. The proposed mechanism enables populating an information model describing the memory usage patterns of applications. The paper also presents techniques for making predictions on the memory usage while the development of the knowledge base is still being carried out. The utility of the mechanism is demonstrated by showing its increasing ability to predict the memory usage for incoming jobs. Also highlighted are the factors that effect the pace in which the system acquires knowledge about the memory usage behaviour.
Advances in radio and digital processing technologies are enabling the construction of radio telescopes that will be able to probe the sky to unprecedented depths at radio wavelengths. With the vast amounts of data that will be produced by such telescopes comes a greater need for a cyberinfrastructure framework to connect the communities of astronomers with the data, processing and visualization tools, and each other. This paper introduces CyberSKA, an on-line, collaborative portal that is aimed at addressing the cyberinfrastructure needs of future radio telescopes such as the Square Kilometer Array (SKA).
Spectral Network (SpecNet) began as a Working Group in 2003 with the goals of integrating remote sensing with biosphere-atmosphere carbon flux measurements and standardizing field optical sampling methods. SpecNet has evolved into an international network of collaborating sites and investigators, with a particular focus on matching optical sampling tools to the temporal and spatial scale of flux measurements and ecological sampling. Current emphasis within the SpecNet community is on greater automation of field optical sampling using simple cost-effective technologies, improving the light-use-efficiency (LUE) model of carbon dioxide flux, consideration of view and illumination angle to improve physiological retrievals, and incorporation of informatics and cyberinfrastructure solutions that address the increasing data dimensionality of cross-site and multiscale sampling. In this review, we summarize recent findings and current directions within the SpecNet community and provide recommendations for the larger remote sensing and flux communities. These recommendations include comparing the LUE model to other flux models driven by remote sensing, considering a wider array of biogenic trace gases in addition to carbon dioxide, adoption of standardized and automated field sensors and sampling protocols where possible, continued development of cyberinfrastructure tools to facilitate data comparison and integration, expanding the network itself so that a greater range of sites are covered by combined optical and flux measurements, and encouraging a broader communication between the flux and remote sensing communities.
Managing the execution of scientific applications in a heterogeneous grid computing environment can be a daunting task, particularly for long running jobs. Increasing fault tolerance by checkpointing and migrating jobs between resources requires expertise and time of the scientist. Automation of such tasks can allow the scientist to focus more on the scientific results and less on the technical details. In this paper a generic framework for managing and automating the execution of jobs is presented. It uses of a variety of information models describing systems, policies, and application details/requirements to make suitable decisions on where and how to run, checkpoint, migrate and reconfigure jobs as needed. To demonstrate the utility of the framework, it is used as part of a simulation study to assess the impact availability of application memory usage information has on meeting the QoS objectives of job submitters and on overall utilization of resources. The study shows that with greater availability of memory usage information, the execution management framework is able to better meet user objectives and improve utilization of resources, particularly
Scientific data continues to grow in volume making the tasks of managing, accessing and sharing such data more challenging. Providing data to scientists via scientific gateways or collaborative portals can aid scientists in achieving these tasks. This paper presents a general data management system that has been built on top of Elgg, an open source social networking platform. The tool enables scientists to upload, browse, view and share a wide variety of scientific data, as well as define and evolve meta data standards in a collaborative manner. The data management system is currently being used as part of GeoChronos, a scientific gateway for Earth observation scientists, for creating and sharing collections of spectral and satellite data.
To manage jobs in multi-institutional grid environments, an automation tool needs to know not only the characteristics of resources, but also whether a job’s credentials will be mapped to accounts on them. Credentials may be mapped to an existing dedicated or shared account on a resource, or a new account may be created. Existing information models provide little account policy information, even though the development of virtual organization and account management tools means that account policies may be increasingly dynamic. Without automation tools being able to understand account policies, projects are unable to take full advantage of modern virtual organization and account management systems. Using advertised account policies, automation tools could consider whether the account creation, access, expiry, and cleanup policies of a service provider make it a good candidate for running particular jobs. Additionally, account renewals could be managed automatically using information in an expiry policy model.
Accessing, running and sharing applications and data presents researchers with many challenges. Cloud computing and social networking technologies have the potential to simplify or eliminate many of these challenges. Cloud computing technologies can provide scientists with transparent and on-demand access to applications served over the Internet in a dynamic and scalable manner. Social networking technologies provide a means for easily sharing applications and data. In this paper we present an on-line/on-demand interactive application service. The service is built on a cloud computing infrastructure that dynamically provisions virtualized application servers based on user demand. An open source social networking platform is leveraged to establish a portal front end that enables applications and results to be easily shared between researchers. Furthermore, the service works with existing/legacy applications without requiring any modifications.
Online social networking has significantly increased in popularity over the past several years, with sites such as Facebook now boasting over 300 million members. Scientific gateways have much to gain by incorporating social networking functionality. There are a variety of approaches for integrating social networking functionality in scientific gateways. These include integration into existing general purpose social networks or developing custom social networks that can be hosted by a third party or in-house. In this paper we compare these different approaches and relate our experience in applying these approaches in the development of GeoChronos, a collaborative gateway for earth observation scientists.
Automating the execution of applications in grid computing environments is a complicated task due to the heterogeneity of computing resources, resource usage policies, and application requirements. Applications differ in memory usage, performance, scalability and storage usage. Having knowledge of this information can aid in matching jobs to resources and in selecting appropriate configuration parameters such as the number of processors to run on and memory requirements for those resources. This paper presents an application memory usage model that can be used to aid in selecting appropriate job configurations for different resources. The model can be used to represent how memory scales with the number of processors, the memory usage of different types of processes, and changes in memory usage during execution. It builds on a previously developed information model used for describing resources, resource usage policies and limited information on applications. An analysis of the memory usage model illustrating its use towards automating job execution in grid computing environments is also presented.
ldquoWeb 2.0rdquo and ldquocloud computingrdquo are revolutionizing the way IT infrastructure is accessed and managed. Web 2.0 technologies such as blogs, wikis and social networking platforms provide Internet users with easier mechanisms to produce Web content and to interact with each other. Cloud computing technologies are aimed at running applications as services over the Internet on a scalable infrastructure. In this paper we explore the advantages of using Web 2.0 and cloud computing technologies in an enterprise setting to provide employees with a comprehensive and transparent environment for utilizing applications. To demonstrate the effectiveness of this approach we have developed an environment that uses a social networking platform to provide access to a legacy application. The application is hosted on an internal cloud computing infrastructure that adapts dynamically to user demands. Initial feedback suggests this approach provides an improved user experience while simplifying management and increasing effective utilization of the underlying IT resources.
Grid computing environments are heterogeneous in terms of the types of computer systems, the applications being run, and the policies governing the use of systems. To support automation in such environments common models to describe systems, applications and scheduling policies are needed. Common models enable the interoperation and reuse of tools transparent to the underlying differences in the environment. This paper relates our experiences in using modelling to support automation in grid environments. We have developed RDF-based models for describing systems, applications and scheduling policy. A GT4-based grid environment consisting of computer systems from across Canada has been established with model information deployed and published via WS MDS. Several examples demonstrating the use of modelling to support automation in the established grid environment are presented.
In the past several years, online social networks (OSNs) such as Facebook and MySpace have become extremely popular with Internet users. Such sites are popular with users because they simplify both communication among "communities" and access to applications. Application developers are attracted to these sites also, as they are able to exploit "word-of-mouth" marketing, which these OSN sites have embodied into their user experience. A challenge for developers though is managing the application, as it is difficult to predict how successful the marketing will be. Our solution combines an OSN, Virtual Appliances, and a utility computing environment together. We demonstrate our solution using the Facebook portal (OSN), the Fire Dynamics Simulator (application), and a utility environment we built using tools such as Condor, Moab and Xen. The application is supported using Virtual Appliances, which interact with our flexible infrastructure to dynamically expand and contract based on user demand. Thus, we are able to make much more efficient use of the underlying physical infrastructure. We believe that our solution also has great potential for enterprise IT environments. Initial feedback suggests combining an OSN with our flexible infrastructure provides a much better user experience than the traditional, standalone use of the (legacy) application, and simplifies the management and increases the effective utilization of the underlying IT resources.