In this paper, we present an experimental study of deterministic non-preemptive multiple workflow scheduling strategies on a Grid. We distinguish twenty five strategies depending on the type and amount of information they require. We analyze scheduling strategies that consist of two and four stages: labeling, adaptive allocation, prioritization, and parallel machine scheduling. We apply these strategies in the context of executing the Cybershake, Epigenomics, Genome, Inspiral, LIGO, Montage, and SIPHT workflows applications. In order to provide performance comparison, we performed a joint analysis considering three metrics. A case study is given and corresponding results indicate that well known DAG scheduling algorithms designed for single DAG and single machine settings are not well suited for Grid scheduling scenarios, where user run time estimates are available. We show that the proposed new strategies outperform other strategies in terms of approximation factor, mean critical path waiting time, and critical path slowdown. The robustness of these strategies is also discussed.
Scientists demand easy-to-use, scalable and flexible infrastructures for sharing, managing and processing their data spread over multiple resources accessible via different technologies and interfaces. In our previous work, we developed the conceptual framework VISPA for addressing these requirements. This paper provides a case study assessing the integrated Rule-Oriented Data System (iRODS) for implementing the key concepts of VISPA. We found that iRODS is already well suited for handling metadata and sharing data. Although it does not directly support provenance information of data and the temporal provisioning of data, basic forms of these capabilities may be provided through its customization mechanisms, ie rules and micro-services.
This paper discusses the problem of planning resource outsourcing and local configurations for infrastructure services that are subject to Service Level Agreements (SLA). The objective of our approach is to minimize implementation and outsourcing costs for reasons of competitiveness, while respecting business policies for profit and risk. We implement a greedy algorithm for outsourcing, using cost and subcontractor reputation as selection criteria; and local resource configurations as a constraint satisfaction problem for acceptable profit and failure risks. Thus, it becomes possible to provide educated price quotes to customers and establish safe electronic contracts automatically. Discarding either local resource provisioning, or outsourcing, models efficiently the specialized cases of infrastructure resellers and isolated infrastructure providers respectively.
Cloud computing effectively implements the vision of utility computing by employing a pay-as-you-go cost model and allowing on-demand (re-)leasing of IT resources. Small or medium-sized Infrastructure-as-a-Service providers, however, find it challenging to satisfy all requests immediately due to their limited resource capacity. In that situation, both providers and customers may benefit greatly from advanced reservation of virtual resources, i.e. virtual machines. In our work, we assume SLA-based resource requests and introduce an advanced reservation methodology during SLA negotiation by using computational geometry. Thereby, we are able to verify, record and manage the infrastructure resources efficiently. Based on that model, service providers can easily verify the available capacity for satisfying the customer's Quality-of-Service requirements. Furthermore, we introduce flexible alternative counter-offers, when the service provider lacks resources. Therefore, our mechanism increases the utilization of the resources and attempts to satisfy as many customers as possible.
Workflow scheduling in Grids becomes an important area as it allows users to process large scale problems in an atomic way. However, validating the performance of workflow scheduling strategies in real production environment cannot be feasibly carried out. The complexity of production systems, dynamicity of Grid execution environments, and the difficulty to reproduce experiments, make workflow scheduling production systems a complex research environment. Instead, this work is based on a trace driven simulator. This work presents workflow scheduling support as an extension to the Teikoku Grid Scheduling Framework (tGSF). tGSF was developed as a response for a standard compliant and trace based Grid scheduling simulation environment. Workflow scheduling is provided via a second layer of Grid scheduler, extensible to new workflow and parallel scheduling strategies. This work also includes a usage case scenario, which illustrates how this extension can be used for quantitative experimental study.
Executing applications in the Grid often requires access to multiple geographically distributed resources. In a Grid environment, these resources belong to different administrative domains, each employing its own scheduling policy. That is, at which time an activity (e.g., compute job, data transfer) is started, is decided by the resource's local management system. In such an environment, the coordinated execution of distributed applications requires guarantees on the quality of service (QoS) of the needed resources. Reserving resources in advance is an accepted means to obtain QoS guarantees from a single provider. The challenge, however, is to coordinate advance reservations of multiple resources. This work presents a system architecture and mechanisms to coordinate multiple advance reservations -- called co-reservations -- for delivering QoS guarantees to complex applications. We formally define the co-reservation problem as an optimization problem. The presented model supports three dimensions of freedom: the start time, the duration and the service level of a reservation. Requests and resources are described in a simple language. After matching the static properties and requirements of either side in a mapping, the reservation mechanism probes information about the future status of the resources. The versatile design of the probing step allows the efficient processing of requests, but also lets the resources express their preferences among the myriads of reservation candidates. Next, the best mapping is found through an implementation of the formal co-reservation model. Then, the mapping has to be secured, i.e., resources need to be allocated to a co-reservation candidate with all-or-nothing semantics. We study several goal-driven sequential and concurrent allocation mechanisms and define schemes for handling allocation failures. Finally, we introduce the concept of virtual resources for seamlessly embedding co-reservations into Grid resource management.
Co-reservations are an efficient means to support guarantees on the allocation of resources in order to execute complex distributed application scenarios. Scheduling multiple moldable co-reservations involves several steps leading to guarantees of resource allocation. Previous work has provided either mechanisms for individual steps only or for selecting co-reservations, but with more limited application scenarios. In this work, we extend previous work by studying the problem as generic optimization problem, by developing a mixed integer linear programming model which allows more freedom in the specification of the goals of requests, resources and the broker, and by performing extensive experiments to analyze important parameters on the time to solve various problem instances.
We propose a mechanism for the co-allocation of multiple resources in Grid environments. By reserving multiple resources in advance, scientific simulations and large-scale data analyses can efficiently be executed with their desired quality-of-service level. Co-allocating multiple Grid resources in advance poses demanding challenges due to the characteristics of Grid environments, which are (1) incomplete status information, (2) dynamic behavior of resources and users, and (3) autonomous resources' management systems. Our co-reservation mechanism addresses these challenges by probing the state of the resources and by enhancing a two-phase commit protocol with timeouts. We performed extensive simulations to evaluate communication overhead of the new protocol and the impact of the timeouts' length on the scheduling of jobs as well as on the utilization of the Grid resources.
Executing complex applications on Grid infrastructures necessitates the guaranteed allocation of multiple resources. Such guarantees are often implemented by means of advance reservations. Reserving resources in advance requires multiple steps – beginning with their description to their actual allocation. In a Grid, a client possesses little knowledge about the future status of resources. Thus, manually specifying successful parameters of a co-reservation is a tedious task. Instead, we propose to parametrize certain reservation characteristics (e.g., the start time) and to let a client define criteria for selecting appropriate values. Then, a Grid reservation service processes such requests by determining the future status of resources and calculating a co-reservation candidate which satisfies the criteria. In this paper, we present the Simple Reservation Language (SRL) for describing the requests, demonstrate the transformation of an example request into an integer program using the Zuse Institute Mathematical Programming Language (ZIMPL) and experimentally evaluate the time needed to find the optimal co-reservation using CPLEX.
We present an architecture for enabling remote access to robotic tele- scopes through the adoption of Grid technology. With this architec- ture, Internet connected robotic telescopes form a global network and are controlled by a global resource management system (scheduler), similar to individual compute resources in a Grid. By virtualizing the access to these telescope resources and by describing them and obser- vation requests in a generic language (RTML). Astronomers are pro- vided with an interface to a telescope network, from which they can get the appropriate resources for their observations. Moreover, new kinds of coordinated observations become feasible, such as multi-wavelength campaigns or immediate and continuous monitoring of transient astro- nomical events. This paper describes the architecture, the processing of observation requests and new research topics in a global network of robotic telescopes.
We present Stellaris, the information service of the community project AstroGrid-D. Stellaris is the core component of the AstroGrid-D middleware that enables scientists to share their resources, provides access to large datasets and integrates instruments such as robotic telescopes. Besides the many diverse types of resources, the information service also supports a wide range of use cases each using a specific schema for the metadata. In addition, Stellaris addresses the distributed and dynamic nature of collaborations in the astronomers’ community. Stellaris satisfies these requirements by adopting RDF and SPARQL for storing and querying metadata. Our paper focuses on the requirements of the community, presents the architecture of the information service in detail and discusses experiences with the prototype already in use by partners within the project.
We present AstroGrid-D, a project bringing together astronomers and experts in Grid technology to enhance astronomic science in many aspects. First, by sharing currently dispersed resources, scientists can calculate their models in more detail. Second, by developing new mechanisms to e‐ciently access and process existing datasets, scientiflc problems can be investigated that were until now impossible to solve. Third, by adopting Grid technology large instruments such as robotic telescopes and complex scientiflc work∞ows from data aquisition to analysis can be managed in an integrated manner. In this paper, we present prominent astronomic use cases, discuss requirements on a Grid middleware and present our approach to extend/augment existing middleware to facilitate the improvements mentioned above.
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 Change History . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
We present a new method for determining placements of flexible reservation requests into a schedule. For each considered placement the what-if method inserts a placeholder into the schedule and simulates the processing of batch jobs currently known to the system. Each placement is evaluated wrt. well-known scheduling metrics. This information may be used by a Grid reservation service to choose the most likely successful placement of a reservation. According to the results of extensive simulations, the what-if method grants more reservations and improves the performance of local jobs compared to our previously used load method.
We present an architectural framework for specifying and processing co-reservations in grid environments. Compared to other approaches, our co-reservation framework is more general. It can be applied to the reservation of applications running concurrently on multiple resources (multi-site applications) and to the planning of job flows, where the components may be linked by some temporal or spatial relationship. We introduce the concept of virtual resources that allows to compose resources by abstracting from their specific features. Among other features, virtual resources enable advanced co-reservation with nested resource levels, resource aggregation, and transparent fault recovery.
We present a framework for the co-ordinated, autonomic management of multiple clusters in a compute center and their integration into a Grid environment. Site autonomy and the automation of administrative tasks are prime aspects in this framework. The system behavior is continuously monitored in a steering cycle and appropriate actions are taken to resolve any problems. All presented components have been implemented in the course of the EU project DataGrid: The Lemon monitoring components, the FT fault-tolerance mechanism, the quattor system for software installation and configuration, the RMS job and resource management system, and the Gridification scheme that integrates clusters into the Grid.
Clusters provide an outstanding cost/performance ratio, but their efficient orchestration, i.e. their cooperative management, maintenance, and use, still poses difficulties. Moreover, many sites operate multiple clusters, each possible running under a different cluster management system. In this paper, we present an architectural scheme for the coordinated management of multiple clusters in a fabric. Our scheme allows different cluster management systems to interact with each other via adaptors, thereby providing interoperability within a single administrative entity, the fabric. Using adaptors, various jobs (grid jobs, local jobs, system maintenance jobs) can be served by the same methods, and existing cluster management software (like LSF, PBS, CCS, etc.) can be extended by additional functions without much modification effort.
The contributions of this work are twofold. First, we describe the design and implementation of a simulation environment for an open-source embedded kernel and an intuitive user interface to complement it. Second, the simulator can be used for embedded program development and research as well as instructional purposes in embedded system classes as a replacement or a complement to hands-on experiments with embedded devices.The technical sections of this article stress the suitability of POSIX Threads (Pthreads) in approximating kernel operations in the simulation environment. We specify the prerequisites for using Pthreads as a means to approximate embedded task execution and suggests an I/O-based representation of device information. The experience gained with a sample implementation stresses the importance of a proper match between a Pthreads implementation and an embedded kernel. We demonstrate the adequacy of both the simulation environment and a graphical user interface to aid program development and debugging. Furthermore, the separation of the simulation component from the user interface provides opportunities to utilize each component separately or even combine them with other components. The simulation environment is publically available, and instructions for installation and use are included in the Appendix. The combination of technical solutions is a contribution toward embedded system simulation, an area that has not been studied much.
We present an architecture for enabling remote access to robotic telescopes through the adoption of Grid technology. With this architecture, Internet connected robotic telescopes form a global network and are controlled by a global resource management system (scheduler), similar to individual compute resources in a Grid. By virtualizing the access to these telescope resources and by describing them and observation requests in a generic language (RTML). Astronomers are provided with an interface to a telescope network, from which they can get the appropriate resources for their observations. Moreover, new kinds of coordinated observations become feasible, such as multi-wavelength campaigns or immediate and continuous monitoring of transient astronomical events. This paper describes the architecture, the processing of observation requests and new research topics in a global network of robotic telescopes.
Ramin Yahyapour合作论文数the new IT and Media Center;University Dortmund3
Thomas Radke合作论文数Max Planck Institute for Gravitational Physics2
Frank Mueller合作论文数Department of Computer Science, North Carolina State University1