This paper describes the problems and explores potential solutions for providing long term storage and access to research outputs, focusing mainly on research data. The ready availability of cloud storage and compute services provides a potentially attractive option for curation and preservation of research information. In contrast to deploying infrastructure within an organisation, which normally requires long lead times and upfront capital investment, cloud infrastructure is available on demand and is highly scalable. However, use of commercial cloud services in particular raises issues of governance, cost-effectiveness, trust and quality of service. We describe a set of in-depth case studies conducted with researchers across the sciences and humanities performing data-intensive research, which demonstrate the issues that need to be considered when preserving data in the cloud. We then describe the design of a repository framework that addresses these requirements. The framework uses hybrid cloud, combining internal institutional storage, cloud storage and cloud-based preservation services into a single integrated repository infrastructure. Allocation of content to storage providers is performed using on a rules-based approach. The results of an evaluation of the proof-of-concept system are described.
The paper describes the investigations and outcomes of the JISC-funded Kindura project, which is piloting the use of hybrid cloud infrastructure to provide repository-focused services to researchers. The hybrid cloud services integrate external commercial cloud services with internal IT infrastructure, which has been adapted to provide cloud-like interfaces. The system provides services to manage and process research outputs, primarily focusing on research data. These services include both repository services, based on use of the Fedora Commons repository, as well as common services such as preservation operations that are provided by cloud compute services. Kindura is piloting the use of the DuraCloud2, open source software developed by DuraSpace, to provide a common interface to interact with cloud storage and compute providers. A storage broker integrates with DuraCloud to optimise the usage of available resources, taking into account such factors as cost, reliability, security and performance. The development is focused on the requirements of target groups of researchers.
We present the architecture and design of a "cloudy" data infrastructure for archiving and backup.By reducing the cloud elasticity, we are able to build a more cost effective service for archiving.DuraSpace and Fedora provide friendly front-ends for users.We have investigated several options for the back end, focusing currently on a federated iRODS infrastructure which will permit automatic replication and metadata extraction.All services will appear as "cloudlike", even internal onesit is thus a hybrid approach that combines the advantages of the commercial/external (public) cloud with an institutional/consortium (private/community) cloud.This project will provide Infrastructure-as-a-Service (IaaS) components, via storage and compute services, but more importantly it will combine these, using DuraCloud and Fedora as enabling technologies, to provide an integrated Software-as-a-Service (SaaS) package of repository-centric services.While this is work in progress, we can already present results.In future work, DuraCloud will be extended to broker between the clouds.
In global Grids, interoperation is important. It enables communities to work together, helps prevent vendor lock-in, and in principle enables “cloud-like” resource provision by permitting different resources to meet needs from other communities. In this paper, we discuss a practical example of achieving interoperation between storage resources, and the lessons learned. The aim is to meet current use cases for interoperation with no additional software development. Apart from the practical results, experiences from this work will be relevant to other interoperation activities.
Scientific facilities, in particular large-scale photon and neutron sources, have demanding requirements to manage the increasing quantities of experimental data they generate in a systematic and secure way. In this paper, we describe the ICAT infrastructure for cataloguing facility-generated experimental data which has been in development within STFC and DLS for several years. We consider the factors which have influenced its design and describe its architecture and metadata model, a key tool in the management of data. We go on to give an outline of its current implementation and use, with plans for its future development.