The Integrated Rule-Oriented Data System (iRODS) is open source data management software used by research organizations and government agencies worldwide. iRODS is released as a production-level distribution aimed at deployment in mission critical environments. It virtualizes data storage resources, so users can take control of their data, regardless of where and on what device the data is stored. As data volumes grow and data services become more complex, iRODS is increasingly important in data management. The development infrastructure supports exhaustive testing on supported platforms; plug-in support for microservices, storage resources, drivers, and databases; and extensive documentation, training and support services. This book is a microservice workbook, with descriptions of the input and output parameters and usage examples for each of the available microservices. Together with the iRODS Primer, a community may use this book to assemble a data management infrastructure that reliably enforces their management policies, automates administrative tasks, and validates assessment criteria. The microservices referenced in this book are supported by iRODS version 4.0, an open source release from the iRODS Consortium, http://irods.org/consortium.
The iRODS system belongs to a class of middleware that we term adaptive middleware. The Adaptive Middleware Architecture (AMA) provides a means for adapting the middleware to meet the needs of the end user community without requiring that they make programming changes. One can view the AMA middleware as a glass box in which users can see how the system works and can tweak the controls to meet their demands. Usually, middleware is the equivalent of a black box for which no changes are programmatically possible to adjust the flow of the operations, except predetermined configuration options that may allow one to set the starting conditions of the middleware.
ROP is a different (although not new) paradigm from normal programming practice. In ROP, the power of controlling the functionality rests more with the users than with system and application developers. Hence, any change to a particular process or policy can be easily constructed by the user, and then tested and deployed without the aid of system and application developers.
This paper describes the architecture of the SDSC Storage Resource Broker (SRB). The SRB is middleware that provides applications a uniform API to access heterogeneous distributed storage resources including, filesystems, database systems, and archival storage systems. The SRB utilizes a metadata catalog service, MCAT, to provide a "collection"- oriented view of data. Thus, data items that belong to a single collection may, in fact, be stored on heterogeneous storage systems. The SRB infrastructure is being used to support digital library projects at SDSC. This paper describes the architecture and various features of the SDSC SRB.
Micro-services are small, well-defined procedures/functions that perform a simple task. Micro-services are developed and made available by system programmers and application programmers and compiled into the iRODS Server code. Users and administrators can chain these Micro-services to implement a function that they want to use or provide for others. In this manner, the users/administrators can have full control over what happens when one performs a macro-level functionality. These macro-level functionalities are called Actions. By having more than one chain of Microservices for an Action, a system can have multiple ways of performing the Action. Using priorities and validation conditions at run-time, the system chooses the “best” Micro-service chain to be executed. There are other caveats to this execution paradigm that were discussed in Chapter 4.
Large ground-based and space-based telescopes are expected to make exciting discoveries in the upcoming decade. These large projects start their construction phase many years before first-light and continue to operate for many years after first-light and usually span multiple countries. The file-storage cyberinfrastructure ("file-storage CI") of these largescale projects has to evolve over several years from a conceptual prototype to a highly flexible data distribution network. During this long period the file-storage CI has to transition into multiple stages, starting with a conceptual prototype before first-light, to a large-scale distributed network in production, and finally into a persistent archive once the project is decommissioned. While the project makes these transitions, the file-storage CI has to incorporate several requirements including but not limited to: Technology Evolution, due to changes in Cyberinfrastructure (CI) software or hardware during the lifetime of the project; International Partnerships that are updated during the various phases of the project; and Data Lifecycle that exists in the project. The file-storage and management software's architecture has to be designed with significant consideration of these requirements for these large projects. In this paper, we provide the generic requirements, for file-storage and management cyberinfrastructure in a large project similar to LSST before first-light.
Large-scale Data Grid Systems (LDGS) facilitate collaborative sharing of large collections (Petabytes and100s of millions of objects) containing files, databases and data streams that are geographically distributed across heterogeneous resources and multiple administrative domains. LDGS provide a “universal view” of the distributed data, resources, users and methods and hide the idiosyncrasies and the heterogeneity of the underlying infrastructure and protocols - enhancing user collaborations. To improve transparency, an “open policy” system is needed by which data providers and administrators can describe the exact processes and policies that implement LDGS services. We consider policies and processes as the essential defining characteristics of a productive LDGS collaboration. We have implemented an LDGS, called integrated Rule-Oriented Data Systems (iRODS), which provides a universal view while enabling an open policy environment for publishing descriptions of the available services. The open policy environment is supported by a distributed workflow/rule engine. The services are encoded as rules in a high-level workflow language that transparently describes the underlying functionality. Well-defined semantics are used to control the composition of the workflow functions, called micro-services, to map to the desired client-level actions. In this paper, we describe the iRODS system from the “universal view” and “open policy” perspective and show its scalability for managing more than 10 million files.
A method is presented to calculate thermodynamic conformational entropy of a biomolecule from molecular dynamics simulation. Principal component analysis (the quasi-harmonic approximation) provides the first decomposition of the correlations in particle motion. Entropy is calculated analytically as a sum of independent quantum harmonic oscillators. The largest classical eigenvalues tend to be more anharmonic and show statistical dependence beyond correlation. Their entropy is corrected using a numerical method from information theory: the k-nearest neighbor algorithm. The method calculates a tighter upper limit to entropy than the quasi-harmonic approximation and is likewise applicable to large solutes, such as peptides and proteins. Together with an estimate of solute enthalpy and solvent free energy from methods such as MMPB/SA, it can be used to calculate the free energy of protein folding as well as receptor-ligand binding constants.
Trusted digital repository audit checklists are now being developed, based on assessments of organizational infrastructure, repository functions, community use, and technical infrastructure. These assessments can be expressed as rules that are applied on state information that define the criteria for trustworthiness. This paper maps the rules to the mechanisms that are needed in a trusted digital repository to minimize risk of both data and state information loss. The required mechanisms have been developed within the Storage Resource Broker data grid technology, and their use is illustrated on existing preservation repository projects.
International data grids are now being built that support joint management of shared collections. An emerging strategy is to build multiple independent data grids, each managed by the local institution. The data grids are then federated to enable controlled sharing of files. We examine the management issues associated with maintaining federations of production data grids, including management of access controls, coordinated sharing of name spaces, replication of data between data grids, and expansion of the data grid federation.
The management of globally distributed data requires capabilities that have been developed in the data grid, digital library, and archivist communities. The storage resource broker incorporates essential features from each of these communities to support international collaborations that share data collections, provide publication environments for scientific data collections, and sustain preservation environments.
The integration of grid, data grid, digital library, and preservation technology has resulted in software infrastructure that is uniquely suited to the generation and management of data. Grids provide support for the organization, management, and application of processes. Data grids manage the resulting digital entities. Digital libraries provide support for the management of information associated with the digital entities. Persistent archives provide long-term preservation. We examine the synergies between these data management systems and the future evolution that is required for the generation and management of information.
Wayne Schroeder合作论文数San Diego Supercomputer Center
University of California, San Diego17
Reagan W. Moore合作论文数Data Intensive Computing
San Diego Supercomputer Center5