Artificial Intelligence and Machine Learning have emerged as a promising approach to scientific investigations, but there is a persistent shortage of high-quality, properly annotated datasets suitable for training models. Here, we outline some of the widely reported characteristics for making AI-ready data and compare that with FAIR data. We discuss the limitations of traditional data repositories and the challenges associated with establishing a data repository that can grow with scientific communities and accommodate rapid evolution in research priorities. Finally, we introduce the SCALE principles for repository design that offer a proven framework for creating sustainable, scalable data repositories that can adapt to infrastructure serves research needs rather than constraining them.
The ability to accumulate and analyze large quantities of data is rapidly becoming a competitive advantage not only in science but in the broader economy as well. Advances such as AlphaFold, the AI-based protein prediction tool and ChatGPT the large language model-based chat bot, have ignited enormous excitement in science and industry for leveraging data and computational techniques to solve important problems. However, what is typically lost in all the excitement is the fact that such startling achievements were only possible after a critical mass of high quality data existed to train models using machine learning algorithms. Both examples relied on open data sources that were generated painstakingly by user communities over the course of decades. We argue that in order to unlock future high impact data science achievements like these will require a culture of and skill set for data management, sharing and reuse. In this paper, we describe our work within the dental, oral, and craniofacial community to create such a data sharing community that has grown out of basic research to increasingly touch on clinical sciences and even commercial applications.
The underlying premise behind eScience is that computational methods and data-driven approaches can contribute to scientific discovery on a par with, or even superior to, traditional experimental methods; that the combination of computers, software, and extant data collections are the modern equivalent to the scientific instruments that have led to our understanding of fundamental laws in physics, chemistry, biology, and other domains. However, a robust methodology for making the results of eScience activities “scientific” is lacking, with significant consequences. In this brief paper we propose a shift in perspective as to what it means to create an eScience-based result and how the scientific validity of eScience experiments might be improved.
Database management systems have been used to great advantage for industry usage scenarios. As science becomes increasingly dependent on carefully organized and curated data to inform and drive new discoveries, the need for database management systems has grown significantly. The long standing “20 questions” method was advocated in early studies of applying relational databases for science in order to elicit requirements for designing and developing information systems for scientific data. It has been observed, however, that database designs become outdated within months of usage leading to degradation in the quality of the database schema. In addition, there is limited evidence that scientists themselves have the tools and processes necessary to develop and maintain scientific databases without reliance on database administrators. Beyond learning to query databases, scientists need tools to create and evolve databases and guidance on how to apply those tools to develop information systems. In this paper, we present a simplified methodology for database evolution for scientists and a case study of database evolution by a scientist in the context of a research database for cell modeling. We include a detailed analysis of the activities and processes employed by the scientist during the schema evolution. Our results show that a scientist can successfully evolve a complex information system driven by new research requirements.
The Common Fund Data Ecosystem (CFDE) has created a flexible system of data federation that enables researchers to discover datasets from across the US National Institutes of Health Common Fund without requiring that data owners move, reformat, or rehost those data. This system is centered on a catalog that integrates detailed descriptions of biomedical datasets from individual Common Fund Programs’ Data Coordination Centers (DCCs) into a uniform metadata model that can then be indexed and searched from a centralized portal. This Crosscut Metadata Model (C2M2) supports the wide variety of data types and metadata terms used by individual DCCs and can readily describe nearly all forms of biomedical research data. We detail its use to ingest and index data from 11 DCCs.
Scientific databases used for organizing, archiving, collaborating and sharing research data depend on a well-defined schema to accurately reflect the scientific domain and on database-driven applications for supporting key user interactions with the database. Applications that interact with a database typically depend on some form of schema mappings, such as object-relational mappings, to inform the application of how to query and manipulate the database. The presence of schema mappings, however, further exacerbates the already difficult task of evolving the database schema. Database migration utilities provide some help by coordinating schema evolution scripts with application code changes, but only automate the simplest schema mapping changes. In this paper, we present an approach to coupled database-application evolution by extending a database evolution language with model management operations. We introduce a novel set of model management operations and define their semantics and then describe how they may be integrated into schema modification operators. We then present an evaluation of the concepts from real-world usage of model mappings in scientific database deployments.
Complex morphological traits are the product of many genes with transient or lasting developmental effects that interact in anatomical context. Mouse models are a key resource for disentangling such effects, because they offer myriad tools for manipulating the genome in a controlled environment. Unfortunately, phenotypic data are often obtained using laboratory-specific protocols, resulting in self-contained datasets that are difficult to relate to one another for larger scale analyses. To enable meta-analyses of morphological variation, particularly in the craniofacial complex and brain, we created MusMorph, a database of standardized mouse morphology data spanning numerous genotypes and developmental stages, including E10.5, E11.5, E14.5, E15.5, E18.5, and adulthood. To standardize data collection, we implemented an atlas-based phenotyping pipeline that combines techniques from image registration, deep learning, and morphometrics. Alongside stage-specific atlases, we provide aligned micro-computed tomography images, dense anatomical landmarks, and segmentations (if available) for each specimen (N = 10,056). Our workflow is open-source to encourage transparency and reproducible data collection. The MusMorph data and scripts are available on FaceBase (www.facebase.org, https://doi.org/10.25550/3-HXMC) and GitHub (https://github.com/jaydevine/MusMorph).
The FaceBase Consortium, funded by the National Institute of Dental and Craniofacial Research of the National Institutes of Health, was established in 2009 with the recognition that dental and craniofacial research are increasingly data-intensive disciplines. Data sharing is critical for the validation and reproducibility of results as well as to enable reuse of data. In service of these goals, data ought to be FAIR: Findable, Accessible, Interoperable, and Reusable. The FaceBase data repository and educational resources exemplify the FAIR principles and support a broad user community including researchers in craniofacial development, molecular genetics, and genomics. FaceBase demonstrates that a model in which researchers "self-curate" their data can be successful and scalable. We present the results of the first 2.5 y of FaceBase's operations as an open community and summarize the data sets published during this period. We then describe a research highlight from work on the identification of regulatory networks and noncoding RNAs involved in cleft lip with/without cleft palate that both used and in turn contributed new findings to publicly available FaceBase resources. Collectively, FaceBase serves as a dynamic and continuously evolving resource to facilitate data-intensive research, enhance data reproducibility, and perform deep phenotyping across multiple species in dental and craniofacial research.
Discovery of new knowledge is increasingly data-driven, predicated on a team's ability to collaboratively create, find, analyze, retrieve, and share pertinent datasets over the duration of an investigation. This is especially true in the domain of scientific discovery where generation, analysis, and interpretation of data are the fundamental mechanisms by which research teams collaborate to achieve their shared scientific goal. Data-driven discovery in general, and scientific discovery in particular, is distinguished by complex and diverse data models and formats that evolve over the lifetime of an investigation. While databases and related information systems have the potential to be valuable tools in the discovery process, developing effective interfaces for data-driven discovery remains a roadblock to the application of database technology as an essential tool in scientific investigations. In this paper, we present a model-adaptive approach to creating interaction environments for data-driven discovery of scientific data that automatically generates interactive user interfaces for editing, searching, and viewing scientific data based entirely on introspection of an extended relational data model. We have applied model-adaptive interface generation to many active scientific investigations spanning domains of proteomics, bioinformatics, neuroscience, occupational therapy, stem cells, genitourinary, craniofacial development, and others. We present the approach, its implementation, and its evaluation through analysis of its usage in diverse scientific settings.
The FaceBase Consortium was established by the National Institute of Dental and Craniofacial Research in 2009 as a 'big data' resource for the craniofacial research community. Over the past decade, researchers have deposited hundreds of annotated and curated datasets on both normal and disordered craniofacial development in FaceBase, all freely available to the research community on the FaceBase Hub website. The Hub has developed numerous visualization and analysis tools designed to promote integration of multidisciplinary data while remaining dedicated to the FAIR principles of data management (findability, accessibility, interoperability and reusability) and providing a faceted search infrastructure for locating desired data efficiently. Summaries of the datasets generated by the FaceBase projects from 2014 to 2019 are provided here. FaceBase 3 now welcomes contributions of data on craniofacial and dental development in humans, model organisms and cell lines. Collectively, the FaceBase Consortium, along with other NIH-supported data resources, provide a continuously growing, dynamic and current resource for the scientific community while improving data reproducibility and fulfilling data sharing requirements.
Persistent identifiers (PIDs) are essential for making data Findable, Accessible, Interoperable, and Reusable, or FAIR. While the advantages of PIDs for data publication and citation are well understood, and Digital Object Identifiers (DOIs) are increasingly applied to data, there are two gaps in the current identifier ecosystem: 1) services that provide a consistent baseline of capabilities encompassing key aspects of the research data lifecycle, including canonical landing pages and machine-readable metadata via the same URL; and 2) support for identifiers to be applied to ephemeral data, particularly as data move across system boundaries, such as during workflows. To address these gaps, we have implemented the FAIR Research Identifiers service. This service supports multiple identifier providers (ARK, Handle, DOIs via DataCite, etc.) and uses Globus Auth to implement a rich user- and group-based authorization model for identifier creation. This paper summarizes the current identifier ecosystem, presents best-practices recommendations for identifier use, and describes our FAIR Research Identifiers service.
In order to conduct research effectively, scientists must be able to access, organize, describe, and produce data as part of their daily research activities. While relational databases are well suited to the tasks of describing and organizing scientific metadata and results, the difficulties of using relational database management systems effectively, have resulted in their limited adoption among scientists. In addition, scientific research is changing steadily with new experimental protocols, instruments, and discoveries that determine what data are generated and how they must be described and organized according to a relational schema. Unfortunately, evolving a schema is one of the most difficult aspects of database usage. The conventional data definition and manipulation languages offer relatively low-level programming abstractions to perform complex database evolution tasks, and therefore require specialized technical skills not possessed by most scientists. A simplified means of expressing database evolution operations would reduce the effort for non-expert users of databases. This paper presents a high-level, user-oriented, schema evolution framework built on a formal algebra of schema modification operators. The approach allows introduction of novel operators as motivated by new requirements and is amenable to well established optimization techniques for efficient planning and execution. We also propose a rigorous evaluation methodology for comparing the user effort of database evolution languages, and we introduce a benchmark for evaluating the execution efficiency of schema evolution expressions. We present the framework and its implementation, and we demonstrate its utility in exemplar use cases and a performance evaluation.
Database evolution is a notoriously difficult task, and it is exacerbated by the necessity to evolve database-dependent applications. As science becomes increasingly dependent on sophisticated data management, the need to evolve an array of database-driven systems will only intensify. In this paper, we present an architecture for data-centric ecosystems that allows the components to seamlessly co-evolve by centralizing the models and mappings at the data service and pushing model-adaptive interactions to the database clients. Boundary objects fill the gap where applications are unable to adapt and need a stable interface to interact with the components of the ecosystem. Finally, evolution of the ecosystem is enabled via integrated schema modification and model management operations. We present use cases from actual experiences that demonstrate the utility of our approach.
Databases are well suited to the task of describing and organizing research datasets, however, the difficulties of using database management systems effectively have resulted in their limited usage among domain scientists. Scientists operate in an environment that is changing steadily with new experimental protocols, instruments, and discoveries that impact what datasets they generate and how they describe and organize them. In order to manage datasets for a scientific application, scientists need to routinely revise their database schemas to reflect these changes. Unfortunately, evolving a database is one of the well-known and most difficult aspects of database usage. The conventional data definition and manipulation languages offer relatively low-level programming abstractions to perform complex database evolution tasks, and therefore require specialized technical skills not possessed by most domain scientists. A simplified means of expressing database evolution operations can reduce the effort of keeping the scientific database in sync with changing requirements. This paper presents a high-level, user-oriented, schema evolution framework with an algebra of specialized schema modification operators. The approach allows introduction of novel operators as motivated by new requirements and is amenable to well established optimization techniques for efficient planning and execution. We present the framework and its implementation, and we demonstrate its utility in an exemplar use case and performance evaluation.
Sharing of bioinformatics data within research communities holds the promise of facilitating more rapid discovery, yet the volume of data is growing at a pace exponentially greater than what traditional biocuration can support. We present here an approach that we have used to empower data producing researchers to curate high quality shared data that is ready for reuse and re-analysis.
Human kidney function is underpinned by approximately 1,000,000 nephrons, although the number varies substantially, and low nephron number is linked to disease. Human kidney development initiates around 4 weeks of gestation and ends around 34-37 weeks of gestation. Over this period, a reiterative inductive process establishes the nephron complement. Studies have provided insightful anatomic descriptions of human kidney development, but the limited histologic views are not readily accessible to a broad audience. In this first paper in a series providing comprehensive insight into human kidney formation, we examined human kidney development in 135 anonymously donated human kidney specimens. We documented kidney development at a macroscopic and cellular level through histologic analysis, RNA in situ hybridization, immunofluorescence studies, and transcriptional profiling, contrasting human development (4-23 weeks) with mouse development at selected stages (embryonic day 15.5 and postnatal day 2). The high-resolution histologic interactive atlas of human kidney organogenesis generated can be viewed at the GUDMAP database (www.gudmap.org) together with three-dimensional reconstructions of key components of the data herein. At the anatomic level, human and mouse kidney development differ in timing, scale, and global features such as lobe formation and progenitor niche organization. The data also highlight differences in molecular and cellular features, including the expression and cellular distribution of anchor gene markers used to identify key cell types in mouse kidney studies. These data will facilitate and inform in vitro efforts to generate human kidney structures and comparative functional analyses across mammalian species.
The foundation of data oriented scientific collaboration is the ability for participants to find, access and reuse data created during the course of an investigation, what has been referred to as the FAIR principles. In this paper, we describe ERMrest, a collaborative data management service that promotes data oriented collaboration by enabling FAIR data management throughout the data life cycle. ERMrest is a RESTful web service that promotes discovery and reuse by organizing diverse data assets into a dynamic entity relationship model. We present details on the design and implementation of ERMrest, data on its performance and its use by a range of collaborations to accelerate and enhance their scientific output.
Database systems are well suited to scientific data management and analysis workloads, however, a database must evolve to keep pace with changing requirements and adjust to changes in the domain conceptualization as applications mature. Evolving a database (i.e., updating its schema and instance data) is one of the greatest challenges in database maintenance and the difficulties are compounded by the lack of sufficient tools to support scientists. This paper presents a schema evolution framework based on an algebraic approach that introduces extended and higher-level composite relational operators tailored to the task of schema evolution. These higher-level operators simplify the task of evolving a database for non-expert users, while enabling efficient evaluation of schema evolution expressions.
The pace of discovery in eScience is increasingly dependent on a scientist's ability to acquire, curate, integrate, analyze, and share large and diverse collections of data. It is all too common for investigators to spend inordinate amounts of time developing ad hoc procedures to manage their data. In previous work, we presented DERIVA, a Scientific Asset Management System, designed to accelerate data driven discovery. In this paper, we report on the use of DERIVA in a number of substantial and diverse eScience applications. We describe the lessons we have learned, both from the perspective of the DERIVA technology, as well as the ability and willingness of scientists to incorporate Scientific Asset Management into their daily workflows.
Creating and maintaining an accurate description of data assets and the relationships between assets is a critical aspect of making data findable, accessible, interoperable, and reusable (FAIR). Typically, such metadata are created and maintained in a data catalog by a curator as part of data publication. However, allowing metadata to be created and maintained by data producers as the data is generated rather then waiting for publication can have significant advantages in terms of productivity and repeatability. The responsibilities for metadata management need not fall on any one individual, but rather may be delegated to appropriate members of a collaboration, enabling participants to edit or maintain specific attributes, to describe relationships between data elements, or to correct errors. To support such collaborative data editing, we have created ERMrest, a relational data service for the Web that enables the creation, evolution and navigation of complex models used to describe and structure diverse file or relational data objects. A key capability of ERMRest is its ability to control operations down to the level of individual data elements, i.e. fine-grained access control, so that many different modes of data-oriented collaboration can be supported. In this paper we introduce ERMRest and describe its fine-grained access control capabilities that support collaborative editing. ERMrest is in daily use in many data driven collaborations and we describe a sample policy that is based on a common biocuration pattern.