Introduction: Significant progress has been made in terms of best practice in research data management for nanosafety. Some of the underlying approaches to date are, however, overly focussed on the needs of specific research projects or aligned to a single data repository, and this "silo" approach is hampering their general adoption by the broader research community and individual labs.Methods: State-of-the-art data/knowledge collection, curation management FAIrification, and sharing solutions applied in the nanosafety field are reviewed focusing on unique features, which should be generalised and integrated into a functional FAIRification ecosystem that addresses the needs of both data generators and data (re)users.Results: The development of data capture templates has focussed on standardised single-endpoint Test Guidelines, which does not reflect the complexity of real laboratory processes, where multiple assays are interlinked into an overall study, and where non-standardised assays are developed to address novel research questions and probe mechanistic processes to generate the basis for read-across from one nanomaterial to another. By focussing on the needs of data providers and data users, we identify how existing tools and approaches can be re-framed to enable "on-the-fly" (meta) data definition, data capture, curation and FAIRification, that are sufficiently flexible to address the complexity in nanosafety research, yet harmonised enough to facilitate integration of datasets from different sources generated for different research purposes. By mapping the available tools for nanomaterials safety research (including nanomaterials characterisation, nonstandard (mechanistic-focussed) methods, measurement principles and experimental setup, environmental fate and requirements from new research foci such as safe and sustainable by design), a strategy for integration and bridging between silos is presented. The NanoCommons KnowledgeBase has shown how data from different sources can be integrated into a one-stop shop for searching, browsing and accessing data (without copying), and thus how to break the boundaries between data silos.Discussion: The next steps are to generalise the approach by defining a process to build consensus (meta)data standards, develop solutions to make (meta)data more machine actionable (on the fly ontology development) and establish a distributed FAIR data ecosystem maintained by the community beyond specific projects. Since other multidisciplinary domains might also struggle with data silofication, the learnings presented here may be transferrable to facilitate data sharing within other communities and support harmonization of approaches across disciplines to prepare the ground for cross-domain interoperability.
The International Union of Pure and Applied Chemistry has a long tradition of supporting the compilation of chemical data and their evaluation through direct projects, nomenclature and terminology work, and partnerships with international scientific bodies, government agencies, and other organizations. The IUPAC Interdivisional Subcommittee on Critical Evaluation of Data has been established to provide guidance on issues related to the evaluation of chemical data. In this first report, we define the general principles of the evaluation of scientific data and describe best practices and approaches to data evaluation in chemistry.
This essay discusses some of the considerations that led to the founding of the [CODATA] Data Science Journal. Three factors were most relevant to the founding. First, there was a need to have a more formal publication mechanism for the papers given at the biennial CODATA International Conferences. Second, there was a pressing need for data science advancements made in one area of scientific data work to be shared with other scientific disciplines. Lastly the increasing number of scientists interested in data, throughout science, and throughout the world, required a more convenient publication outlet. Thus arose the Data Science Journal.
In this paper, categorization of nanomaterials is examined from four perspectives; context, criteria for success, ensuring measurements are relevant, and the life cycle of a nanomaterial. For each perspective, its relevance to categorization is discussed as well as the difficulties it presents. For example, while the context of assessing potential harm to living things and the environment is clearly important, other contexts are often needed and require different categorization schemes. Understanding what success means for a categorization scheme, within its target context, is critical to making sure a categorization is actually useful. The complexity of nanomaterials and their interactions makes generating and collecting the required data and metadata to support categorization a challenge. Finally, the transformation a nanomaterial undergoes through its lifetime, including the testing process, present additional challenges to accurate categorization. How these factors impact development of usable categorization schemes is analyzed.
The emergence of nanoinformatics as a key component of nanotechnology and nanosafety assessment for the prediction of engineered nanomaterials (NMs) properties, interactions, and hazards, and for grouping and read-across to reduce reliance on animal testing, has put the spotlight firmly on the need for access to high-quality, curated datasets. To date, the focus has been around what constitutes data quality and completeness, on the development of minimum reporting standards, and on the FAIR (findable, accessible, interoperable, and reusable) data principles. However, moving from the theoretical realm to practical implementation requires human intervention, which will be facilitated by the definition of clear roles and responsibilities across the complete data lifecycle and a deeper appreciation of what metadata is, and how to capture and index it. Here, we demonstrate, using specific worked case studies, how to organise the nano-community efforts to define metadata schemas, by organising the data management cycle as a joint effort of all players (data creators, analysts, curators, managers, and customers) supervised by the newly defined role of data shepherd. We propose that once researchers understand their tasks and responsibilities, they will naturally apply the available tools. Two case studies are presented (modelling of particle agglomeration for dose metrics, and consensus for NM dissolution), along with a survey of the currently implemented metadata schema in existing nanosafety databases. We conclude by offering recommendations on the steps forward and the needed workflows for metadata capture to ensure FAIR nanosafety data.
New nanomaterials make comprehensive testing more difficult and costly. Categorization technology is being applied to nanomaterials, such as expansion of quantitative structure activity relationship technology (QSAR), to reduce testing burdens. Success is dependent critically on accurate categorization. To date nanomaterials categorization schemes have used chemical composition, shape, size, size distribution, surface coatings, and other properties. No categorization procedure, however, has yet been successful in predicting important properties, especially those related to risk to health or environment. In this paper, we define a set of goals for nanomaterials categorization and discuss them in detail. One problem is the lack of precision in describing nanomaterials well enough to correlate specific features with important properties to determine cause and effect. We present a systematic approach to describing nanomaterials that supports categorization and identifies the information needed. We clearly differentiate between an individual piece of nanomaterial (nano-object) and collections of nano-objects and highlight the importance of differentiating among different types of nanomaterials, including changes during their life cycle. We discuss how this description system supports more precise nanomaterials characterization by enabling more precise correlation of properties to specific features, an important component of successful categorization. We conclude with thoughts about the importance of using natural language in description systems and categorization.
Science today is rapidly becoming both multi-disciplinary and data-driven. These two trends pose new challenges to the capture, management, sharing, and dissemination of research data. Multi-disciplinary science means diverse data generation communities and equally diverse user groups. Data-driven means that sharing data among different communities is more important than ever because of the growth of modeling and knowledge discovery. Nanotechnology is a prime example, involving chemistry, physics, materials science, toxicology, environmental science, and many other disciplines. During the past few years, CODATA has created an international, multi-disciplinary Working Group that has developed a number of critically important metadata standards to facilitate sharing nanomaterials data. In this paper, we discuss the challenges faced in starting and executing this work, as well as the approaches taken to make progress on producing internationally accepted metadata standards. Many of these approaches are directly applicable to other multi-disciplinary subjects.
Providing better availability to materials data has recently gained new momentum. Many successes abound—large numbers of individual materials databases exist, powerful modeling and data analysis approaches have been developed, and Web-based technologies are available. At the same time, challenges remain: one-stop access is lacking, use of multiple databases at the same time is virtually impossible, using shared data is difficult, and understanding data quality is very hard. In this paper, we review the successes and challenges of accessing digital materials data, especially as new initiatives are starting. We also identify insights from previous work that provide guidance to future progress, including adherence to the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles, in achieving this dream.
Many groups within the broad field of nanoinformatics are already developing data repositories and analytical tools driven by their individual organizational goals. Integrating these data resources across disciplines and with non-nanotechnology resources can support multiple objectives by enabling the reuse of the same information. Integration can also serve as the impetus for novel scientific discoveries by providing the framework to support deeper data analyses. This article discusses current data integration practices in nanoinformatics and in comparable mature fields, and nanotechnology-specific challenges impacting data integration. Based on results from a nanoinformatics-community-wide survey, recommendations for achieving integration of existing operational nanotechnology resources are presented. Nanotechnology-specific data integration challenges, if effectively resolved, can foster the application and validation of nanotechnology within and across disciplines. This paper is one of a series of articles by the Nanomaterial Data Curation Initiative that address data issues such as data curation workflows, data completeness and quality, curator responsibilities, and metadata.
Nanoparticle zeta-potentials are relatively easy to measure, and have consistently been proposed in guidance documents as a particle property that must be included for complete nanoparticle characterization. There is also an increasing interest in integrating data collected on nanomaterial properties and behavior measured in different systems (e.g. in vitro assays, surface water, soil) to identify the properties controlling nanomaterial fate and effects, to be able to integrate and reuse datasets beyond their original intent, and ultimately to predict behaviors of new nanomaterials based on their measured properties (i.e. read across), including zeta-potential. Several confounding factors pose difficulty in taking, integrating and interpreting this measurement consistently. Zeta-potential is a modeled quantity determined from measurements of the electrophoretic mobility in a suspension, and its value depends on the nanomaterial properties, the solution conditions, and the theoretical model applied. The ability to use zeta-potential as an explanatory variable for measured behaviors in different systems (or potentially to predict specific behaviors) therefore requires robust reporting with relevant meta-data for the measurement conditions and the model used to convert mobility measurements to zeta-potentials. However, there is currently no such standardization for reporting in the nanoEHS literature. The objective of this tutorial review is to familiarize the nanoEHS research community with the zeta-potential concept and the factors that influence its calculated value and interpretation, including the effects of adsorbed macromolecules. We also provide practical guidance on the precision of measurement, interpretation of zeta-potential as an explanatory variable for processes of interest (e.g. toxicity, environmental fate), and provide advice for addressing common challenges associated with making meaningful zeta-potential measurements using commercial instruments. Finally, we provide specific guidance on the parameters that need to be reported with zetapotential measurements to maximize interpretability and to support scientific synthesis across data sets.