Cognitive neuroscience routinely collects large datasets, yet data integration, curation, version control, and publishing are rarely automated. Here, we present such a workflow, offering open science by design, with an implementation in the Python based `LSLAutoBIDS` open-source package. We first describe our exemplary workflow based on LabStreamingLayer (to integrate), BIDS (to transform), DataLad (to version), and Dataverse (to publish), before discussing the place such tools can have in future data collection efforts.
High-dimensional datasets are becoming an increasingly vital asset in the machine learning and AI domains due to their ability to capture large volumes of structured, self-descriptive information. Research data repositories play a key role in supporting the documentation, discoverability, and reuse of such data. This paper highlights the growing presence of high-dimensional formats—particularly NetCDF and HDF5—underscoring the need for better infrastructure to support them. We use a large language model to assess the quality of metadata associated with these files and find that user-provided metadata often falls short when compared to the richness of embedded metadata within the files themselves. In response, we implement an automated metadata extraction process during file ingestion, offering a practical pathway to FAIR-ify high-dimensional data. Our empirical analysis and technical solution are demonstrated through integration with the Dataverse data repository platform.
Biocatalysis needs improved reproducibility and quality of research reporting. Our interdisciplinary team has developed a flexible and extensible metadata catalogue based on STRENDA guidelines, essential for describing complex experimental setups in biocatalysis. The catalogue is available online via GitHub for community use.
Characterizing the dependence of the thermophysical properties of complex liquid mixtures on parameters such as composition and temperature is pivotal to the choice of an optimal solvent in process engineering. Therefore, it is indispensable to perform comprehensive parameter studies for the exploration of design space. Molecular simulation is a powerful tool for the prediction of properties under conditions that have not yet been explored experimentally. However, simulation results have to be calibrated with published experimental data. In order to make experimental and simulated data available to a comprehensive analysis, we developed a data management and analysis platform based on the standard data exchange format ThermoML. The practicability of integrating thermophysical data from experiment and simulation was demonstrated for two binary mixtures, methanol-water and glycerol-water, by systematically studying the dependence of densities and diffusion coefficients from water content and temperature. Experimental data was extracted manually from literature. The same parameter space was explored by comprehensive molecular dynamics simulations, whose results were directly transferred to the analysis platform. The usefulness of data integration was illustrated by assessing the transferability of the force fields, which had been developed for pure compounds at a specific temperature to different compositions and temperatures, and by analyzing the excess mixing properties as a measure of non-ideality of methanol-water and glycerol-water mixtures. The core of the data management and analysis platform is the newly developed Python library pyThermoML, which represents metadata, the parameters and the experimentally determined or simulated properties as Python data classes. The feasibility of a seamless data flow from data acquisition to a comprehensive data analysis and publication on Dataverse was demonstrated. Because the Dataverse datasets are in ThermoML format, the data is findable, accessible, interoperable, and reusable (FAIR).
Research Data Management (RDM) has gained significant traction in recent years, being essential to allowing research data to be, e.g., findable, accessible, interoperable, and reproducible (FAIR), thereby fostering collaboration or accelerating scientific findings. We present solutions for RDM developed within the DFG-Funded Cluster of Excellence EXC2075 Data-Integrated Simulation Science (SimTech). After an introduction to the scientific context and challenges faced by simulation scientists, we outline the general data management infrastructure and present tools that address these challenges. Exemplary domain applications demonstrate the use and benefits of the proposed data management software solutions. These are complemented by additional measures for enablement and dissemination to foster the adoption of these techniques.
The design of biocatalytic reaction systems is highly complex owing to the dependency of the estimated kinetic parameters on the enzyme, the reaction conditions, and the modeling method. Consequently, reproducibility of enzymatic experiments and reusability of enzymatic data are challenging. We developed the XML-based markup language EnzymeML to enable storage and exchange of enzymatic data such as reaction conditions, the time course of the substrate and the product, kinetic parameters and the kinetic model, thus making enzymatic data findable, accessible, interoperable and reusable (FAIR). The feasibility and usefulness of the EnzymeML toolbox is demonstrated in six scenarios, for which data and metadata of different enzymatic reactions are collected and analyzed. EnzymeML serves as a seamless communication channel between experimental platforms, electronic lab notebooks, tools for modeling of enzyme kinetics, publication platforms and enzymatic reaction databases. EnzymeML is open and transparent, and invites the community to contribute. All documents and codes are freely available at https://enzymeml.org .
A modular research data management toolbox based on the programming language Python, the widely used computing platform Jupyter Notebook, the standardized data exchange format for analytical data (AnIML) and the generic repository Dataverse has been established and applied to analyze small-angle X-ray scattering (SAXS) data according to the FAIR data principles (findable, accessible, interoperable and reusable). The SAS-tools library is a community-driven effort to develop tools for data acquisition, analysis, visualization and publishing of SAXS data. Metadata from the experiment and the results of data analysis are stored as an AnIML document using the novel Python-native pyAnIML API. The AnIML document, measured raw data and plots resulting from the analysis are combined into an archive in OMEX format and uploaded to Dataverse using the novel easyDataverse API, which makes each data set accessible via a unique DOI and searchable via a structured metadata block. SAS-tools is applied to study the effects of alkyl chain length and counterions on the phase diagrams of alkyltrimethyl-ammonium surfactants in order to demonstrate the feasibility and usefulness of a scalable data management workflow for experiments in physical chemistry.
Wir erkennen die Wichtigkeit von Forschungsdaten und -software für unsere Forschungsprozesse an und ordnen die Veröffentlichung von Forschungsdaten und -software als wesentlichen Bestandteil der wissenschaftlichen Publikationstätigkeit ein. Dafür unterstützen wir als Verbund unsere Forschenden im Umgang mit Daten und Software nach den FAIR-Prinzipien in Einvernehmen mit dem DFG-Kodex “Leitlinien zur Sicherung guter wissenschaftlicher Praxis”. In Zusammenarbeit mit unseren Institutionen und Fachcommunities stellen wir adäquate Forschungsdatenmanagement-Werkzeuge und -Dienste bereit und befähigen unsere Forschenden zum Umgang damit. Dabei bauen wir vorzugsweise auf existierenden Angeboten auf und bemühen uns im Gegenzug um deren Anpassung an unsere Bedürfnisse. Wir streben Maßnahmen für die Definition und Sicherstellung der Qualität unserer Forschungsdaten und -software an. Wir verwenden vorzugsweise existierende Daten-/Metadatenstandards und vernetzen uns nach Möglichkeit für die Erstellung und Implementierung neuer Standards mit entsprechenden nationalen und internationalen Initiativen. Wir verfolgen die Entwicklungen im Bereich des Forschungsdaten- und -softwaremanagements und prüfen neu entstehende Empfehlungen und Richtlinien zeitnah auf ihre Umsetzbarkeit.
EnzymeML is an XML–based data exchange format that supports the comprehensive documentation of enzymatic data by describing reaction conditions, time courses of substrate and product concentrations, the kinetic model, and the estimated kinetic constants. EnzymeML is based on the Systems Biology Markup Language, which was extended by implementing the STRENDA Guidelines. An EnzymeML document serves as a container to transfer data between experimental platforms, modelling tools, and databases. EnzymeML supports the scientific community by introducing a standardised data exchange format to make enzymatic data findable, accessible, interoperable, and reusable according to the FAIR data principles. An Application Programming Interface in Python and Java supports the integration of applications. The feasibility of a seamless data flow using EnzymeML is demonstrated by creating an EnzymeML document from a structured spreadsheet or from a STRENDA DB database entry, by kinetic modelling using the modelling platform COPASI, and by uploading to the enzymatic reaction kinetics database SABIO-RK.
In order to make thermophysical properties of complex liquid mixtures available to a comprehensive analysis, we developed a data management and analysis platform based on the standard data exchange format ThermoML. The practicability of integrating thermophysical data from experiments and simulations was demonstrated for two binary mixtures, methanol-water and glycerol-water, by systematically studying the dependence of densities and diffusion coefficients from water content over the whole composition range and temperatures between 278.15 and 318.15 K. Experimental data were extracted manually from the literature. The same parameter space was explored by comprehensive molecular dynamics simulations, whose results were directly transferred to the analysis platform. The benefit of data integration was illustrated by assessing the transferability of the force fields, which had been developed for pure compounds to different compositions and temperatures, and by analyzing the excess mixing properties as a measure of nonideality of methanol-water and glycerol-water mixtures. The core of the data management and analysis platform is the newly developed Python library pyThermoML, which represents metadata, the parameters, and the experimentally determined or simulated properties as Python data classes. The feasibility of a seamless data flow from data acquisition to a comprehensive data analysis was demonstrated. PyThermoML enables interoperability and reusability of the datasets. The publication of ThermoML documents on the Dataverse installation of the University of Stuttgart (DaRUS) makes thermophysical data findable and accessible and thus FAIR.
Experimental data on thermophysical properties of solvent mixtures such as aqueous deep eutectic solvents (DES) are scattered in scientific publications. It is desirable to integrate thermophysical properties with parameters that describe a DES in a human and machine readable format. On the basis of the Chemical Markup Language (CML), a standardized exchange format for density, viscosity, conductivity, and water activity of solvent mixtures was established and applied to represent published data on choline chloride/glycerol/water mixtures. In total, 300 different data sets served as a basis for data analysis by machine learning. Gradient tree boosting (GB) was used to predict thermophysical properties from the collected parameters, resulting in an excellent correlation between predicted, experimental, and simulation data. The experimental viscosity data was modeled assuming an Arrhenius dependency on temperature and by determining two parameters (eta(0) and E-eta). Integration of experimental and simulation data into a standardized exchange format makes data findable, accessible, interoperable, and reusable and enables machine learning methods. To facilitate data exchange, we recommend researchers publish experimental and simulated data on thermophysical properties as a CML-formatted Supporting Information with an associated digital object identifier (DOI).