Scientific data across physics, materials science, and materials engineering often lacks adherence to FAIR principles (Barker et al., 2022; Jacobsen et al., 2020; M. D. Wilkinson et al., 2016; S. R. Wilkinson et al., 2025) due to incompatible instrument-specific formats and diverse standardization practices. pynxtools is a Python software development framework with a command line interface (CLI) that standardizes data conversion for scientific experiments in materials science to the NeXus format (Klosowski et al., 1997; Könnecke, 2006; Könnecke et al., 2015) across diverse scientific domains. NeXus defines data storage specifications for different experimental techniques through application definitions. pynxtools provides a fixed, versioned set of NeXus application definitions that ensures convergence and alignment in data specifications across, among others, atom probe tomography, electron microscopy, optical spectroscopy, photoemission spectroscopy, scanning probe microscopy, and X-ray diffraction. Through its modular plugin architecture pynxtools provides conversion of data and metadata from instruments and electronic lab notebooks to these unified definitions, while performing validation to ensure data correctness and NeXus compliance. pynxtools can be integrated directly into Research Data Management Systems (RDMS) to facilitate parsing and normalization. We detail one example for the RDM system NOMAD. By simplifying the adoption of NeXus, the framework enables true data interoperability and FAIR data management across multiple experimental techniques.
FAIRmat develops concepts for paving the way to enable FAIR research data in solid-state physics. For selected theoretical data in this field, the NOMAD portal has developed mature concepts and technological solutions for storing data according to the FAIR principles. Extending this approach to experimental data is challenging due to their diversity and missing standards. In this paper we present our comprehensive approach to establish FAIR data in the field of experimental solid-state physics despite its heterogeneity. The concept includes elaboration of standards, community building and methods that facilitate the community’s transition to FAIR standards.
Spectroscopy and X-ray diffraction techniques encode ample information on investigated samples. The ability of rapidly and accurately extracting these enhances the means to steer the experiment, as well as the understanding of the underlying processes governing the experiment. It improves the efficiency of the experiment, and maximizes the scientific outcome. To address this, we introduce and validate three frameworks based on self-supervised learning which are capable of classifying 1D spectral curves using data transformations preserving the scientific content and only a small amount of data labeled by domain experts. In particular, in this work we focus on the identification of phase transitions in samples investigated by x-ray powder diffraction. We demonstrate that the three frameworks, based either on relational reasoning, contrastive learning, or a combination of the two, are capable of accurately identifying phase transitions. Furthermore, we discuss in detail the selection of data augmentation techniques, crucial to ensure that scientifically meaningful information is retained.
Scientific research is becoming increasingly data centric, which requires more effort to manage, share, and publish data.NOMAD is a web-based platform that provides research data management (RDM) for materials-science data. In addition to core RDM functions like uploading and sharing files, NOMAD automatically extracts structured data from supported file formats, normalizes, and converts data from these formats. NOMAD provides an extendable framework for managing not just files, but structured machine-actionable harmonized and inter-operable data. This is the basis for a faceted search with domain-specific filters, a comprehensive API, structured data entry via customizable ELNs, integrated data-analysis and machine-learning tools. NOMAD is run as a free public service and can additionally be operated by research institutes. Connecting NOMAD installations through the public services will allow a federated data infrastructure to share data between research institutes and further harmonize RDM within a large research domain such as materials science.
The expansive production of data in materials science, their widespread sharing and repurposing requires educated support and stewardship. In order to ensure that this need helps rather than hinders scientific work, the implementation of the FAIR-data principles ( Findable, Accessible, Interoperable, and Reusable ) must not be too narrow. Besides, the wider materials-science community ought to agree on the strategies to tackle the challenges that are specific to its data, both from computations and experiments. In this paper, we present the result of the discussions held at the workshop on “Shared Metadata and Data Formats for Big-Data Driven Materials Science”. We start from an operative definition of metadata, and the features that a FAIR-compliant metadata schema should have. We will mainly focus on computational materials-science data and propose a constructive approach for the FAIRification of the (meta)data related to ground-state and excited-states calculations, potential-energy sampling, and generalized workflows. Finally, challenges with the FAIRification of experimental (meta)data and materials-science ontologies are presented together with an outlook of how to meet them.
Markus Scheidgen 1*¶, Lauri Himanen 1*, Alvin Noe Ladines 1*, David Sikter 1*, Mohammad Nakhaee 1*, Ádám Fekete 1*, Theodore Chang 1*, Amir Golparvar 1*, José A. Márquez 1, Sandor Brockhauser 1, Sebastian Brückner 2, Luca M. Ghiringhelli 1, Felix Dietrich 3, Daniel Lehmberg 3, Thea Denell 1, Andrea Albino 1, Hampus Näsström 1, Sherjeel Shabih 1, Florian Dobener 1, Markus Kühbach 1, Rubel Mozumder 1, Joseph F. Rudzinski 1, Nathan Daelman 1, José M. Pizarro 1, Martin Kuban 1, Cuauhtemoc Salazar 1, Pavel Ondračka 4, Hans-Joachim Bungartz 3, and Claudia Draxl 1
Characterizing microstructure-material-property relations calls for software tools which extract point-cloud- and continuum-scale-based representations of microstructural objects. Application examples include atom probe, electron, and computational microscopy experiments. Mapping between atomic- and continuum-scale representations of microstructural objects results often in representations which are sensitive to parameterization; however assessing this sensitivity is a tedious task in practice. Here, we show how combining methods from computational geometry, collision analyses, and graph analytics yield software tools for automated analyses of point cloud data for reconstruction of three-dimensional objects, characterization of composition profiles, and extraction of multi-parameter correlations via evaluating graph-based relations between sets of meshed objects. Implemented for point clouds with mark data, we discuss use cases in atom probe microscopy that focus on interfaces, precipitates, and coprecipitation phenomena observed in different alloys. The methods are expandable for spatio-temporal analyses of grain fragmentation, crystal growth, or precipitation.
Journal Article Development of a FAIR Data Management Infrastructure Get access Sherjeel Shabih, Sherjeel Shabih Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Corresponding author: sherjeel.shabih@hu-berlin.de Search for other works by this author on: Oxford Academic Google Scholar Markus Kühbach, Markus Kühbach Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Markus Scheidgen, Markus Scheidgen Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Lauri Himanen, Lauri Himanen Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Sandor Brockhauser, Sandor Brockhauser Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Benedikt Haas, Benedikt Haas Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Christoph Koch Christoph Koch Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Microscopy and Microanalysis, Volume 28, Issue S1, 1 August 2022, Pages 2930–2932, https://doi.org/10.1017/S1431927622010996 Published: 01 August 2022
The European X-ray Free Electron Laser (XFEL) and Linac Coherent Light Source (LCLS) II are extremely intense sources of X-rays capable of generating Serial Femtosecond Crystallography (SFX) data at megahertz (MHz) repetition rates. Previous work has shown that it is possible to use consecutive X-ray pulses to collect diffraction patterns from individual crystals. Here, we exploit the MHz pulse structure of the European XFEL to obtain two complete datasets from the same lysozyme crystal, first hit and the second hit, before it exits the beam. The two datasets, separated by <1 µs, yield up to 2.1 Å resolution structures. Comparisons between the two structures reveal no indications of radiation damage or significant changes within the active site, consistent with the calculated dose estimates. This demonstrates MHz SFX can be used as a tool for tracking sub-microsecond structural changes in individual single crystals, a technique we refer to as multi-hit SFX.
Spectroscopy experiment techniques are widely used and produce a huge amount of data especially in facilities with very high repetition rates. In High Energy Density (HED) experiments with high-density materials, changes in pressure will cause changes in the spectral peak. Immediate feedback on the actual status (e.g. time-resolved status of the sample) would be essential to quickly judge how to proceed with the experiment. The two major spectral changes we aim to capture are either the change of intensity distribution (e.g., drop or appearance) of peaks at certain locations, or the shift of those on the spectrum. In this work, we apply recent popular machine learning/deep learning models to HED experimental spectra data classification. The models we presented range from supervised deep neural networks (state-of-the-art LSTM-based model and Transformer-based model) to unsupervised spectral clustering algorithm. These are the common architectures for time series processing. The PCA method is used as data preprocessing for dimensionality reduction. Three different ML algorithms are evaluated and compared for the classification task. The results show that all three methods can achieve 100
In scientific research, spectroscopy and diffraction experimental techniques are widely used and produce huge amounts of spectral data. Learning patterns from spectra is critical during these experiments. This provides immediate feedback on the actual status of the experiment (e.g., time-resolved status of the sample), which helps guide the experiment. The two major spectral changes what we aim to capture are either the change in intensity distribution (e.g., drop or appearance) of peaks at certain locations, or the shift of those on the spectrum. This study aims to develop deep learning (DL) classification frameworks for one-dimensional (1D) spectral time series. In this work, we deal with the spectra classification problem from two different perspectives, one is a general two-dimensional (2D) space segmentation problem, and the other is a common 1D time series classification problem. We focused on the two proposed classification models under these two settings, the namely the end-to-end binned Fully Connected Neural Network (FCNN) with the automatically capturing weighting factors model and the convolutional SCT attention model. Under the setting of 1D time series classification, several other end-to-end structures based on FCNN, Convolutional Neural Network (CNN), ResNets, Long Short-Term Memory (LSTM), and Transformer were explored. Finally, we evaluated and compared the performance of these classification models based on the High Energy Density (HED) spectra dataset from multiple perspectives, and further performed the feature importance analysis to explore their interpretability. The results show that all the applied models can achieve 100% classification confidence, but the models applied under the 1D time series classification setting are superior. Among them, Transformer-based methods consume the least training time (0.449 s). Our proposed convolutional Spatial-Channel-Temporal (SCT) attention model uses 1.269 s, but its self-attention mechanism performed across spatial, channel, and temporal dimensions can suppress indistinguishable features better than others, and selectively focus on obvious features with high separability.
The European XFEL is a hard X-ray free-electron laser (FEL) based on a high-electron-energy superconducting linear accelerator. The superconducting technology allows for the acceleration of many electron bunches within one radio-frequency pulse of the accelerating voltage and, in turn, for the generation of a large number of hard X-ray pulses. We report on the performance of the European XFEL accelerator with up to 5,000 electron bunches per second and demonstrating a full energy of 17.5 GeV. Feedback mechanisms enable stabilization of the electron beam delivery at the FEL undulator in space and time. The measured FEL gain curve at 9.3 keV is in good agreement with predictions for saturated FEL radiation. Hard X-ray lasing was achieved between 7 keV and 14 keV with pulse energies of up to 2.0 mJ. Using the high repetition rate, an FEL beam with 6 W average power was created. The first operation of the European X-ray free-electron laser facility accelerator based on superconducting technology is reported. The maximum electron energy is 17.5 GeV. A laser average power of 6 W is achieved at a photon energy of 9.3 keV.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Nowadays there is a growing need for user friendly workflow editors in all fields of scientific research. A special interest group is present at big physics research facilities where instrumentation is mostly controlled by a robust and reliable low level control software solution. Different types of specific experiments using predetermined automated protocols and on-line data processing with real-time feedback require a more flexible and abstract high level control system[1]. Beside flexibility and dynamism, easy usability is also required for researchers collaborating from several different fields. Tentatively, to test the ease and flexible usability, the Kepler workflowengine was integrated with TANGO[2]. It enables researchers to automate and document experiment protocols without any programming skill. The X-ray crystallography laboratory at the Biological Research Center of Hungarian Academy of Science (BRC) has implemented an example crystallographic workflow to test the integrated system. This development was performed in cooperation with ELI-ALPS.
We provide a detailed description of a serial femtosecond crystallography (SFX) dataset collected at the European X-ray free-electron laser facility (EuXFEL). The EuXFEL is the first high repetition rate XFEL delivering MHz X-ray pulse trains at 10 Hz. The short spacing (<1 µs) between pulses requires fast flowing microjets for sample injection and high frame rate detectors. A data set was recorded of a microcrystalline mixture of at least three different jack bean proteins (urease, concanavalin A, concanavalin B). A one megapixel Adaptive Gain Integrating Pixel Detector (AGIPD) was used which has not only a high frame rate but also a large dynamic range. This dataset is publicly available through the Coherent X-ray Imaging Data Bank (CXIDB) as a resource for algorithm development and for data analysis training for prospective XFEL users.
EXPERIMENTS AT THE EUXFEL Dall'Antonia, Fabio; Beg, Marijan; Bergemann, Martin; Bielecki, Johan; Bondar, Valerii; Carinan, Cammille; Costa Junior, Raul; Danilevski, Cyril; Ehsan, Wajid; Esenov, Sergey; Fabbri, Riccardo; Flucke, Gero; Fullà-Marsà, Daniel; Giovanetti, Gabriele; Göries, Dennis; Hickin, David; Jarosiewicz, Tobiasz; Kamil, Ebad; Kirienko, Yury; Kirkwood, Henry; Klimovskaia, Anna; Kluyver, Thomas; Mamchyk, Denys; Michelat, Thomas; Mohacsi, Istvan; Parenti, Andrea; Rosca, Robert; Rück, Denivy; Santos, Hugo; Schaffer, Robert; Silenzi, Alessandro; Spirzewski, Michal; Trojanowski, Sebastian; Youngman, Christopher; Zhu, Jun; Mancuso, Adrian P.; Fangohr, Hans; Brockhauser, Sandor: European XFEL GmbH, Schenefeld, GER
Jupyter notebooks are executable documents that are displayed in a web browser. The notebook elements consist of human-authored contextual elements and computer code, and computer-generated output from executing the computer code. Such outputs can include tables and plots. The notebook elements can be executed interactively, and the whole notebook can be saved, re-loaded and re-executed, or converted to read-only formats such as HTML, LaTeX and PDF. Exploiting these characteristics, Jupyter notebooks can be used to improve the effectiveness of computational and data exploration, documentation, communication, reproducibility and re-usability of scientific research results. They also serve as building blocks of remote data access and analysis as is required for facilities hosting large data sets and initiatives such as the European Open Science Cloud (EOSC). In this contribution we report from our experience of using Jupyter notebooks for data analysis at research facilities, and outline opportunities and future plans.