The lack of machine-readable experimental data impedes data-driven discoveries in catalysis research. To advance the FAIR principles — guidelines to improve the Findability, Accessibility, Interoperability and Reuse of digital assets — we have developed a catalysis plugin (called the Catalysis App) for the NOMAD platform that supports standardized data upload and features integrated visualization. This infrastructure provides a robust foundation for machine-learning workflows and the direct comparison of experimental data with theory.
The field of catalysis currently lacks a structured repository for experimental data, and the publication of machine-readable datasets remains uncommon. To address this gap and support FAIR (Findable, Accessible, Interoperable, and Reusable) data principles, we introduce a new plugin within the NOMAD platform for managing and publishing heterogeneous catalysis data, the nomad_catalysis plugin. This plugin enables the upload of structured experimental data and metadata with built-in visualization and alignment to the community-developed vocabulary Voc4Cat, ensuring long-term interpretability. In addition to facilitating efficient data sharing, the catalysis app offers intuitive search functionality, enabling researchers to quickly identify relevant catalytic reactions, catalyst materials, reaction conditions and kinetic properties. This infrastructure lays the foundation for advanced data analytics and machine learning applications, supporting more efficient and reproducible catalyst development.
Scientific data across physics, materials science, and materials engineering often lacks adherence to FAIR principles (Barker et al., 2022; Jacobsen et al., 2020; M. D. Wilkinson et al., 2016; S. R. Wilkinson et al., 2025) due to incompatible instrument-specific formats and diverse standardization practices. pynxtools is a Python software development framework with a command line interface (CLI) that standardizes data conversion for scientific experiments in materials science to the NeXus format (Klosowski et al., 1997; Könnecke, 2006; Könnecke et al., 2015) across diverse scientific domains. NeXus defines data storage specifications for different experimental techniques through application definitions. pynxtools provides a fixed, versioned set of NeXus application definitions that ensures convergence and alignment in data specifications across, among others, atom probe tomography, electron microscopy, optical spectroscopy, photoemission spectroscopy, scanning probe microscopy, and X-ray diffraction. Through its modular plugin architecture pynxtools provides conversion of data and metadata from instruments and electronic lab notebooks to these unified definitions, while performing validation to ensure data correctness and NeXus compliance. pynxtools can be integrated directly into Research Data Management Systems (RDMS) to facilitate parsing and normalization. We detail one example for the RDM system NOMAD. By simplifying the adoption of NeXus, the framework enables true data interoperability and FAIR data management across multiple experimental techniques.
With the rapidly increasing amount of materials data being generated in a variety of projects, efficient and accurate classification of atomistic structures is essential. A current barrier to effective database queries lies in the often ambiguous, inconsistent, or completely missing classification of existing data, highlighting the need for standardized, automated, and verifiable classification methods. This work proposes a robust solution for identifying and classifying a wide spectrum of materials through an iterative technique, called symmetry-based clustering (SBC). Because SBC is not a machine learning-based method, it requires no prior training. Instead, it identifies clusters in atomistic systems by automatically recognizing common unit cells. We demonstrate the potential of SBC to provide automated, reliable classification and to reveal well-known symmetry properties of various materials. Even noisy systems are shown to be classifiable, showing the suitability of our algorithm for real-world data applications. The software implementation is provided in the open-source Python package, MatID, exploiting synergies with popular atomic-structure manipulation libraries and extending the accessibility of those libraries through the NOMAD platform.
Science is and always has been based on data, but the terms "data-centric" and the "4th paradigm of" materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of Artificial Intelligence (AI) and its subset Machine Learning (ML), has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.
Here, we present the outcomes from the second Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry, which engaged participants across global hybrid locations, resulting in 34 team submissions. The submissions spanned seven key application areas and demonstrated the diverse utility of LLMs for applications in (1) molecular and material property prediction; (2) molecular and material design; (3) automation and novel interfaces; (4) scientific communication and education; (5) research data management and automation; (6) hypothesis generation and evaluation; and (7) knowledge extraction and reasoning from scientific literature. Each team submission is presented in a summary table with links to the code and as brief papers in the appendix. Beyond team results, we discuss the hackathon event and its hybrid format, which included physical hubs in Toronto, Montreal, San Francisco, Berlin, Lausanne, and Tokyo, alongside a global online hub to enable local and virtual collaboration. Overall, the event highlighted significant improvements in LLM capabilities since the previous year's hackathon, suggesting continued expansion of LLMs for applications in materials science and chemistry research. These outcomes demonstrate the dual utility of LLMs as both multipurpose models for diverse machine learning tasks and platforms for rapid prototyping custom applications in scientific research.
The Open Databases Integration for Materials Design (OPTIMADE) application programming interface (API) empowers users with holistic access to a growing federation of databases, enhancing the accessibility and discoverability of materials and chemical data. Since the first release of the OPTIMADE specification (v1.0), the API has undergone significant development, leading to the upcoming v1.2 release, and has underpinned multiple scientific studies. In this work, we highlight the latest features of the API format, accompanying software tools, and provide an update on the implementation of OPTIMADE in contributing materials databases. We end by providing several use cases that demonstrate the utility of the OPTIMADE API in materials research that continue to drive its ongoing development.
Scientific research is becoming increasingly data centric, which requires more effort to manage, share, and publish data.NOMAD is a web-based platform that provides research data management (RDM) for materials-science data. In addition to core RDM functions like uploading and sharing files, NOMAD automatically extracts structured data from supported file formats, normalizes, and converts data from these formats. NOMAD provides an extendable framework for managing not just files, but structured machine-actionable harmonized and inter-operable data. This is the basis for a faceted search with domain-specific filters, a comprehensive API, structured data entry via customizable ELNs, integrated data-analysis and machine-learning tools. NOMAD is run as a free public service and can additionally be operated by research institutes. Connecting NOMAD installations through the public services will allow a federated data infrastructure to share data between research institutes and further harmonize RDM within a large research domain such as materials science.
Scientific research is becoming increasingly data centric, which requires more effort to manage, share, and publish data.NOMAD is a web-based platform that provides research data management (RDM) for materials-science data. In addition to core RDM functions like uploading and sharing files, NOMAD automatically extracts structured data from supported file formats, normalizes, and converts data from these formats. NOMAD provides an extendable framework for managing not just files, but structured machine-actionable harmonized and inter-operable data. This is the basis for a faceted search with domain-specific filters, a comprehensive API, structured data entry via customizable ELNs, integrated data-analysis and machine-learning tools. NOMAD is run as a free public service and can additionally be operated by research institutes. Connecting NOMAD installations through the public services will allow a federated data infrastructure to share data between research institutes and further harmonize RDM within a large research domain such as materials science.
The expansive production of data in materials science, their widespread sharing and repurposing requires educated support and stewardship. In order to ensure that this need helps rather than hinders scientific work, the implementation of the FAIR-data principles ( Findable, Accessible, Interoperable, and Reusable ) must not be too narrow. Besides, the wider materials-science community ought to agree on the strategies to tackle the challenges that are specific to its data, both from computations and experiments. In this paper, we present the result of the discussions held at the workshop on “Shared Metadata and Data Formats for Big-Data Driven Materials Science”. We start from an operative definition of metadata, and the features that a FAIR-compliant metadata schema should have. We will mainly focus on computational materials-science data and propose a constructive approach for the FAIRification of the (meta)data related to ground-state and excited-states calculations, potential-energy sampling, and generalized workflows. Finally, challenges with the FAIRification of experimental (meta)data and materials-science ontologies are presented together with an outlook of how to meet them.
Markus Scheidgen 1*¶, Lauri Himanen 1*, Alvin Noe Ladines 1*, David Sikter 1*, Mohammad Nakhaee 1*, Ádám Fekete 1*, Theodore Chang 1*, Amir Golparvar 1*, José A. Márquez 1, Sandor Brockhauser 1, Sebastian Brückner 2, Luca M. Ghiringhelli 1, Felix Dietrich 3, Daniel Lehmberg 3, Thea Denell 1, Andrea Albino 1, Hampus Näsström 1, Sherjeel Shabih 1, Florian Dobener 1, Markus Kühbach 1, Rubel Mozumder 1, Joseph F. Rudzinski 1, Nathan Daelman 1, José M. Pizarro 1, Martin Kuban 1, Cuauhtemoc Salazar 1, Pavel Ondračka 4, Hans-Joachim Bungartz 3, and Claudia Draxl 1
We develop a materials descriptor based on the electronic density-of-states (DOS) and investigate the similarity of materials based on it. As an application example, we study the Computational 2D Materials Database (C2DB) that hosts thousands of two-dimensional materials with their properties calculated by density-functional theory. Combining our descriptor with a clustering algorithm, we identify groups of materials with similar electronic structure. We introduce additional descriptors to characterize these clusters in terms of crystal structures, atomic compositions, and electronic configurations of their members. This allows us to rationalize the found (dis)similarities and to perform an automated exploratory and confirmatory analysis of the C2DB data. From this analysis, we find that the majority of clusters consist of isoelectronic materials sharing crystal symmetry, but we also identify outliers, i.e., materials whose similarity cannot be explained in this way.
The prosperity and lifestyle of our society are very much governed by achievements in condensed matter physics, chemistry and materials science, because new products for sectors such as energy, the environment, health, mobility and information technology (IT) rely largely on improved or even new materials. Examples include solid-state lighting, touchscreens, batteries, implants, drug delivery and many more. The enormous amount of research data produced every day in these fields represents a gold mine of the twenty-first century. This gold mine is, however, of little value if these data are not comprehensively characterized and made available. How can we refine this feedstock; that is, turn data into knowledge and value? For this, a FAIR (findable, accessible, interoperable and reusable) data infrastructure is a must. Only then can data be readily shared and explored using data analytics and artificial intelligence (AI) methods. Making data 'findable and AI ready' (a forward-looking interpretation of the acronym) will change the way in which science is carried out today. In this Perspective, we discuss how we can prepare to make this happen for the field of materials science.
Journal Article Development of a FAIR Data Management Infrastructure Get access Sherjeel Shabih, Sherjeel Shabih Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Corresponding author: sherjeel.shabih@hu-berlin.de Search for other works by this author on: Oxford Academic Google Scholar Markus Kühbach, Markus Kühbach Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Markus Scheidgen, Markus Scheidgen Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Lauri Himanen, Lauri Himanen Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Sandor Brockhauser, Sandor Brockhauser Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Benedikt Haas, Benedikt Haas Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Christoph Koch Christoph Koch Humboldt Universität zu Berlin, Institut für Physik & IRIS, Adlershof, Berlin, Germany Search for other works by this author on: Oxford Academic Google Scholar Microscopy and Microanalysis, Volume 28, Issue S1, 1 August 2022, Pages 2930–2932, https://doi.org/10.1017/S1431927622010996 Published: 01 August 2022
In recent years, we have been witnessing a paradigm shift in computational materials science. In fact, traditional methods, mostly developed in the second half of the XXth century, are being complemented, extended, and sometimes even completely replaced by faster, simpler, and often more accurate approaches. The new approaches, that we collectively label by machine learning, have their origins in the fields of informatics and artificial intelligence, but are making rapid inroads in all other branches of science. With this in mind, this Roadmap article, consisting of multiple contributions from experts across the field, discusses the use of machine learning in materials science, and share perspectives on current and future challenges in problems as diverse as the prediction of materials properties, the construction of force-fields, the development of exchange correlation functionals for density-functional theory, the solution of the many-body problem, and more. In spite of the already numerous and exciting success stories, we are just at the beginning of a long path that will reshape materials science for the many challenges of the XXIth century.
The Open Databases Integration for Materials Design (OPTIMADE) consortium has designed a universal application programming interface (API) to make materials databases accessible and interoperable. We outline the first stable release of the specification, v1.0, which is already supported by many leading databases and several software packages. We illustrate the advantages of the OPTIMADE API through worked examples on each of the public materials databases that support the full API specification.
1 Institut de la Matière Condensée et des Nanosciences, Université catholique de Louvain, Chemin des Étoiles 8, Louvain-la-Neuve 1348, Belgium 2 Theory of Condensed Matter Group, Cavendish Laboratory, University of Cambridge, J. J. Thomson Avenue, Cambridge, CB3 0HE, United Kingdom 3 Theory and Simulation of Materials (THEOS), Faculté des Sciences et Techniques de l’Ingénieur, École Polytechnique Fédérale de Lausanne, CH-1015 Lausanne, Switzerland 4 Lawrence Berkeley National Laboratory, Berkeley, CA, USA 5 Fritz-Haber-Institut der Max-Planck-Gesellschaft, Faradayweg 4-6, 14195, Berlin, Germany 6 Humboldt-Universität zu Berlin, Institut für Physik and IRIS Adlershof, 12489 Berlin, Germany 7 Polyneme LLC, New York, NY, USA 8 Department of Physics, King’s College London, Strand, London WC2R 2LS, United Kingdom 9 Department of Physics and Namur Institute of Structured Materials, University of Namur, Rue de Bruxelles 51, 5000 Namur, Belgium DOI: 10.21105/joss.03458