Every decision made during a machine learning pipeline has an impact on the outcome. Feature selection can reduce overfitting and focus models on the attributes that matter most, and sample selection can reduce bias to ensure models recognise patterns comprehensively. eXplainable AI (XAI) can provide quantitative ways of evaluating the impact of these decisions, and help ensure the right data is used for training models predicting structure property relationships. In this paper we explore the use of residual decomposition with Shapely values to identify which nanoparticle shapes are most influential in predicting charge transfer properties of gold nanoparticles and how they impact the ability to predict the properties of the different morphologies.
Explainable artificial intelligence methods are increasingly essential for extracting physical insight from high-throughput computational materials data. We apply the Shapley behavioral transformation framework to two MXene datasets to characterize how individual compositions contribute to dataset statistics. By decomposing variance, skewness, kurtosis, and entropy using Shapley values, we create four complementary “behavioral spaces” that reveal compositional patterns invisible in raw feature representations. All behavioral spaces exhibit strong clustering tendency, with the skewness space providing the clearest chemical interpretation. Regional analysis in skewness space identifies six distinct behavioral zones with characteristic property distributions, including a fluorine-terminated MXenes cluster with elevated intercalation voltages and deep valence band positions. Critically, no MXenes appear as an outlier in all four behavioral spaces, demonstrating that each statistic captures genuinely complementary aspects of distributional behavior. The maximum cross-space agreement is 3/4 spaces, achieved by Sc2C in the electrochemistry dataset and by eight compositions in the electronic structure dataset. Among these consistent outliers, Y2CBr2 and Hf3C2(NH)2 exhibit conduction band positions favorable for photocatalytic hydrogen evolution. This model-agnostic framework complements predictive machine learning by answering not “what property will this composition have?” but “how does this composition contribute to the statistical structure of the dataset?” The latter question is directly relevant for identifying synthesis priorities and understanding structure–property relationships in complex materials spaces.
We review molecular and polymer representations for machine learning, from descriptors and fingerprints to learned embeddings, highlighting how hierarchy, stochasticity, and data limitations shape polymer informatics.
In cheminformatics, the explainability of machine learning models is important for interpreting complex chemical data, deriving new chemical insights, and building trust in predictive models. However, cheminformatics datasets often exhibit clustered distributions, while traditional explanation methods might overlook intra-cluster variations and complicate the extraction of meaningful explanations. Additionally, diverse representations (tabular, sequence, image, and graph) yield divergent explanations. To address these issues, we propose a novel approach termed regional explanation, designed as an intermediate-level interpretability method that bridges the gap between local and global explanations. This approach systematically reveals how explanations and feature importance vary across data clusters. Using 2 public datasets, a graphene oxide nanoflakes dataset and QM9, with natural clustering properties, we comprehensively evaluate 4 molecular representations through tabular, sequence, image, and graph regional explanation, providing practical guidelines for representation selection. Our analysis illuminates complex, nonlinear relationships between molecular structures and predicted properties within clusters; explores the interplay among molecular features, feature importance, and target properties across distinct regions of chemical space; and advances the interpretability of machine learning models for complex molecular systems.
Artificial intelligence is a powerful tool for materials discovery and design, but like most scientific approaches, there are many methods to choose from, and it is not immediately clear which one is right for a particular problem. Materials design is largely concerned with identifying promising materials with desirable properties, which is a job for machine learning, or identifying promising experimental conditions to make these hypothetical materials a reality, which is a job for statistical learning. In both cases, there are numerous algorithms, each with unique capabilities, advantages and disadvantages. In this review, we examine and compare some useful approaches that are less commonly used in the community; machine learning to select the right materials, and statistical learning to recommend how to make it. We focus on introducing and discussing high‐level approaches that can be used in combination with a wide variety of specific supervised and unsupervised learning algorithms and provide examples of where they have been successfully used in the past to study the structure, properties, and processing of materials. We conclude by showing they can be used together to develop an insightful materials design strategies, backed‐up by informative and actionable experimental tactics.
Machine learning methods have been remarkably successful in material science, providing novel scientific insights, guiding future laboratory experiments, and accelerating materials discovery. Despite the promising performance of these models, understanding the decisions they make is also essential to ensure the scientific value of their outcomes. However, there is a recent and ongoing debate about the diversity of explanations, which potentially leads to scientific inconsistency. This Perspective explores the sources and implications of these diverse explanations in ML applications for physical sciences. Through three case studies in materials science and molecular property prediction, we examine how different models, explanation methods, levels of feature attribution, and stakeholder needs can result in varying interpretations of ML outputs. Our analysis underscores the importance of considering multiple perspectives when interpreting ML models in scientific contexts and highlights the critical need for scientists to maintain control over the interpretation process, balancing data-driven insights with domain expertise to meet specific scientific needs. By fostering a comprehensive understanding of these inconsistencies, we aim to contribute to the responsible integration of eXplainable artificial intelligence into physical sciences and improve the trustworthiness of ML applications in scientific discovery.
Inverse models are a valuable tool for designing new engineering materials based on a set of property requirements. In the case of metal alloys, the properties are also determined by the microstructure, which is a function of the post‐synthesis processing methods as well as the composition. Measuring the microstructure is destructive and labor intensive, so predicting the post‐synthesis processing methods directly has some advantages. However, process‐structure‐property, or inverse property‐structure‐process relationships are challenging for machine learning, due to the scarcity and imbalance of alloy data, and the fact that the mechanical properties of novel alloys are unknown. In this study of Mg alloys, guided oversampling and imputation with the Synthetic Minority Oversampling Technique is used to overcome the former issues, and semi‐supervised learning with label spreading is used to overcome the latter. This resulted in area under the receiver operator characteristic curves exceeding 0.91 for binary classifiers of casting, extrusion, or wrought shaping. The pipeline is general, and the final models can be used to recommend a processing method based on combinations alloying metals and the ductility, yield strength and ultimate tensile strength, in advance of synthesis.
This perspective addresses the topic of harnessing the tools of artificial intelligence (AI) for boosting innovation in functional materials design and engineering as well as discovering new materials for targeted applications in energy storage, biomedicine, composites, nanoelectronics or quantum technologies. It gives a current view of experts in the field, insisting on challenges and opportunities provided by the development of large materials databases, novel schemes for implementing AI into materials production and characterization as well as progress in the quest of simulating physical and chemical properties of realistic atomic models reaching the trillion atoms scale and with near ab initio accuracy.
Explainable artificial intelligence (XAI) techniques are increasingly becoming integrated into many scientific workflows. However, the involvement of such methods typically manifests in the later stages of the workflow such as in model fitting and selection. The pre-processing and exploration stages are equally important when considering the data analysis and scientific workflow as a whole. This paper frames the initial data analysis problem as an XAI task where we seek to find patterns using interpretable transformations of data feature values using Shapley Values. We demonstrate that our methodology ties in with many existing approaches and results that uncover previously unseen patterns in three materials science datasets.
Primary care clinicians play a key role in asthma and asthma exacerbation management worldwide because most patients with asthma are treated in primary care settings. The high burden of asthma exacerbations persists and important practice gaps remain, despite continual advances in asthma care. Lack of primary care-specific guidance, uncontrolled asthma, incomplete assessment of exacerbation and asthma control history, and reliance on systemic corticosteroids or short-acting beta2-agonist-only therapy are challenges clinicians face today with asthma care. Evidence supports the use of inhaled corticosteroids (ICS) + fast-acting bronchodilator treatments when used as needed in response to symptoms to improve asthma control and reduce rates of exacerbations, and the symptoms that occur leading up to an asthma exacerbation provide a window of opportunity to intervene with ICS. Incorporating patient perspectives and preferences when designing asthma regimens will help patients be more engaged in their therapy and may contribute to improved adherence and outcomes. This expert consensus contains 10 Best Practice Advice Points from a panel of primary care clinicians and a patient representative, formed in collaboration with the International Primary Care Respiratory Group (IPCRG), a clinically led charitable organization that works locally and globally in primary care to improve respiratory health. The panel met virtually and developed a series of best practice statements, which were drafted and subsequently voted on to obtain consensus. Primary care clinicians globally are encouraged to review and adapt these best practice advice points on preventing and managing asthma exacerbations to their local practice patterns to enhance asthma care within their practice.
Conflicting explanations, arising from different attribution methods or model internals, limit the adoption of machine learning models in safety-critical domains. We turn this disagreement into an advantage and introduce EXplanation AGREEment (EXAGREE), a two-stage framework that selects a Stakeholder-Aligned Explanation Model (SAEM) from a set of similar-performing models. The selection maximizes Stakeholder-Machine Agreement (SMA), a single metric that unifies faithfulness and plausibility. EXAGREE couples a differentiable mask-based attribution network (DMAN) with monotone differentiable sorting, enabling gradient-based search inside the constrained model space. Experiments on six real-world datasets demonstrate simultaneous gains of faithfulness, plausibility, and fairness over baselines, while preserving task accuracy. Extensive ablation studies, significance tests, and case studies confirm the robustness and feasibility of the method in practice.
Predicting the properties for unseen materials exclusively on the basis of the chemical formula before synthesis and characterization has advantages for research and resource planning. This can be achieved using suitable structure-free encoding and machine learning methods, but additional processing decisions are required. In this study, we compare a variety of structure-free materials encodings and machine learning algorithms to predict the structure/property relationships of battery materials. It was found that the physical units used to measure the property labels have an important impact on the predictive ability of the models, regardless of the computational approach. Property labels with respect to weight give excellent performance, but property labels with respect to volume cannot be predicted with confidence using only chemical information, even when the underlying physical characteristics are the same. These results contrast with previous studies of unsupervised learning and classification, where structure-free encoding excelled, and highlight how the structural features or property labels of materials are represented plays an important role in the predictive ability of machine learning models.
Electron microscopy, a sub-field of microanalysis, is critical to many fields of research. The widespread use of electron microscopy for imaging molecules and materials has had an enormous impact on our understanding of countless systems and has accelerated impacts in drug discovery and materials design, for electronic, energy, environment and health applications. With this success a bottleneck has emerged, as the rate at which we can collect data has significantly exceeded the rate at which we can analyze it. Fortunately, this has coincided with the rise of advanced computational methods, including data science and machine learning. Deep learning (DL), a sub-field of machine learning capable of learning from large quantities of data such as images, is ideally suited to overcome some of the challenges of electron microscopy at scale. There are a variety of different DL approaches relevant to the field, with unique advantages and disadvantages. In this review, we describe some well-established methods, with some recent examples, and introduce some new methods currently emerging in computer science. Our summary of DL is designed to guide electron microscopists to choose the right DL algorithm for their research and prepare for their digital future.
Machine learning is proving to be an ideal tool for materials design, capable of predicting forward structure-property relationships, and inverse property-structure relationships. However, it has yet to be used extensively for materials engineering challenges, predicting post-processing/structure relationships, and has yet to be used for to predict structure/post-processing relationships for inverse engineering. This is often due to the lack of sufficient metadata, and the overall scarcity and imbalance of processing data in many domains. This topic is explored in the current study using binary and multi-class classification to predict the appropriate post-synthesis processing conditions for aluminium alloys, based entirely on the alloying composition. The data imbalance was addressed using a new guided oversampling strategy that improves model performance by simultaneously balancing the classes and avoiding noise that contributes to over-fitting. This is achieved by through the deliberate but strategic introduction of not-a-numbers (NaNs) and the use of algorithms that naturally avoid them during learning. The outcome is the successful training of highly accurate binary classifiers, with significant reductions in false negatives and/or false positives with respect to the classifiers trained on the original data alone. Superior results were obtained for models predicting whether alloys should be solutionised or aged, post-synthesis, by guiding the re-balancing of the classes based on features (metals) that are highly ranked by the classifier, and then doubling the size of the data set via interpolation. Overall, this strategy has the greatest impact on tasks with a Shannon Diversity Index greater than 1 or less than 0.5, but can be applied to any prediction of post-processing conditions as part of an inverse engineering workflow.
The application of supervised machine learning to the study of catalytic metal nanoparticles has been shown to deliver excellent performance for a range of predictive tasks. However, this success assumes that the particles have been thoroughly characterised and that the property labels are known. Even in exclusively computational studies, the labelling of metal nanoparticles remains the bottleneck for most machine learning studies due to either high computational costs or low relevance to the experimental properties of interest. To facilitate more widespread use of machine learning in catalysis, a computationally affordable strategy to describe metal nanoparticles by a label that is relevant to their catalytic activities is needed. In this study we propose an entirely data-driven approach that can be automated to characterise the patterns and catalytic activities of the surface atoms of simulated metal nanoparticles, and evaluate its utility for catalytic applications. The structural patterns and catalytic activities of the surface atoms of simulated metal nanoparticles are characterised by an automatable data-driven unsupervised machine learning approach.
Different prediction models might perform equally well (Rashomon set) in the same task, but offer conflicting interpretations and conclusions about the data. The Rashomon effect in the context of Explainable AI (XAI) has been recognized as a critical factor. Although the Rashomon set has been introduced and studied in various contexts, its practical application is at its infancy stage and lacks adequate guidance and evaluation. We study the problem of the Rashomon set sampling from a practical viewpoint and identify two fundamental axioms - generalizability and implementation sparsity that exploring methods ought to satisfy in practical usage. These two axioms are not satisfied by most known attribution methods, which we consider to be a fundamental weakness. We use the norms to guide the design of an ϵ-subgradient-based sampling method. We apply this method to a fundamental mathematical problem as a proof of concept and to a set of practical datasets to demonstrate its ability compared with existing sampling methods.
The fractal dimension of a surface allows its degree of roughness to be characterized quantitatively. However, limited effort is attempted to calculate the fractal dimension of surfaces computed from precisely known atomic coordinates from computational biomolecular and nanomaterial studies. This work proposes methods to estimate the fractal dimension of the surface of any 3D object composed of spheres, by representing the surface as either a voxelized point cloud or a mathematically exact surface, and computing its box-counting dimension. Sphractal is published as a Python package that provides these functionalities, and its utility is demonstrated on a set of simulated palladium nanoparticle data.
The design of aluminium alloys often encounters a trade-off between strength and ductility, making it challenging to achieve desired properties. Adding to this challenge is the broad range of alloying elements, their varying concentrations, and the different processing conditions (features) available for alloy production. Traditionally, the inverse design of alloys using machine learning involves combining a trained regression model for the prediction of properties with a multi-objective genetic algorithm to search for optimal features. This paper presents an enhancement in this approach by integrating data-driven classes to train class-specific regressors. These models are then used individually with genetic algorithms to search for alloys with high strength and elongation. The results demonstrate that this improved workflow can surpass traditional class-agnostic optimisation in predicting alloys with higher tensile strength and elongation.