While machine learning promises to accelerate materials discovery, its opaque nature risks undermining both scientific rigor and trust in its predictions. This has motivated the development and use of eXplainable Artificial Intelligence (XAI) methods, which aim to elucidate the decision-making logic behind these intelligent systems. In this paper, we provide a critical review of recent advances in XAI applied to materials science, based on a systematic analysis of more than 140 publications. Our review identifies conceptual ambiguities in XAI terminology and clarifies the distinction between self-explanatory and post-hoc approaches. In this regard, we introduce a taxonomy that organizes current families of XAI methods and highlights key methodological innovations in the field. Our analysis highlights SHAP’s dominance, such that it has become the gold standard for XAI in materials science. This adoption raises both opportunities and concerns, as graph-based and perturbation-based approaches continue to emerge. Finally, we present current limitations and open research questions about the use of XAI in the field, particularly regarding how we evaluate and trust explanations. Moving from prediction to understanding is now the central challenge for applying machine learning to materials science.
ABSTRACT Generative artificial intelligence (GAI) methods have shown strong potential to accelerate drug discovery projects and de novo drug design. Yet, the fast pace of new proposals and the complexity of existing GAI methods configures a difficult scenario for professionals that aim to understand and effectively exploit these new tools. These issues are further exacerbated due to the plethora of different architectures for generative models, molecular representations, generation objectives and evaluation metrics. Moreover, other relevant aspects for understanding the inner workings of these models and their applications are scarcely addressed in the literature. For this reason, we propose different dimensions to taxonomically organize the wide range of GAI approaches used for drug discovery. These dimensions include the underlying computational model, the representation of molecules, and the building block for molecular generation. We also describe advantages and limitations of each family of methods, along with a comprehensive survey of metrics to evaluate performances from multiple perspectives. Besides, we explore how the concept of applicability domain applies to conditional generative methods, particularly in terms of quantifying the uncertainty associated with the generated molecules. As a final contribution, we describe and classify different approaches currently used in GAI‐driven drug design to enhance explainability. Finally, we discuss open challenges to strengthen the adoption of these generative models within research and industry. This article is categorized under: Technologies > Machine Learning Application Areas > Health Care
The Anatomical Therapeutic Chemical (ATC) code is a drug classification system that indicates the therapeutic potential use of a compound. Predicting ATC codes for drugs using automatic approaches is key to guide clinical trials and for drug repurposing. However, such automatic assignment is challenging due to the hierarchical organization of the code in four levels, possible polypharmacological behavior, and the imbalance and scarcity of annotated data in relation to the large number of compounds and possible ATC codes a drug may have. In this work, we propose a novel multimodal generative approach for predicting ATC codes, which leverages molecular information using a sequence-to-sequence architecture. Our hypothesis explores the idea that describing the chemical structure of the input compounds using two different representations, i.e., modes, the SMILES code and its molecular descriptors, provides complementary information, hence improving the accuracy of the predictions. Furthermore, given the multilabel nature of generative sequence-based models, we also present an additional prediction method to determine when to stop generating ATC labels for each compound. We compared the performance of our proposed methods against several baselines, both for new drugs and in drug repurposing tasks. In all of these cases, the superior performance of our multimodal proposals is clearly demonstrated. The source code and different data sets used to train and evaluate the models are made publicly available.
Predicting the ductile behavior of thermoplastic materials is a significant challenge for both industry and research. In this study, we present a predictive model that classifies polymers based on their ductility degree, which is a relationship created in the present work to approach this problem. It comes from relating two critical and measurable properties from the tensile test. This target was discretized into three classes, more ductile, intermediate, and less ductile. The feature selection process employed for finding the most relevant molecular descriptors for the predictive model used two approaches: a classical one and an expert-guided one. A new metric, called relaxed %CC, was presented to prioritize the models that reduce the misclassification between the extremes of the ductile scale, which is considered more important than confusion with the intermediate class. Our final model was able to successfully classify polymers, achieving a precision rate of 0.91, an %CC of 89.47 % (traditional accuracy), and a relaxed %CC of 100 % (no extremes confusion). This approach has the potential to help both the industry and R&D by selecting polymers with suitable ductile properties for specific applications during the design stage before their synthesis.
Pharmacovigilance performed from social media data is an active research field that contributes to the automatic detection of adverse drug reactions (ADRs) of medications and vaccines. Natural language processing techniques combined with machine learning models are used to perform the challenging task of analyzing heterogeneous short text content. This study explores the application of state-of-the-art transfer learning approaches for classifying Spanish tweets to identify mentions of ADRs as a result of COVID-19 vaccination. We created a corpus of 1332 tweets about COVID-19 post-vaccination adverse reactions and employed language models for text classification. Preliminary results suggest that these models achieve superior performances in terms of F1 score compared to traditional machine learning models.
Deep learning has had tremendous impact on numerous scientific and technological fields. We review key concepts and methods in deep learning with core applications in drug design and development, while introducing the main types of neural architectures. We emphasize how deep learning can be used for decision support and decision-making in drug development, discussing these advancements in the context of two case studies.
The likelihood of being diagnosed with thyroid cancer has increased in recent years; it is the fastest-expanding cancer in the United States and it has tripled in the last three decades. In particular, Papillary Thyroid Carcinoma (PTC) is the most common type of cancer affecting the thyroid. It is a slow-growing cancer and, thus, it can usually be cured. However, given the worrying increase in the diagnosis of this type of cancer, the discovery of new genetic markers for accurate treatment and prognostic is crucial. In the present study, the aim is to identify putative genes that may be specifically relevant in PTC through bioinformatic analysis of several gene expression public datasets and clinical information. Two datasets from Gene Expression Omnibus (GEO) and The Cancer Genome Atlas (TCGA) dataset were studied. Statistics and machine learning methods were sequentially employed to retrieve a final small cluster of genes of interest: PTGFR, ZMAT3, GABRB2, and DPP6. Kaplan-Meier plots were employed to assess the expression levels regarding overall survival and relapse-free survival. Furthermore, a manual bibliographic search for each gene was carried out, and a Protein-Protein Interaction (PPI) network was built to verify existing associations among them, followed by a new enrichment analysis. The results revealed that all the genes are highly relevant in the context of thyroid cancer and, more particularly interesting, PTGFR and DPP6 have not yet been associated with the disease up to date, thus making them worthy of further investigation as to their relationship to PTC.
Matrices that cannot be handled using conventional clustering, regression or classification methods are often found in every big data research area. In particular, datasets with thousands or millions of rows and less than a hundred columns regularly appear in biological so-called omic problems. The effectiveness of conventional data analysis approaches is hampered by this matrix structure, which necessitates some means of reduction. An evolutionary method called PreCLAS is presented in this article. Its main objective is to find a submatrix with fewer rows that exhibits some group structure. Three stages of experiments were performed. First, a benchmark dataset was used to assess the correct functionality of the method for clustering purposes. Then, a microarray gene expression data matrix was used to analyze the method’s performance in a simple classification scenario, where differential expression was carried out. Finally, several classification methods were compared in terms of classification accuracy using an RNA-seq gene expression dataset. Experiments showed that the new evolutionary technique significantly reduces the number of rows in the matrix and intelligently performs unsupervised row selection, improving classification and clustering methods.
Artificial intelligence (AI) is an emerging technology that is revolutionizing the discovery of new materials. One key application of AI is virtual screening of chemical libraries, which enables the accelerated discovery of materials with desired properties. In this study, we developed computational models to predict the dispersancy efficiency of oil and lubricant additives, a critical property in their design that can be estimated through a quantity named blotter spot. We propose a comprehensive approach that combines machine learning techniques with visual analytics strategies in an interactive tool that supports domain experts' decision-making. We evaluated the proposed models quantitatively and illustrated their benefits through a case study. Specifically, we analyzed a series of virtual polyisobutylene succinimide (PIBSI) molecules derived from a known reference substrate. Our best-performing probabilistic model was Bayesian Additive Regression Trees (BART), which achieved a mean absolute error of 5.50±0.34 and a root mean square error of 7.56±0.47, as estimated through 5-fold cross-validation. To facilitate future research, we have made the dataset, including the potential dispersants used for modeling, publicly available. Our approach can help accelerate the discovery of new oil and lubricant additives, and our interactive tool can aid domain experts in making informed decisions based on blotter spot and other key properties.
This work describes a hybrid methodology that combines machine learning and expert intervention to improve the predictive modeling of high-interest properties of polymeric materials. Although these materials have many advantages, developing a new material with specific properties, from a new molecular structure, is a great challenge and a costly and time-consuming task. The demand for materials with very specific properties continues to grow, so machine learning techniques have been applied for the prediction of these properties. The hybrid methodology was developed in an evolutionary way from expert intervention at the end of the machine learning process to a more determinative intervention throughout the cycle. This allows for more robust and reliable models for the design of new materials, which can help designers obtain property profiles for prototypes prior to the synthesis stage, saving time and resources.
Artificial intelligence (AI) is having a growing impact in many areas related to drug discovery. However, it is still critical for their adoption by the medicinal chemistry community to achieve models that, in addition to achieving high performance in their predictions, can be trusty explained to the end users in terms of their knowledge and background. Therefore, the investigation and development of explainable artificial intelligence (XAI) methods have become a key topic to address this challenge. For this reason, a comprehensive literature review about explanation methodologies for AI based models, focused in the field of drug discovery, is provided. In particular, an intuitive overview about each family of XAI approaches, such as those based on feature attribution, graph topologies, or counterfactual reasoning, oriented to a wide audience without a strong background in the AI discipline is introduced. As the main contribution, we propose a new taxonomy of the current XAI methods, which take into account specific issues related with the typical representations and computational problems study in the design of molecules. Additionally, we also present the main visualization strategies designed for supporting XAI approaches in the chemical domain. We conclude with key ideas about each method category, thoroughly providing insightful analysis about the guidelines and potential benefits of their adoption in medical chemistry. This article is categorized under: Data Science > Artificial Intelligence/Machine Learning
The classification of human cancers constitutes to date a significant challenge in the context of microarray data analysis. The discovery of gene hallmarks for biological processes involves the examination of large gene expression matrices in a broad and massively parallel manner. In this article, a comprehensive and comparative analysis of thyroid cancer datasets is presented, including stages for feature selection, hypothesis testing, and classification. Also, datasets are integrated, and results for this integration are reported and analyzed. To conclude, text mining is used to investigate some biological information regarding the main resulting characteristic genes. Some genes found during the research, HINT3 in particular, appear to be worth to be further studied.
The artificial intelligence-based prediction of the mechanical properties derived from the tensile test plays a key role in assessing the application profile of new polymeric materials, especially in the design stage, prior to synthesis. This strategy saves time and resources when creating new polymers with improved properties that are increasingly demanded by the market. A quantitative structure-property relationship (QSPR) model for tensile strength at break is presented in this work. The QSPR methodology applied here is based on machine learning tools, visual analytics methods, and expert-in-the-loop strategies. From the whole study, a QSPR model composed of five molecular descriptors that achieved a correlation coefficient of 0.9226 is proposed. We applied visual analytics tools at two levels of analysis: a more general one in which models are discarded for redundant information metrics and a deeper one in which a chemistry expert can make decisions on the composition of the model in terms of subsets of molecular descriptors, from a physical-chemical point of view. In this way, with the present work, we close a contribution cycle to polymer informatics, providing QSPR models oriented to the prediction of mechanical properties related to the tensile test.
With the consolidation of deep learning in drug discovery, several novel algorithms for learning molecular representations have been proposed. Despite the interest of the community in developing new methods for learning molecular embeddings and their theoretical benefits, comparing molecular embeddings with each other and with traditional representations is not straightforward, which in turn hinders the process of choosing a suitable representation for Quantitative Structure-Activity Relationship (QSAR) modeling. A reason behind this issue is the difficulty of conducting a fair and thorough comparison of the different existing embedding approaches, which requires numerous experiments on various datasets and training scenarios. To close this gap, we reviewed the literature on methods for molecular embeddings and reproduced three unsupervised and two supervised molecular embedding techniques recently proposed in the literature. We compared these five methods concerning their performance in QSAR scenarios using different classification and regression datasets. We also compared these representations to traditional molecular representations, namely molecular descriptors and fingerprints. As opposed to the expected outcome, our experimental setup consisting of over $25 000$ trained models and statistical tests revealed that the predictive performance using molecular embeddings did not significantly surpass that of traditional representations. Although supervised embeddings yielded competitive results compared with those using traditional molecular representations, unsupervised embeddings tended to perform worse than traditional representations. Our results highlight the need for conducting a careful comparison and analysis of the different embedding techniques prior to using them in drug design tasks and motivate a discussion about the potential of molecular embeddings in computer-aided drug design.
The Ames mutagenicity test constitutes the most frequently used assay to estimate the mutagenic potential of drug candidates. While this test employs experimental results using various strains of Salmonella typhimurium, the vast majority of the published in silico models for predicting mutagenicity do not take into account the test results of the individual experiments conducted for each strain. Instead, such QSAR models are generally trained employing overall labels (i.e., mutagenic and nonmutagenic). Recently, neural-based models combined with multitask learning strategies have yielded interesting results in different domains, given their capabilities to model multitarget functions. In this scenario, we propose a novel neural-based QSAR model to predict mutagenicity that leverages experimental results from different strains involved in the Ames test by means of a multitask learning approach. To the best of our knowledge, the modeling strategy hereby proposed has not been applied to model Ames mutagenicity previously. The results yielded by our model surpass those obtained by single-task modeling strategies, such as models that predict the overall Ames label or ensemble models built from individual strains. For reproducibility and accessibility purposes, all source code and datasets used in our experiments are publicly available.
Polymer informatics is an emerging discipline that has benefited from the strong development that data science has experienced over the last decade. Machine learning methods are useful to infer QSPR (Quantitative Structure-Property Relationships) models that allow predicting mechanical properties related to the industrial profile of polymeric materials based on their structural repeating units (SRUs). Nonetheless, the chemical structure of the SRU is only one of the many factors that affects the industrial usefulness of a polymer. Other equally relevant factors are polymer molecular weight, molecular weight distribution, and production method, which are related to the inherent polydispersity of this kind of material. For this reason, the computational characterization used for the building of QSPR models for predicting mechanical properties should consider these main factors. The aim of this paper is to highlight recent advances in data science to address the inclusion of polydispersity information of polymeric materials in QSPR modeling. We present two dimensions of discussion: data representation and algorithmic issues. In the first one, we examine how different strategies can be applied to include polydispersity data in the molecular descriptors that characterize the polymers. We explain two data representation approaches designed by our group, named as trivalued and multivalued molecular descriptors. In the second dimension, we discuss algorithms proposed to deal with these new molecular descriptor representations during the construction of the QSPR models. Thus, we present here a comprehensible and integral methodology to address the challenges that polydispersity generates in the QSPR modeling of mechanical properties of polymers.
Food informatics is having an increasing impact on the food industry and improving the quality of end products, as well as the efficiency of manufacturing processes. In the case of winemaking, a particular application of interest for food informatics is the sensory analysis of wines. This problem can benefit from the strong development that machine learning has achieved in recent decades. However, these data-driven techniques require accurate and sufficient information to generate models capable of predicting the sensory profile of wines. A review of the sensory analysis and volatile composition of wines is presented in this work, along with significant studies on the use of machine learning models to predict wine-related characteristics such as the antioxidant activity of polyphenols of wine and aroma compounds. In this sense, data from a sensory panel and analytical technology were gathered. This literature review reveals the lack of a homogeneous and sufficiently large database of sensory analysis related to the volatile composition of wines to develop machine learning models. However, among artificial intelligence approaches, the application of quantitative structure-odour relationship (QSOR) models is currently gaining importance. Recent studies show that it would be possible to predict quantitatively the sensory analysis of wines by QSOR models, using general volatile composition information. Therefore, the purpose of this review is to identify key aspects and guidelines for the development of this area.
The Polymer Maker SMILES-based (PolyMaS) software was used to generate linear macromolecules from the repeating structural units (SRU) of polymers without limiting their length and molar mass. The SRU input is stored in the SMILES code available on the Internet. PolyMaS makes head-tail junctions to the desired length of the macromolecule.
In the modern drug discovery process, medicinal chemists deal with the complexity of analysis of large ensembles of candidate molecules. Computational tools, such as dimensionality reduction (DR) and classification, are commonly used to efficiently process the multidimensional space of features. These underlying calculations often hinder interpretability of results and prevent experts from assessing the impact of individual molecular features on the resulting representations. To provide a solution for scrutinizing such complex data, we introduce ChemVA, an interactive application for the visual exploration of large molecular ensembles and their features. Our tool consists of multiple coordinated views: Hexagonal view, Detail view, 3D view, Table view, and a newly proposed Difference view designed for the comparison of DR projections. These views display DR projections combined with biological activity, selected molecular features, and confidence scores for each of these projections. This conjunction of views allows the user to drill down through the dataset and to efficiently select candidate compounds. Our approach was evaluated on two case studies of finding structurally similar ligands with similar binding affinity to a target protein, as well as on an external qualitative evaluation. The results suggest that our system allows effective visual inspection and comparison of different high-dimensional molecular representations. Furthermore, ChemVA assists in the identification of candidate compounds while providing information on the certainty behind different molecular representations.
The aim of industry 4.0 is to promote productivity and innovation by incorporating emerging IT technologies, where machine learning is playing a central role in this industrial revolution. In this sense, the production of new materials could take advantage of novel virtual testing approaches based on data science for supporting the design of new polymers. Nevertheless, the lack of data for learning virtual testing models constitutes a hard challenge for progressing in these innovative techniques. Therefore, it is especially important to create reliable databases for polymer study and make them available to the scientific community. In this work, we have focused on the generation of a trustworthy database of Refractive Index (RI) of synthetic polymers. This paper details the different types of errors found in the data source and the corrections made during the curation and cleaning of this database. Additionally, some Quantitative Structure-Property Relationship models for predicting RI, inferred without domain expert intervention, are presented and discussed for illustrating how virtual testing can be applied using this database.
Carlos M. Lorenzetti合作论文数Universidad Nacional del Sur3
Marc Strickert合作论文数Institute of Plant Genetics and Crop Plant Research2