Delineating the superconducting order parameters is a pivotal task in investigating superconductivity for probing pairing mechanisms, as well as their symmetry and topology. Point-contact Andreev reflection (PCAR) measurement is a simple yet powerful tool for identifying the order parameters. The PCAR spectra exhibit significant variations depending on the type of the order parameter in a superconductor, including its magnitude (Delta), as well as temperature, interfacial quality, Fermi velocity mismatch, and other factors. The information on the order parameter can be obtained by finding the combination of these parameters, generating a theoretical spectrum that fits a measured experimental spectrum. However, due to the complexity of the spectra and the high dimensionality of parameters, extracting the fitting parameters is often time-consuming and labor-intensive. In this study, we employ a convolutional neural network (CNN) algorithm to create models for rapid and automated analysis of PCAR spectra of various superconductors with different pairing symmetries (conventional s-wave, chiral px + ipy-wave, and dx2-y2 -wave). The training datasets are generated based on the Blonder-TinkhamKlapwijk (BTK) theory and further modified and augmented by selectively incorporating noise and peaks according to the bias voltages. This approach not only replicates the experimental spectra but also brings the model's attention to important features within the spectra. The optimized models provide fitting parameters for experimentally measured spectra in less than 100 ms per spectrum. Our approaches and findings pave the way for rapid and automated spectral analysis which will help accelerate research on superconductors with complex order parameters.
Link prediction is a key network analysis technique that infers missing or future relations between nodes in a graph, based on observed patterns of connectivity. Scientific literature networks and knowledge graphs are typically large, sparse, and noisy, and often contain missing links, potential but unobserved connections, between concepts, entities, or methods. Here, we present an AI-driven hierarchical link prediction framework that integrates matrix factorization to infer hidden associations and steer discovery in complex material domains. Our method combines Hierarchical Nonnegative Matrix Factorization (HNMFk), Boolean matrix factorization (BNMFk) with automatic model selection. These discrete factors are then fused with Logistic matrix factorization (LMF), we use to construct a three-level topic tree from a 46,862-document corpus focused on 73 transition-metal dichalcogenides (TMDs). This class of materials has been studied in a variety of physics fields and has a multitude of current and potential applications. An ensemble BNMFk + LMF approach fuses discrete interpretability with probabilistic scoring. The resulting HNMFk clusters map each material onto coherent research themes, such as superconductivity, energy storage, and tribology, and highlight missing or weakly connected links between topics and materials, suggesting novel hypotheses for cross-disciplinary exploration. We validate our method by removing publications about superconductivity in wellknown superconductors, and demonstrate that the model correctly predicts their association with the superconducting TMD clusters. This highlights the ability of the method to find hidden connections in a graph of material to latent topic associations built from scientific literature. This is especially useful when examining a diverse corpus of scientific documents covering the same class of phenomena or materials but originating from distinct communities and perspectives. The inferred links generating new hypotheses, produced by our method, are exposed through an interactive Streamlit dashboard, designed for scientific discovery.
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external document retrieval to provide domain-specific or up-to-date knowledge. The effectiveness of RAG depends on the relevance of retrieved documents, which is influenced by the semantic alignment of embeddings with the domain's specialized content. Although full fine-tuning can align language models to specific domains, it is computationally intensive and demands substantial data. This paper introduces Hierarchical Embedding Alignment Loss (HEAL), a novel method that leverages hierarchical fuzzy clustering with matrix factorization within contrastive learning to efficiently align LLM embeddings with domain-specific content. HEAL computes level/depth-wise contrastive losses and incorporates hierarchical penalties to align embeddings with the underlying relationships in label hierarchies. This approach enhances retrieval relevance and document classification, effectively reducing hallucinations in LLM outputs. In our experiments, we benchmark and evaluate HEAL across diverse domains, including Healthcare, Material Science, Cyber-security, and Applied Maths.
FeSe is one of the most enigmatic superconductors. Among the family of iron-based compounds, it has the simplest chemical makeup and structure, and yet it displays superconducting transition temperature ( T c ) spanning 0 to 15 K for thin films, while it is typically 8 K for single crystals. This large variation of T c within one family underscores a key challenge associated with understanding superconductivity in iron chalcogenides. Here, using a dual-beam pulsed laser deposition (PLD) approach, we have fabricated a unique lattice-constant gradient thin film of FeSe which has revealed a clear relationship between the atomic structure and the superconducting transition temperature for the first time. The dual-beam PLD that generates laser fluence gradient inside the plasma plume has resulted in a continuous variation in distribution of edge dislocations within a single film, and a precise correlation between the lattice constant and T c has been observed here, namely, T c ∝ c - c 0 , where c is the c-axis lattice constant (and c 0 is a constant). This explicit relation in conjunction with a theoretical investigation indicates that it is the shifting of the d xy orbital of Fe which plays a governing role in the interplay between nematicity and superconductivity in FeSe.Supplementary Information:The online version contains supplementary material available at 10.1007/s44214-024-00058-0.
We present a machine learning framework for modeling the absorption properties of azobenzene molecules – an important class of organic compounds with many potential photochemical applications. The framework utilizes predictors based on the chemical composition and structure of each molecule and consists of separate regression models trained to predict the absorption at distinct wavelengths, covering the UV and visible light ranges. Despite the relatively small size of the dataset (330 molecule-absorption spectrum pairs), the models were able to learn to accurately predict the absorption at fixed wavelengths, as well as the position and intensity of the maximum absorption. These predictions can be used to rapidly screen thousands of candidate molecules for a variety of potential applications, reducing the need for time-consuming and expensive experiments or first-principles computations.
Transition metal carbides and nitrides have unique mechanical and chemical characteristics. At low temperatures many of them also exhibit superconductivity, which can be controlled by substitutions into both the transition metal and carbon/nitrogen sites. To investigate the factors governing the superconducting state, we apply machine learning methods. We collected a dataset containing 147 materials, which was used to create a pipeline for predicting their superconducting critical temperature. When this pipeline is applied to a randomly selected test set, it shows a good performance, with R2 of 0.82 and RMSE of 1.9 K. To explore the limits of the machine learning approach, we also use it to predict entire substitution series within the dataset. This represents a realistic test for the predictive models, which can be extremely useful when applied to new substitutions in materials systems. The performance of the pipeline in this case is much more uneven, with good predictions for some series, while for others the model shows minimal predictive power. We discuss possible reasons for these results, as well as methods to estimate the performance of machine learning on new substitution series.
In the biopharmaceutical industry, chromatography resins have a finite number of uses before they start to age and degrade, typically due to losses of ligand integrity and/or density. The "health" of a column is predicted and validated by running multiple cycles on representative scale-down models and can be followed by real-time on-going validation during commercial production. Principal Component Analysis (PCA), Partial Least Square (PLS), Similarity Scores and Single One Point-MultiParameter Technique (SOP-MPT) along with machine learning principles were applied to explore the hypothesis that there is predictive capability of latent variables in chromatography absorbance profiles for process performance (step yield) and product quality (aggregates, fragments, host cell proteins (HCP) and DNA, and Protein A ligand). The first stage of this study is described in this paper: a MabSelect SuRe™ chromatography column was cycled with a method to establish the "normal" baseline for process performance and product quality, followed by runs using a harsher NaOH Cleaning in Place (CIP) procedure (with a higher NaOH concentration than that recommended by the vendor) to accelerate resin degradation. The different mathematical analytical tools correlated with resin degradation of the column (reflected in decreasing step yield and binding capacity with increasing running cycle), specifically when using the Wash, Elution and Strip phases of the chromatography method. Monomer, HCP and DNA content were not significantly impacted and therefore a correlation with product quality was inconsequential. Importantly, this work shows proof-of-concept that while more traditional methods of measuring resin integrity such as the height equivalent to a theoretical place (HETP) and Asymmetry (As) measurements could not detect changes in the integrity of the resin, PCA, PLS, Similarity Scores and SOP-MPT (to a lesser extent) applied to the absorbance data were capable of anticipating issues in the chromatography bed by identifying atypical outcomes.
We have developed a phase mapping method based on machine learning analysis of reflection high-energy electron diffraction (RHEED) images. RHEED produces diffraction patterns containing a wealth of static and dynamic information and is commonly used to determine the growth rate, the growth mode, and the surface morphology of epitaxial thin films. However, the ability to extract quantitative structural information from the RHEED patterns that appear during film growth is limited by the lack of versatile and automated analysis techniques. We have created a deep learning-based analysis method for automating the identification of different RHEED pattern types that occur during the growth of a material. Our approach combines several supervised and unsupervised machine learning techniques and permits the extraction of quantitative phase composition information. We applied this method to the mapping of the structural phase diagram of FexOy thin films grown by pulsed laser deposition as a function of growth temperature and oxygen pressure close to the hematite-magnetite phase boundary. The in situ RHEED-based mapping method produces results that are qualitatively similar to postsynthesis x-ray diffraction analysis.
Designing materials with advanced functionalities is the main focus of contemporary solid-state physics and chemistry. Research efforts worldwide are funneled into a few high-end goals, one of the oldest, and most fascinating of which is the search for an ambient temperature superconductor (A-SC). The reason is clear: superconductivity at ambient conditions implies being able to handle, measure and access a single, coherent, macroscopic quantum mechanical state without the limitations associated with cryogenics and pressurization. This would not only open exciting avenues for fundamental research, but also pave the road for a wide range of technological applications, affecting strategic areas such as energy conservation and climate change. In this roadmap we have collected contributions from many of the main actors working on superconductivity, and asked them to share their personal viewpoint on the field. The hope is that this article will serve not only as an instantaneous picture of the status of research, but also as a true roadmap defining the main long-term theoretical and experimental challenges that lie ahead. Interestingly, although the current research in superconductor design is dominated by conventional (phonon-mediated) superconductors, there seems to be a widespread consensus that achieving A-SC may require different pairing mechanisms. In memoriam, to Neil Ashcroft, who inspired us all.
We analyze a corpus consisting of more than 17,000 abstracts in the general field of superconductivity, extracted from the arXiv – an online repository of scientific articles. We utilize a recently developed topic modeling method called SeNMFk, extending the standard Non-negative Matrix Factorization (NMF) methods by incorporating the semantic structure of the text, and adding a robust system for determining the number of topics. With SeNMFk, we were able to extract coherent topics validated by human experts. From these topics, a few are relatively general and cover broad concepts, while the majority can be precisely mapped to particular scientific effects or measurement techniques. The topics also differ by ubiquity, with only three topics prevalent in almost 40% of the abstract, while each specific topic tends to dominate a small subset of the abstracts. These results demonstrate the ability of SeNMFk to produce a layered and nuanced analysis of large scientific corpora.
Artificial intelligence and machine learning are becoming indispensable tools in many areas of physics, including astrophysics, particle physics, and climate science. In the arena of quantum materials, the rise of new experimental and computational techniques has increased the volume and the speed with which data are collected, and artificial intelligence is poised to impact the exploration of new materials such as superconductors, spin liquids, and topological insulators. This review outlines how the use of data-driven approaches is changing the landscape of quantum materials research. From rapid construction and analysis of computational and experimental databases to implementing physical models as pathfinding guidelines for autonomous experiments, we show that artificial intelligence is already well on its way to becoming the lynchpin in the search and discovery of quantum materials. Quantum materials host many exotic properties, which might be utilized for new electronic devices. Here, artificial intelligence for the discovery of quantum materials is discussed, covering both materials and property prediction, and high-throughput synthesis.
Topic modeling, or identifying the set of topics that occur in a collection of articles, is one of the primary objectives of text mining. One of the big challenges in topic modeling is determining the correct number of topics: underestimating the number of topics results in a loss of information, i.e., omission of topics, underfitting, while overestimating leads to noisy and unexplainable topics and overfitting. In this paper, we consider a semantic-assisted non-negative matrix factorization (NMF) topics model, which we call SeNMFk, based on Kullback-Leibler(KL) divergence and integrated with a method for determining the number of latent topics. SeNMFk involves (i) creating a random ensemble of pairs of matrices whose mean is equal to the initial words-by-documents matrix representing the text corpus and the Shifted Positive Pointwise Mutual Information (SPPMI) matrix, which encodes the context information, respectively, and (ii) jointly factorizing each of these pairs with different number of topics to acquire sets of latent topics that are stable to noise. We demonstrate the performance of our method by identifying the number of topics in several benchmark text corpora, when compared to other state-of-the-art techniques. We also show that the number of document classes in the input text corpus may differ from the number of the extracted latent topics, but these classes can be retrieved by clustering the column-vectors of one of the factor matrices. Additionally, we introduce a software called pyDNMFk to estimate the number of topics. We demonstrate that our unsupervised method, SeNMFk, not only determines the correct number of topics, but also extracts topics with a high coherence and accurately classifies the documents of the corpus.
Solutions to many of the world's problems depend upon materials research and development. However, advanced materials can take decades to discover and decades more to fully deploy. Humans and robots have begun to partner to advance science and technology orders of magnitude faster than humans do today through the development and exploitation of closed-loop, autonomous experimentation systems. This review discusses the specific challenges and opportunities related to materials discovery and development that will emerge from this new paradigm. Our perspective incorporates input from stakeholders in academia, industry, government laboratories, and funding agencies. We outline the current status, barriers, and needed investments, culminating with a vision for the path forward. We intend the article to spark interest in this emerging research area and to motivate potential practitioners by illustrating early successes. We also aspire to encourage a creative reimagining of the next generation of materials science infrastructure. To this end, we frame future investments in materials science and technology, hardware and software infrastructure, artificial intelligence and autonomy methods, and critical workforce development for autonomous research.
Structure is the most basic and important property of crystalline solids; it determines directly or indirectly most materials characteristics. However, predicting crystal structure of solids remains a formidable and not fully solved problem. Standard theoretical tools for this task are computationally expensive and at times inaccurate. Here we present an alternative approach utilizing machine learning for crystal structure prediction. We developed a tool called Crystal Structure Prediction Network (CRYSPNet) that can predict the Bravais lattice, space group, and lattice parameters of an inorganic material based only on its chemical composition. CRYSPNet consists of a series of neural network models, using as inputs predictors aggregating the properties of the elements constituting the compound. It was trained and validated on more than 100,000 entries from the Inorganic Crystal Structure Database. The tool demonstrates robust predictive capability and outperforms alternative strategies by a large margin. Made available to the public (at https://github.com/AuroraLHT/cryspnet), it can be used both as an independent prediction engine or as a method to generate candidate structures for further computational and/or experimental validation.
Non-negative Matrix Factorization (NMF) models the topics of a text corpus by decomposing the matrix of term frequency-inverse document frequency (TF-IDF) representation, X, into two low-rank non-negative matrices: W , representing the topics and H, mapping the documents onto space of topics. One challenge, common to all topic models, is the determination of the number of latent topics (aka model determination). Determining the correct number of topics is important: underestimating the number of topics results in a poor topic separation, under-fitting, while overestimating leads to noisy topics, over-fitting. Here, we introduce SeNMFk, a semantic-assisted NMF-based topic modeling method, which incorporates semantic correlations in NMF by using a word-context matrix, and employs a method for determination of the number of latent topics. SeNMFk first creates a random ensemble of matrices based on the initial TF-IDF matrix and a word-context matrix, and then applies a coupled factorization to acquire sets of stable coherent topics that are robust to noise. The latent dimension is determined based on the stability of these topics. We show that SeNMFk accurately determines the number of high-quality topics in benchmark text corpora, which leads to an accurate document clustering.
Applications are being found for superconducting materials in a rapidly growing number of technological areas, and the search for novel superconductors remains a major scientific task. However, the steady increase in the complexity of candidate materials presents a big challenge to researchers. In particular, conventional experimental methods are not well suited to an efficient search for candidates in a compositional space growing exponentially with the number of elements; neither do they permit a quick extraction of reliable multidimensional phase diagrams delineating the physical parameters that control superconductivity. New research paradigms that can boost the speed and the efficiency of research into superconducting materials are urgently needed. High-throughput methods for the rapid screening and optimization of materials have aided the acceleration of research in bioinformatics and the pharmaceutical industry, yet remain rare in quantum materials research. In this paper we briefly review the history of high-throughput research and then focus on some recent applications of this paradigm in superconductivity research. We consider the role these methods can play in all stages of materials development, including high-throughput computation, synthesis, characterization, and the emerging field of machine learning for materials. The high-throughput paradigm will undoubtedly become an indispensable tool in superconductivity research in the near future.
High‐dimensional datasets are becoming ubiquitous in many applications and therefore unsupervised tensor methods to interrogate them are needed. Here, we report a new unsupervised machine learning (ML) approach (NTFk) based on nonnegative tensor factorization integrated with a custom k‐means clustering. We demonstrate the ability of NTFk to extracting temporal and spatial features of phase separation of copolymers as they are modeled by self‐consistent field theory. Microphase separation of block copolymers has been extensively studied both experimentally and theoretically. However, the interpretation of computer simulations and/or experimental data, representing temporal and spatial changes of molecular species concentration is still a challenging task. Thus, extracting the phase diagram from simulations or experimental data as well as the interpretation of data requires discernment of the model/experimental parameters (such as, temperature, concentrations, the number of molecular species and the interaction between species) impact on the microphase separation process. An attractive and unique aspect of the introduced ML method is that it ensures the nonnegativity of the extracted latent features. Nonnegativity is an essential constraint needed to obtain interpretable and sparse latent features that are parts‐based representation of the data. The custom clustering in NTFk serves to estimate the number of latent features in the data.
Machine learning technologies are expected to be great tools for scientific discoveries. In particular, materials development (which has brought a lot of innovation by finding new and better functional materials) is one of the most attractive scientific fields. To apply machine learning to actual materials development, collaboration between scientists and machine learning is becoming inevitable. However, such collaboration has been restricted so far due to black box machine learning, in which it is difficult for scientists to interpret the data-driven model from the viewpoint of material science and physics. Here, we show a material development success story that was achieved by good collaboration between scientists and one type of interpretable (explainable) machine learning called factorized asymptotic Bayesian inference hierarchical mixture of experts (FAB/HMEs). Based on material science and physics, we interpreted the data-driven model constructed by the FAB/HMEs, so that we discovered surprising correlation and knowledge about thermoelectric material. Guided by this, we carried out actual material synthesis that led to identification of a novel spin-driven thermoelectric material with the largest thermopower to date.