Large language models (LLMs) are rapidly changing how researchers in materials science and chemistry discover, organize, and act on scientific knowledge. This paper analyzes a broad set of community-developed LLM applications in an effort to identify emerging patterns in how these systems can be used across the scientific research lifecycle. We organize the projects into two complementary categories: Knowledge Infrastructure, systems that structure, retrieve, synthesize, and validate scientific information; and Action Systems, systems that execute, coordinate, or automate scientific work across computational and experimental environments. The submissions reveal a shift from single-purpose LLM tools toward integrated, multi-agent workflows that combine retrieval, reasoning, tool use, and domain-specific validation. Prominent themes include retrieval-augmented generation as grounding infrastructure, persistent structured knowledge representations, multimodal and multilingual scientific inputs, and early progress toward laboratory-integrated closed-loop systems. Together, these results suggest that LLMs are evolving from general-purpose assistants into composable infrastructure for scientific reasoning and action. This work provides a community snapshot of that transition and a practical taxonomy for understanding emerging LLM-enabled workflows in materials science and chemistry.
Advances in generative artificial intelligence are transforming how metal-organic frameworks (MOFs) are designed and discovered. This Perspective introduces the shift from laborious enumeration of MOF candidates to generative approaches that can autonomously propose and synthesize in the laboratory new porous reticular structures on demand. We outline the progress of employing deep learning models, such as variational autoencoders, diffusion models, and large language model-based agents, that are fueled by the growing amount of available data from the MOF community and suggest novel crystalline materials designs. These generative tools can be combined with high-throughput computational screening and even automated experiments to form accelerated, closed-loop discovery pipelines. The result is a new paradigm for reticular chemistry in which AI algorithms more efficiently direct the search for high-performance MOF materials for clean air and energy applications. Finally, we highlight remaining challenges such as synthetic feasibility, dataset diversity, and the need for further integration of domain knowledge.
Emerging energy and electronic systems rely on the thermodynamic properties of chemical and cooling fluids. These properties are a function of both chemical structure and temperature. For instance, the dynamic viscosity of a fluid can vary by orders of magnitude across the operating range of a cooling system. However, capturing this behavior remains a challenge for experimental and modelling approaches. Machine learning models, although powerful for fixed temperatures, fail to generalize across temperatures due to a lack of data and a lack of embedded physical constraints. Here, we introduce a physics-informed machine learning framework that incorporates established physical relationships, such as the Arrhenius equation or Clausius-Clapeyron, to capture both chemical diversity and temperature dependence. We demonstrate that decoupling chemistry from thermodynamic conditions enables accurate prediction of temperature-dependent dynamic viscosity for both pure compounds and binary mixtures, which we validated with new experimental data. Through a materials-discovery campaign for cooling applications, we show that neglecting temperature effects can cause relative efficiency errors exceeding an order of magnitude, leading to inaccurate materials ranking and suboptimal fluid selection. Finally, we extend the framework to other properties, such as vapor pressure and diffusion coefficient, highlighting a generalizable strategy for accelerating fluid property prediction and design for sustainable technologies.
A new deep learning method enables molecular dynamics simulations over longer time scales while still achieving accurate physical property prediction.
Cocrystal formation is a widely used strategy in solid-state chemistry and pharmaceutical development to improve the solubility, stability, and bioavailability of molecules with otherwise poor physicochemical properties. Identifying viable coformer combinations remains laborious and uncertain. A key but underappreciated challenge is that experimental databases overwhelmingly report successful cocrystals, while unsuccessful attempts are rarely documented, creating biased data sets that cause many machine-learning models to make overly optimistic and unreliable predictions when applied to new chemical systems. Here, we address this limitation by reframing cocrystal prediction as a learning problem with missing negative information and by adopting a conservative strategy that focuses on identifying molecular pairs that are very unlikely to form cocrystals. We leverage multiple, independent molecular descriptions─including structural, electronic, and physicochemical characteristics─that provide complementary views for identifying reliable negatives, and use their agreement to exclude implausible combinations from large sets of untested pairs. These highly confident pseudonegative examples are then used to mitigate data imbalance and to fine-tune a pretrained graph attention network for cocrystal prediction. Across large and chemically diverse data sets, this data-centric strategy significantly improves the reliability and generalization of cocrystal prediction models compared with existing deep-learning approaches, demonstrating that carefully correcting for missing negative information is critical for making computational screening more realistic and more useful for guiding future experimental discovery.
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. Here, we introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, material science, and physics. For this benchmark, domain experts define research projects of genuine value and interest, and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, execute simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific “superintelligence”. Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery- relevant evaluation of LLMs and charts practical paths to advance their development toward scientific innovation.
Self-driving laboratories (SDLs) integrating automation and data-driven control are increasingly used in materials synthesis. Existing SDLs still face challenges in hardware integration and stable flow operation. Here, we report a modular fluidic-microwave SDL for the automated synthesis of perovskite nanocrystals (PNCs) with programmable thermal profiles and inline photoluminescence monitoring. Its constant-pressure fluidics and batch microwave design decouple reaction temperature from residence time, suppress flow pulsation, and yield reproducible low-noise data sets. Using CsPbX3 (X = Cl, Br, I), the platform achieves relative standard deviations of 1.53% in fwhm, 0.06% in peak wavelength, and 1.85% in PL intensity across runs. Automated parameter screening identifies an optimal synthesis window (120-140 °C, 28 °C min-1, 120-180 s, Pb/Cs = 2-3) that produces phase-pure nanocrystals with narrow emission line widths (∼18-19 nm). This SDL provides a reproducible, programmable basis for closed-loop optimization and AI-guided nanomaterials synthesis.
Universal machine learning interatomic potentials (uMLIPs) have emerged as powerful tools for accelerating atomistic simulations, offering scalable and efficient modeling with accuracy close to quantum calculations. However, their reliability and effectiveness in practical, real-world applications remain an open question. Metal-organic frameworks (MOFs) and related nanoporous materials are highly porous crystals with critical relevance in carbon capture, energy storage, and catalysis applications. Modeling nanoporous materials presents distinct challenges for uMLIPs due to their diverse chemistry, structural complexity, including porosity and coordination bonds, and the absence from existing training datasets. Here, we introduce MOFSimBench, a benchmark for evaluating uMLIPs on key materials modeling tasks for nanoporous materials, including structural optimization, molecular dynamics (MD) stability, bulk property prediction, and host-guest interactions. Evaluating 20 models from various architectures, we find that top-performing uMLIPs consistently outperform classical force fields and fine-tuned machine learning potentials across all tasks, demonstrating their readiness for deployment in nanoporous materials modeling. Our analysis highlights that data quality plays a more critical role than model architecture in determining performance across all evaluated uMLIPs. We release our modular and extensible benchmarking framework at https://github.com/AI4ChemS/mofsim-bench, providing an open resource to guide the adoption for nanoporous materials modeling and further development of uMLIPs.
Bayesian optimization (BO) is increasingly used in molecular optimization and in guiding self-driving laboratories for automated materials discovery. A crucial aspect of BO is how molecules and materials are represented as feature vectors, where both the completeness and compactness of these representations can influence the efficiency of the optimization process. Traditionally, a fixed representation is chosen by expert chemists or applying data-driven feature selection methods on available labeled datasets. However, when dealing with novel optimization tasks, prior knowledge or large datasets are often unavailable, and relying on these even can introduce bias into the search process. In this work, we demonstrate a Feature Adaptive Bayesian Optimization (FABO) framework, which integrates feature selection in the Bayesian optimization process with Gaussian processes to dynamically adapt material representations throughout the optimization cycles. We demonstrate the effectiveness of this adaptive approach across several molecular optimization tasks, including the discovery of high-performing metal-organic frameworks (MOFs) in three distinct tasks, each involving unique property distributions and requiring a distinct representation. Our results show that the adaptive nature of the representation leads to outperforming random search baseline and scenarios where prior knowledge of the feature space is available. Notably, for known optimization tasks, FABO automatically identifies representations that are aligned with human chemical intuition, validating its utility for optimization tasks where such insights are not available in advance. Lastly, we show how a suboptimal representation, e.g., when missing key features, can adversely impact BO performance, highlighting the importance of starting from a full feature set and adapt it to different tasks. Our findings highlight FABO as a robust approach for navigating large, complex materials search spaces in automated discovery campaigns.
Large Language Models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 34 total projects developed during the second annual Large Language Model Hackathon for Applications in Materials Science and Chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.
We present MOFA, an open-source generative AI (GenAI) plus simulation workflow for high-throughput generation of metal-organic frameworks (MOFs) on large-scale high-performance computing (HPC) systems. MOFA addresses key challenges in integrating GPU-accelerated computing for GPU-intensive GenAI tasks, including distributed training and inference, alongside CPU- and GPU-optimized tasks for screening and filtering AI-generated MOFs using molecular dynamics, density functional theory, and Monte Carlo simulations. These heterogeneous tasks are unified within an online learning framework that optimizes the utilization of available CPU and GPU resources across HPC systems. Performance metrics from a 450-node (14,400 AMD Zen 3 CPUs + 1800 NVIDIA A100 GPUs) supercomputer run demonstrate that MOFA achieves high-throughput generation of novel MOF structures, with CO$_2$ adsorption capacities ranking among the top 10 in the hypothetical MOF (hMOF) dataset. Furthermore, the production of high-quality MOFs exhibits a linear relationship with the number of nodes utilized. The modular architecture of MOFA will facilitate its integration into other scientific applications that dynamically combine GenAI with large-scale simulations.
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.
The success of diffusion models in the field of image processing has propelled the creation of software such as Dall-E, Midjourney and Stable Diffusion, which are tools used for text-to-image generations. Mapping this workflow onto materials discovery, a new diffusion model was developed for the generation of pure silica zeolite, marking it the first application of diffusion models to porous materials. Our model demonstrates the ability to generate novel crystalline porous materials that are not present in the training dataset, while exhibiting exceptional performance in inverse design tasks targeted on various chemical properties including the void fraction, Henry coefficient and heat of adsorption. Comparing our model with a Generative Adversarial Network (GAN) revealed that the diffusion model outperforms the GAN in terms of structure validity, exhibiting an over 2,000-fold improvement in performance. We firmly believe that diffusion models (along with other deep generative models) hold immense potential in revolutionizing the design of new materials, and anticipate the wide extension of our model to other classes of porous materials.
We present an updated version of the Computation-Ready, Experimental (CoRE) Metal-Organic Framework (MOF) database, which includes a curated set of computation-ready MOF crystal structures designed for high-throughput computational materials discovery. Data collection and curation procedures were improved from the previous version to enable more frequent updates in the future. Machine-learning-predicted properties, such as stability metrics and heat capacities, are included in the dataset to streamline screening activities. An updated version of MOFid was developed to provide detailed information on metal nodes, organic linkers, and topologies of an MOF structure. DDEC6 partial atomic charges of MOFs were assigned based on a machine-learning model. Gibbs ensemble Monte Carlo simulations were used to classify the hydrophobicity of MOFs. The finalized dataset was subsequently used to perform integrated material-process screening for various carbon-capture conditions using high-fidelity temperature-swing adsorption (TSA) simulations. Our workflow identified multiple MOF candidates that are predicted to outperform CALF-20 for these applications.
The widespread adoption of electric vehicles (EVs) depends on batteries with higher energy density and faster charging. Achieving these goals requires effective thermal management, with immersion cooling emerging as a promising solution. However, the lack of a standardized metric for screening immersion cooling fluids remains a challenge, in part because performance is a function of fluid properties and geometry. This study introduces a figure of merit (FOM) to quantify the thermal performance of dielectric fluids for the immersion cooling of the most common cylindrical EV battery packs. We combined computational fluid dynamics (CFD) and active learning techniques to screen over 1600 fluid mixtures drawn from liquid families potentially suitable for EV immersion cooling. Active learning drove the FOM with 15 data points-achieving a tenfold improvement in computational efficiency compared to random sampling. The resulting model was predictive (R2 = 0.99) across diverse dielectric fluids and binary mixtures, with silicon oils and essential oils achieving the highest FOM values. This work provides a metric-driven approach to rapidly screening and ranking fluids, facilitating the accelerated discovery of high-performance thermal fluids tailored to specific EV applications.
Artificial intelligence (AI) is transforming materials research in metal-organic frameworks (MOFs), where models trained on structured computational data routinely predict new materials and optimize their properties. This raises a central question: What if we could leverage the full breadth of MOF knowledge, not just structured datasets, but also the scientific literature? For human researchers, the literature remains the primary source of knowledge, yet much of its content, including experimental data and expert insight, remains underutilized by AI systems. We introduce MOF-ChemUnity, a structured, extensible, and scalable knowledge graph that unifies MOF chemical data by linking literature-derived insights to crystal structures and computational datasets. By disambiguating MOF names in the literature and connecting them to crystal structures in the Cambridge Structural Database, MOF-ChemUnity unifies experimental and computational sources and enables cross-document knowledge extraction and linking. We showcase how this enables multi-property machine learning across simulated and experimental data, compilation of complete synthesis records for individual compounds by aggregating information across multiple publications, and expert-guided materials recommendations via structure-based embeddings. When used as a knowledge source to augment large language models (LLMs), MOF-ChemUnity enables a literature-informed AI assistant that operates over the full scope of MOF knowledge. Expert evaluations show improved accuracy, interpretability, and trustworthiness across tasks such as retrieval, inference of structure-property relationships, and materials recommendation, outperforming standard LLMs. This work lays the foundation for literature-informed materials discovery, enabling both human scientists and AI systems to reason over the full landscape of MOF knowledge in a new way.
Large language models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 32 total projects developed during the second annual LLM hackathon for applications in materials science and chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.
Artificial intelligence (AI) is transforming research in metal-organic frameworks (MOFs), where models trained on structured computational data routinely predict new materials and optimize their properties. This raises a central question: What if we could leverage the full breadth of MOF knowledge, not just structured data sets, but also the scientific literature? For researchers, the literature remains the primary source of knowledge, yet much of its content, including experimental data and expert insight, remains underutilized by AI systems. We introduce MOF-ChemUnity, a structured, extensible, and scalable knowledge graph that unifies MOF data by linking literature-derived insights to crystal structures and computational data sets. By disambiguating MOF names in the literature and connecting them to crystal structures in the Cambridge Structural Database, MOF-ChemUnity unifies experimental and computational sources and enables cross-document knowledge extraction and linking. We showcase how this enables multiproperty machine learning across simulated and experimental data, compilation of complete synthesis records for individual compounds by aggregating information across multiple publications, and expert-guided materials recommendations via structure-based machine learning descriptors for pore geometry and chemistry. When used as a knowledge source to augment large language models (LLMs), MOF-ChemUnity enables a literature-informed AI assistant that operates over the full scope of MOF knowledge. Expert evaluations show improved accuracy, interpretability, and trustworthiness across tasks such as retrieval, inference of structure-property relationships, and materials recommendation, outperforming standard LLMs. This work lays the foundation for literature-informed materials discovery, enabling both scientists and AI systems to reason over the full existing knowledge in a new way.
In this work, we present the latest advancements from our PrISMa (Process-Informed design of tailor-made Sorbent Materials) platform, where we seamlessly connect quantum calculations, molecular simulations, process design, techno-economic assessment (TEA), and life cycle assessment (LCA), to provide insights and guide the selection of optimal sorbent-based capture technologies. The performance of 1200 Metal-Organic Frameworks (MOFs) materials for over 60 case studies are investigated, covering different CO2 sources, regions, and technologies. We demonstrate how the PrISMa platform can inform multiple stakeholders about key issues of most interest in carbon capture applications. This holistic approach serves to de-risk investments and establishes a common basis for identifying the optimal path forward in carbon capture technologies.
Every year, researchers create hundreds of thousands of new materials, each with unique structures and properties. For example, over 5000 new metal-organic frameworks (MOFs) were reported in the past year alone. While these materials are often synthesized for specific applications, they may have potential uses in entirely different domains. However, linking these new materials to their best applications remains a significant challenge. In this study, we demonstrate a multimodal approach that uses the information available as soon as a MOF is synthesized, specifically its powder X-ray diffraction pattern (PXRD) and the chemicals used in its synthesis, to predict its potential properties and uses. By self-supervised pretraining of this model on crystal structures accessible from MOF databases, our model achieves accurate predictions for various properties, across pore structure, chemistry-reliant, and quantum-chemical properties, even when small data is available. We further assess the robustness of this method in the presence of experimental measurement imperfections. Utilizing this approach, we create a synthesis-to-application map for MOFs, offering insights into optimal material classes for diverse applications. Finally, by augmenting this model with a recommendation system, we identify promising MOFs for applications that are different from the originally reported applications. We provide this tool as an open source code and a web app to accelerate the matching of new materials with their potential industrial applications.