The design of small molecules with tailored properties is a central goal in chemistry and materials science. Recent advances in machine learning provide powerful tools to accelerate the pace of discovery. One promising avenue for acceleration involves the use of generative models that propose novel candidates for diverse optimization tasks. Despite their promise, these methods are often evaluated solely using computational benchmarks, and many studies fail to advance proposed candidates to experimental validation in the wet lab. A key reason for this gap, the elephant in the room, is the limited synthesizability of the generated molecules. In response, the community has recently developed various strategies to address this challenge and incorporate synthesizability into generative design workflows. In this opinion, we provide a comprehensive overview of recent contributions that explicitly tackle molecular synthesizability, highlighting notable advances. We also discuss key limitations of current approaches and outline promising directions for future research.
Accurate crystal structure prediction (CSP) is essential for discovering novel materials. Although various CSP methods have been developed, systematic benchmarks and quantitative comparisons remain limited. In this study, we evaluate eight CSP approaches: the evolutionary algorithm USPEX, three versions of ab initio random sampling with symmetry constraints (USPEX, AIRSS, and PyXtal codes), generative machine learning models (generative adversarial neural network, GAN, and variational autoencoder, VAE), and two very different template-based structure generators (random topological structure generator of USPEX and CSPML). These tests are done through a case study on vanadium oxide systems V5O8 and V3O4. Our results show that most of the compared methods are capable of identifying low-energy and metastable structures when sufficient sampling is performed. However, each method exhibits distinct strengths and trade-offs in terms of accuracy, efficiency, structural diversity, and symmetry character. Notably, more established traditional methods, such as USPEX and AIRSS, offer robust performance across system sizes, while ML-based approaches demonstrate rapid structure generation with minimal sampling, albeit with a greater reliance on the quality of the training data and post-processing approaches. Interestingly, different versions of random sampling show very different performances. This study underscores the complementary nature of traditional and ML-based CSP strategies and provides practical guidance for selecting appropriate methods based on the complexity of the system, the computational resources available, and the specific discovery objectives.
Systematically exploring the multidimensional parameter space of metal-organic framework (MOF) crystallization remains challenging due to limited adoption of high-throughput (HT), automated experimental workflows. MOF development is dominated by manual synthesis and characterization methods and a trial-and-error approach, and the integration of HT MOF synthesis with HT characterization and analysis is uncommon. Here, we present a practical HT MOF discovery workflow that combines automated solvothermal synthesis with three scalable characterization methods. First, we characterize bulk structure using HT powder X-ray diffraction (PXRD) and rapid matching of PXRD data to reported MOF crystal structures. We also employ a custom machine-learning based computer vision (CV) model to identify MOF formation rates from images of sample vials. Finally, we develop a HT X-ray fluorescence (XRF) method to quantify elemental ratios in bimetallic MOF samples. As a case study, we investigate the crystallization of rare earth (RE) MOFs, systematically probing the effects of reaction conditions such as metal identity, linker structure, temperature, and acid concentration. We then leverage these insights to demonstrate a proof-of-concept selective crystallization from a mixed RE solution. Using our HT workflow, we performed 1488 unique MOF crystallization reactions and characterized the resulting samples through the collection of >800 PXRD patterns, CV analysis of >13 000 images, and elemental analysis measurements of 144 bimetallic crystallization reactions. We identified 5 previously unreported rare earth MOFs (NU-2501-NU-2505) and characterized their structures with single-crystal X-ray diffraction (SCXRD) and microcrystal electron diffraction (MicroED). Our HT approach enabled us to construct phase diagrams mapping out the crystallization preferences and formation kinetics for 18 distinct RE-MOF products. By unifying automated MOF synthesis with multimodal characterization, we demonstrate the efficient exploration of a complex synthetic landscape, generating insights into MOF structure, crystallization kinetics, and composition.
Alán Aspuru-Guzik is a professor of chemistry and computer science at the University of Toronto, a Canada 150 Research Chair in Theoretical Chemistry, CIFAR AI Chair at the Vector Institute, and director of the Acceleration Consortium. His research advances self-driving laboratories, quantum information, machine learning, molecular design, and autonomous discovery for energy and electronics. He co-founded Zapata AI, Kebotix, and Axiomatic AI; is the editor-in-chief of Digital Discovery; and is an APS, AAAS, and Royal Society of Canada fellow. Varinia Bernales is an assistant professor in computer science and chemistry at the University of Toronto and co-director of the Matter Lab, where she develops autonomous research agents for molecular and material discovery. Her work integrates quantum chemistry; machine learning; scientific automation; and artificial intelligence (AI) safety, alignment, and governance with applications in catalysis, materials, and chemistry. Prior to her current role, she worked at Dow Chemical and Underwriters Laboratories Research Institutes.
Large language models (LLMs) can plan scientific workflows and generate code, but these capabilities do not specify how scientific state is validated, transferred and recorded across heterogeneous computational and experimental operations. Here we present El Agente Gráfico, a semantic execution runtime for scientific agents that uses typed execution graphs to enforce admissible scientific state transitions, record provenance and confine model judgement to explicit decision points. Using the same top-level LLM and task-specific rubrics on six university-level quantum chemistry exercises, El Agente Gráfico improved performance while reducing model cost by approximately 80
Generative models for matter are often evaluated as samplers over output representations, and their latent spaces are commonly used as proxies for navigating chemical space. Much less is known about how these models internally arrange discrete chemical identities within those representations. We study this arrangement by making molecular identity explicit and pulling it back through the generative process. Through these pullbacks we probe the regions that generate the same object, exposing the trained model's internal repertoire: a fixed partition that determines which objects (novel or not) the model can produce. Across three molecular generative architectures, we find that this repertoire is arranged into piecewise-constant regions separated by recurring coarse-to-fine boundaries. Its organization depends on the representation probed, the identity convention, decoder stochasticity, and the metric used to compare coordinates. During training, local chemical organization stabilizes while the number of distinct molecular identities represented within each neighborhood continues to change. Internal organization must therefore be characterized, rather than assumed, before a generative space can be treated as chemically navigable.
Coupled-cluster (CC) theory is often considered the gold standard of quantum chemistry, but its high computational cost limits routine access to accurate energies, forces and response properties. While the right-hand T-amplitudes determine the correlated wavefunction, many practically important observables additionally require the left-hand Λ-amplitudes. We introduce MōLe-Λ, an extension of Molecular Orbital Learning (MōLe) that predicts the full ground-state coupled-cluster singles and doubles (CCSD) response state by jointly learning right-hand amplitudes (T_1,T_2) and left-hand amplitudes (Λ_1,Λ_2) from localized Hartree–Fock molecular orbitals. Architecturally, MōLe-Λ extends MōLe with Λ_1 and Λ_2 readouts that mirror the symmetry constraints of the T_1 and T_2 heads, while preserving the original equivariant orbital encoder, odd sign-equivariant decoding, locality and size-extensivity. The resulting model yields accurate CC-quality energies and forces, while simultaneously recovering dipoles, quadrupoles, polarizabilities, the electron density, and 2-electron observables such as the pair density. We show that MōLe-Λ further extends the speed advantage of MōLe over full CCSD while substantially expanding the accessible properties, providing a route to wavefunction-level surrogate models for correlated quantum chemistry.
Precise liquid handling is an essential operation for self-driving laboratories. In 2023, we introduced the digital pipette, a low-cost, 3D-printed device that enables accurate liquid transfer by robotic arms. However, the initial version lacked mechanisms to prevent cross-contamination when handling multiple liquids. In this commit paper, we present the digital pipette v2, an updated design that mitigates contamination risk by allowing robotic arms to exchange pipette tips. The new hardware achieves liquid handling accuracy within the permissible error range defined by ISO 8655-2, supporting a broader range of experiments involving multiple liquids.
We have recently demonstrated the ability of using self-driving laboratories for AI-driven searches of organic emitters for solid-state lasing devices. Our past workflow featured solubility challenges for such large molecular moieties. In this next-generation study, we return to the drawing board to explore a family of compounds that are much solution processable and composed of a set of electronic cores that provide a broader color response. Out of 252 potential candidates, and with guidance from DFT calculations, we selectively perform a comprehensive study exploring 51 fluorene-based A-B-A-type organic laser oligomers, armed with our self-driving lab. The candidates range from simple hydrocarbon molecules to complex heteroatom-mixed molecules. As a result of this study, we highlight diketopyrrolopyrrole and benzodiazole derivatives for their largely red-shifted emissions. Furthermore, we investigate the effect of color change arising from heteroatom permutation, fluorine addition, thiophene coupling, and a combination of fluorine addition and thiophene coupling. Amplified spontaneous emission (ASE) measurements in the solid state further corroborate the lasing potential of selected candidates, reinforcing their suitability for future device applications. The computational study with density functional theory confirms the experimental results.
Molecular post-modification design strategies that enable low-temperature pyrolysis of polystyrene (PS) remain an underexplored area. Conventional pyrolysis of PS demands heating above 400 °C, creating economic barriers to commercial-scale monomer recovery. Here, we demonstrate the post-functionalization of the PS backbone with a labile C-S bond, specifically a trifluoromethylthio group (-SCF3), to accelerate the depolymerization of PS at lower temperatures. A previously established small-molecule trifluoromethylthiolation reaction was adapted to PS through solvent screening and reaction optimization. Across a wide range of molecular weights (M n = 1.12-110 kg mol-1), including consumer-grade samples, thermogravimetric analysis demonstrates that PS-SCF3 exhibits an onset degradation temperature 10-20 °C lower and a greater mass loss of 10-35% over 20 hours at 300 °C compared to pristine PS. Flynn-Ozawa-Wall analysis reveals that the average apparent activation energy for depolymerization of PS-SCF3 is approximately 11 kJ mol-1 lower than that of pristine PS. To assess the potential industrial relevance of this protocol, pyrolysis of several consumer-grade PS samples and their post-modified PS-SCF3 analogues was performed at 300 °C; PS-SCF3 samples were found to afford higher styrene recovery relative to pristine PS. This study explores the potential of backbone post-functionalization of PS as a strategy to accelerate depolymerization at lower temperatures and shorter timescales, enabling greater styrene recovery and advancing progress toward a circular economy for plastics.
Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for this task, yet no systematic evaluation exists of how well vision-language models (VLMs) interpret them. We introduce QCalEval, the first VLM benchmark for quantum calibration plots: 243 samples across 87 scenario types from 22 experiment families, spanning superconducting qubits and neutral atoms, evaluated on six question types in both zero-shot and in-context learning settings. The best general-purpose zero-shot model reaches a mean score of 72.3, and many open-weight models degrade under multi-image in-context learning, whereas frontier closed models improve substantially. A supervised fine-tuning ablation at the 9-billion-parameter scale shows that SFT improves zero-shot performance but cannot close the multimodal in-context learning gap. As a reference case study, we release NVIDIA Ising Calibration 1, an open-weight model based on Qwen3.5-35B-A3B that reaches 74.7 zero-shot average score.
Quantum chemistry calculations are a key component of the materials discovery process. The results from first-principles explorations enable the prediction of material properties prior to experimental validation. Despite their impact, the practical use of first-principles methods remains limited by the expertise required to design, execute, and troubleshoot complex computational workflows. Even when workflows are successfully built, they are sometimes rigid and not adaptable to different use cases. Recent advances in large language models (LLMs) and agentic systems offer a pathway to flexibly automate these processes and lower barriers to entry. Here, we introduce El Agente Sólido, a hierarchical multi-agent framework for automating solid-state quantum chemistry workflows using the open-source Quantum ESPRESSO simulation package. The framework translates high-level scientific objectives expressed in natural language into end-to-end computational pipelines that include structure generation, input file construction, workflow execution, and post-processing analysis. El Agente Sólido integrates density functional theory with phonon calculations and machine-learning interatomic potentials to enable efficient and physically consistent simulations. Extensive benchmarking and case studies demonstrate that El Agente Sólido reliably executes a wide range of solid-state calculations, highlighting its potential to improve reproducibility and accelerate computational materials discovery
The extracellular matrix (ECM) exhibits tissue-specific viscoelasticity with a unique combination of elasticity and viscosity. Hydrogels with independently controlled elastic modulus and stress relaxation have been developed. Yet, time- and labor-efficient identification of formulations used for the preparation of multicomponent hydrogels recapitulating the mechanical properties of a specific tissue remains a challenge. Conventional bottom-up screening of hydrogel formulations is resource-intensive, especially for navigating multicomponent hydrogels. Here, we present an active learning framework based on multi-objective Bayesian optimization to accelerate the discovery of biomimetic fibrous hydrogels replicating the elastic and viscous properties of the ECM of healthy and diseased tissues. In the data-driven approach, we iteratively navigated the multidimensional chemical space, while simultaneously optimizing hydrogel's conflicting mechanical properties and mapping the boundaries of material feasibility using sparse experimental feedback. The experimental data were used to train a digital twin model and unveil the interplay between the fabrication conditions and hydrogel properties. The decoupled effects of the elasticity and viscosity were established for cell activation, proliferation, and morphogenesis. This work shows the potential of using the data-driven approach, leveraging even sparse experimental data to engineer hydrogels and optimize diverse biomaterials for tissue engineering and regenerative medicine.
Advances in generative artificial intelligence are transforming how metal-organic frameworks (MOFs) are designed and discovered. This Perspective introduces the shift from laborious enumeration of MOF candidates to generative approaches that can autonomously propose and synthesize in the laboratory new porous reticular structures on demand. We outline the progress of employing deep learning models, such as variational autoencoders, diffusion models, and large language model-based agents, that are fueled by the growing amount of available data from the MOF community and suggest novel crystalline materials designs. These generative tools can be combined with high-throughput computational screening and even automated experiments to form accelerated, closed-loop discovery pipelines. The result is a new paradigm for reticular chemistry in which AI algorithms more efficiently direct the search for high-performance MOF materials for clean air and energy applications. Finally, we highlight remaining challenges such as synthetic feasibility, dataset diversity, and the need for further integration of domain knowledge.
Discrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align composed diffusion dynamics with the intended target. However, these methods are fundamentally limited by working on a per-sample basis, treating each generated state monolithically and ignoring the potential spatial or functional specializations of different experts. In this work, we address this limitation by proposing FactorDiff - a factor-wise composition framework for diffusion models. We posit that samples can be further decomposed into smaller factors, and propose a sampling process that dynamically routes each factor to the most relevant expert. We instantiate this framework with spatial/pixel-level compositions and validate it on the ARC-AGI benchmark, demonstrating that simple factor-specific routing consistently outperforms complex global scalar weighting schemes on tasks that require logical consistency and spatial disentanglement.
Discrete diffusion models have recently emerged as a promising alternative to the autoregressive approach for generating discrete sequences. Sample generation via gradual denoising or demasking processes allows them to capture hierarchical non-sequential interdependencies in the data. These custom processes, however, do not assume a flexible control over the distribution of generated samples. We propose Discrete Feynman-Kac Correctors, a framework that allows for controlling the generated distribution of discrete masked diffusion models at inference time. We derive Sequential Monte Carlo (SMC) algorithms that, given a trained discrete diffusion model, control the temperature of the sampled distribution (i.e. perform annealing), sample from the product of marginals of several diffusion processes (e.g. differently conditioned processes), and sample from the product of the marginal with an external reward function, producing likely samples from the target distribution that also have high reward. Notably, our framework does not require any training of additional models or fine-tuning of the original model. We illustrate the utility of our framework in several applications including: efficient sampling from the annealed Boltzmann distribution of the Ising model, improving the performance of language models for code generation and amortized learning, as well as reward-tilted protein sequence generation.
Advances in high-throughput instrumentation and laboratory automation are revolutionizing materials synthesis by enabling the rapid generation of large libraries of novel materials. However, efficient characterization of these synthetic libraries remains a significant bottleneck in the discovery of new materials. Traditional characterization methods are often limited to sequential analysis, making them time-intensive and cost-prohibitive when applied to large sample sets. In the same way that chemists interpret visual indicators to identify promising samples, computer vision (CV) is an efficient approach to accelerate materials characterization across varying scales when visual cues are present. CV is particularly useful in high-throughput synthesis and characterization workflows, as these techniques can be rapid, scalable, and cost-effective. Although there is a set of growing examples in the literature, we have found a lack of resources where newcomers interested in the field could get a hold of a practical way to get started. Here, we aim to fill that identified gap and present a structured tutorial for experimentalists to integrate computer vision into high-throughput materials research, providing a detailed roadmap from data collection to model validation. Specifically, we describe the hardware and software stack required for deploying CV in materials characterization, including image acquisition, annotation strategies, model training, and performance evaluation. As a case study, we demonstrate the implementation of a CV workflow within a high-throughput materials synthesis and characterization platform to investigate the crystallization of metal–organic frameworks (MOFs). By outlining key challenges and best practices, this tutorial aims to equip chemists and materials scientists with the necessary tools to harness CV for accelerating materials discovery.
The growing demand for energy-efficient processes to support a sustainable future drives the need for research to rapidly explore chemical and material space through accelerated catalyst discovery initiatives. Recent breakthroughs in high-throughput experimental and computational methods are transforming the catalysis field, surpassing traditional approaches to manipulating variables in catalytic processes. Key advancements in innovation include the integration of machine learning for efficient catalyst screening, high-throughput experimentation, data-driven methodologies employing comprehensive databases, and in situ and in operando techniques for realistic observations. This progress has undoubtedly been intertwined with a collaborative framework across disciplines, reshaping catalyst discovery methods in both industry and academia. This Opinion article presents a multifaceted perspective from coauthors with expertise spanning various stages of the Technology Readiness Level spectrum, highlighting both opportunities and persistent challenges in integrating computational and experimental approaches in catalysis. These challenges span from obtaining high-quality experimental data, scaling simulations to industrially relevant materials and process conditions to navigating the complexity and predictive accuracy of computational models.
The traditional development of novel metal-organic frameworks (MOFs) is often hindered by challenges such as synthetic accessibility and time- and resource-intensive experimentation. High-throughput, automated experimental and computational techniques have enabled rapid chemical space exploration and theoretical MOF design. When combined with artificial intelligence (AI), these methods can be used to lead autonomous laboratories to new frontiers for MOF discovery, where these materials can be designed for a specific application, efficiently synthesized, characterized, and evaluated. This perspective highlights the role of AI in advancing automated MOF synthesis and characterization, computational MOF design and screening, and the integration of these approaches within autonomous workflows to ultimately enable the MOF laboratories of the future.