Autonomous systems integrating machine learning (ML) and laboratory automation are transforming synthetic chemistry by enabling closed-loop experimentation and discovery. In this review, we examine the state-of-the-art in autonomous systems for organic synthesis, with a focus on the components, configurations, and ML algorithms that enable automated reaction planning, execution, and optimization. We survey representative systems that span applications from reaction discovery to molecular optimization, comparing flow and batch configurations and identifying trends in system design. Emphasis is placed on the critical bottlenecks of purification and analytical measurement, particularly structural elucidation of unexpected products-areas that currently constrain autonomous platforms. We describe recent advances in chromatographic method development, structural elucidation from mass spectrometry and nuclear magnetic resonance, and novel ML-based approaches to quantify complex mixtures without calibration. By focusing on enabling technologies in chemical analysis, we identify opportunities for ML and automation to expand beyond domain-specific platforms and accelerate the pace of synthetic discovery.
We demonstrate the usefulness of general atom- and bond-level density functional theory (DFT) descriptors to enhance the performance of neural networks for general reaction condition prediction. We treat condition prediction as a multiclass classification task and report the performance of neural networks and random forests as evaluated by 5-fold cross-validation on a 69,935 reaction data set with 296 distinct single-component reaction condition classes and varying input embedding compositions. We show that by combining structural and general DFT descriptors, models with up to 71% fewer trainable parameter than their purely structural counterparts can provide comparable or superior weighted precision, top-1 and top-3 accuracies. Moreover, we report improvements of up to 5, 10, and 11% in weighted precision, top-1 accuracy and F1 score, respectively, for neural networks trained on hybrid representations which combine general DFT and structural descriptors, when compared to structural models with equivalent architectures and input sizes. Remarkably, the best performing neural network trained on hybrid embeddings outperforms the best purely structural model investigated despite the latter benefiting from of an embedding strategy with 267 times more data points than the one used for generating and embedding hybrid descriptors, with both strategies being unsupervised learning algorithms that share considerable conceptual and architectural similarities.
ConspectusThe advancement of machine learning and the availability of large-scale reaction datasets have accelerated the development of data-driven models for computer-aided synthesis planning (CASP) in the past decade. In this Account, we describe the range of data-driven methods and models that have been incorporated into the newest version of ASKCOS, an open-source software suite for synthesis planning that we have been developing since 2016. This ongoing effort has been driven by the importance of bridging the gap between research and development, making research advances available through a freely available practical tool. ASKCOS integrates modules for retrosynthetic planning, modules for complementary capabilities of condition prediction and reaction product prediction, and several supplementary modules and utilities with various roles in synthesis planning. For retrosynthetic planning, we have developed an Interactive Path Planner (IPP) for user-guided search as well as a Tree Builder for automatic planning with two well-known tree search algorithms, Monte Carlo Tree Search (MCTS) and Retro*. Four one-step retrosynthesis models covering template-based and template-free strategies form the basis of retrosynthetic predictions and can be used simultaneously to combine their advantages and propose diverse suggestions. Strategies for assessing the feasibility of proposed reaction steps and evaluating the full pathways are built on top of several pioneering efforts that we have made in the subtasks of reaction condition recommendation, pathway scoring and clustering, and the prediction of reaction outcomes including the major product, impurities, site selectivity, and regioselectivity. In addition, we have also developed auxiliary capabilities in ASKCOS based on our past and ongoing work for solubility prediction and quantum mechanical descriptor prediction, which can provide more insight into the suitability of proposed reaction solvents or the hypothetical selectivity of desired transformations. For each of these capabilities, we highlight its relevance in the context of synthesis planning and present a comprehensive overview of how it is built on top of not only our work but also of other recent advancements in the field. We also describe in detail how chemists can easily interact with these capabilities via user-friendly interfaces. ASKCOS has assisted hundreds of medicinal, synthetic, and process chemists in their day-to-day tasks by complementing expert decision making and route ideation. It is our belief that CASP tools are an important part of modern chemistry research and offer ever-increasing utility and accessibility.
The advancement of machine learning and the availability of large-scale reaction datasets have accelerated the development of data-driven models for computer-aided synthesis planning (CASP) in the past decade. Here, we detail the newest version of ASKCOS, an open source software suite for synthesis planning that makes available several research advances in a freely available, practical tool. Four one-step retrosynthesis models form the basis of both interactive planning and automatic planning modes. Retrosynthetic planning is complemented by other modules for feasibility assessment and pathway evaluation, including reaction condition recommendation, reaction outcome prediction, and auxiliary capabilities such as solubility prediction and quantum mechanical descriptor prediction. ASKCOS has assisted hundreds of medicinal, synthetic, and process chemists in their day-to-day tasks, complementing expert decision making. It is our belief that CASP tools like ASKCOS are an important part of modern chemistry research, and that they offer ever-increasing utility and accessibility.
This manuscript presents machine learning models for Pd-catalyzed C-N couplings constructed using a large, pharmaceutically relevant, structurally diverse dataset (4204 unique products) generated de novo using high-throughput experimentation. The dataset generation was enabled by the discovery of novel nanomole scale compatible automation friendly C-N coupling reaction conditions using LiOTMS as the base. The large dataset enabled the systematic evaluation of model performance using five different data-splitting strategies that were carefully designed to assess the models' ability to both interpolate and extrapolate. The models exhibit high predictive performance across all splits as gauged by standard metrics. In addition, the models predicted with high accuracy the outcome of validation libraries that were outside the scope of the training set. Employing these models in the context of medicinal chemistry campaigns should result in significant enrichment of successful C-N couplings.
Carbohydrates are an abundant, inexpensive and renewable biomass feedstock that could be a cornerstone for sustainable chemical manufacturing, but scalable and environmentally friendly methods that leverage these feedstocks are lacking. For example, 1-allyl sorbitol is the foundational building block for the polypropylene clarifying agent Millad NX 8000, which is produced on the multi-metric ton scale annually, but the manufacturing process at present requires superstoichiometric amounts of tin(1,2). The NX 8000 additives dominate about 80% of the global clarified polypropylene market(3) and are used in concentrations of 0.01-1% during polypropylene production to improve its transparency and resistance to high temperatures, translating to 300-30,000 metric tons annually. The market volume of polypropylene in 2022 was approximately 79.01 million metric tons (MMT), with demand expected to rise by nearly 33% to 105 MMT by 2030 (ref. 4). The cost and sustainability benefits of clarified polypropylene are driving this demand, necessitating more clarifying agents(5). Here we report a high-yielding allylation of unprotected carbohydrates in water using a catalytic amount of indium metal and either allylboronic acid or the pinacol ester (allylBpin) as donors. Aldohexoses, aminohexoses, ketohexoses and aldopentoses are all allylated in high yield under mild conditions and the indium metal is recoverable and reusable with no loss of catalytic activity. Leveraging these features, this process was translated to a scalable continuous synthesis of 1-allyl sorbitol in flow(6) with high yield and productivity through Bayesian optimization of reaction parameters.
The identification of suitable reaction conditions is a crucial step in organic synthesis. Computer- aided synthesis planning promises to improve the efficiency of chemistry and enable robot-assisted workflows, but there remains a gap in bridging computational tools with experimental execution due to the challenge of reaction condition prediction. The conditions used to carry out a reaction consist of qualitative details, such as the discrete identities of “above-the-arrow” agents (catalysts, additives, solvents, etc.) as well as quantitative details, such as temperature and concentrations of both reactants (product contributing) and agents. These procedural aspects of organic chemistry exert a direct influence over the outcome of a chemical transformation and must be provided in any hypothetical autonomous synthesis workflow. In this work, we push beyond qualitative reaction condition recommendation by developing a data-driven framework that incorporates quantitative details, specifically equivalence ratios. We frame the condition recommendation problem as four sub-tasks: predicting agent identities, reaction temperature, reactant amounts, and agent amounts, and evaluate our model accordingly. We demonstrate improved performance over popularity and nearest neighbor baselines and highlight the model’s practical utility for predicting conditions in diverse reaction classes via representative case studies.
Different experiments of differing fidelities are commonly used in the search for new drug molecules. In classic experimental funnels, libraries of molecules undergo sequential rounds of virtual, coarse, and refined experimental screenings, with each level balanced between the cost of experiments and the number of molecules screened. Bayesian optimization offers an alternative approach, using iterative experiments to locate optimal molecules with fewer experiments than large-scale screening, but without the ability to weigh the costs and benefits of different types of experiments. In this work, we combine the multifidelity approach of the experimental funnel with Bayesian optimization to search for drug molecules iteratively, taking full advantage of different types of experiments, their costs, and the quality of the data they produce. We first demonstrate the utility of the multifidelity Bayesian optimization (MF-BO) approach on a series of drug targets with data reported in ChEMBL, emphasizing what properties of the chemical search space result in substantial acceleration with MF-BO. Then we integrate the MF-BO experiment selection algorithm into an autonomous molecular discovery platform to illustrate the prospective search for new histone deacetylase inhibitors using docking scores, single-point percent inhibitions, and dose-response IC50 values as low-, medium-, and high-fidelity experiments. A chemical search space with appropriate diversity and fidelity correlation for use with MF-BO was constructed with a genetic generative algorithm. The MF-BO integrated platform then docked more than 3,500 molecules, automatically synthesized and screened more than 120 molecules for percent inhibition, and selected a handful of molecules for manual evaluation at the highest fidelity. Many of the molecules screened have never been reported in any capacity. At the end of the search, several new histone deacetylase inhibitors were found with submicromolar inhibition, free of problematic hydroxamate moieties that constrain the use of current inhibitors.
Functionalization of lead compounds to create analogs is a challenging step in discovering new molecules with desired properties and it is conducted throughout the chemical industry, including pharmaceuticals and agrochemicals. The process can be time-consuming and expensive, requiring expert intuition and experience. To help address synthesis planning challenges in late-stage functionalization, we have developed a molecular similarity approach that proposes single-step functionalization reactions based on analogy to precedent reactions. The developed approach mimics reaction strategies and suggests co-reactants defined implicitly by a corpus of known reactions. Using ca. 348 k reactions from the patent literature as a knowledge base, the recorded products or close analogs are among the top 20 proposed products in 74% of ∼44 k test reactions. The combinatorial growth inherent in recursive applications of the tool allows the enumeration of chemical libraries surrounding a target compound of interest. Moreover, each step of the resulting library synthesis leverages common chemical transformations reported in the literature accessible to most chemists.
Electrochemical C-H oxidation reactions offer a sustainable route to functionalize hydrocarbons, yet the identification of competent substrates and their synthesis optimization remains challenging. Here, we report an integrated approach combining machine learning (ML) and large language models (LLMs) to streamline the exploration of electrochemical C-H oxidation reactions. Utilizing a batch rapid screening electrochemical platform, we evaluated a wide range of reactions, initially classifying substrates by their reactivity, while LLMs text-mined literature data to augment the training set. The resulting ML models, one for reactivity prediction and the other one for site selectivity, both achieved high accuracy (>90%) and enabled virtual screening of a large set of commercially available molecules. To optimize reaction conditions of substrates of interest upon the screening, LLMs were prompted to generate code to iteratively improve yield, lowering the barrier for scientists to access ML programs, and this strategy efficiently identified high-yield conditions for eight drug-like substances or intermediates. Notably, we benchmarked the accuracy and reliability of 10 different LLMs, including llama, Claude, and GPT-4, on generating and executing codes related to ML based on natural language prompts given by chemists to showcase their tool-making and tool-using capabilities and potentials for accelerating research across four diverse tasks. In addition, we collected an experimental benchmark dataset comprising 1071 reaction conditions and yields for electrochemical C-H oxidation reactions, and our findings revealed that integrating LLMs and ML outperformed using either method alone. We envision that this combined approach offers a robust and generalizable pathway for advancing synthetic chemistry research
The mechanism of Pd-catalyzed amination of five-membered heteroaryl halides was investigated by integrating experimental kinetic analysis with kinetic modeling through predictive testing and likelihood ratio analysis, revealing an atypical productive coupling pathway and multiple off-cycle events. The GPhos-supported Pd catalyst, along with the moderate-strength base NaOTMS, was previously found to promote efficient coupling between five-membered heteroaryl halides and secondary amines. However, slight deviations from the optimal concentration, temperature, and/or solvent resulted in significantly lower yields, contrary to typical reaction optimization trends. We found that the coupling of 4-bromothiazole with piperidine proceeds through an uncommon mechanism in which the NaOTMS base, rather than the amine, binds first to the oxidative addition complex; the resulting OTMS-bound Pd species is the resting state. Formation of the Pd-amido complex via base/amine exchange was identified as the turnover-limiting step, unlike other reported catalyst systems for which reductive elimination is turnover-limiting. We determined that the amine-bound Pd complex, usually an on-cycle intermediate, is instead a reversibly generated off-cycle species, and that base-mediated decomposition of 4-bromothiazole is the primary irreversible catalyst deactivation pathway. Predictive testing and kinetic modeling were key to the identification of these off-cycle processes, providing insight into minor mechanistic pathways that are difficult to observe experimentally. Collectively, this report reveals the unique enabling features of the Pd-GPhos/NaOTMS system, implementing mechanistic insights to improve the yields of particularly challenging coupling reactions. Moreover, these findings highlight the utility of applying predictive tests to kinetic models for the rapid evaluation of mechanistic possibilities in small-molecule catalytic systems.
Executive Editor Maria Southall welcomes Dionisios Vlachos as the new Editor-in-Chief of Reaction Chemistry & Engineering and pays tribute to the leadership and many contributions of departing inaugural Editor-in-Chief Klavs Jensen.
Reaction screening and high-throughput experimentation (HTE) coupled with liquid chromatography (HPLC, UHPLC) are becoming more important than ever in synthetic chemistry. With growing number of experiments, it is increasingly difficult to ensure correct peak identification and integration, especially due to unknown side components which often overlap with the peaks of interest. We developed a comprehensive Python package with web-based graphical user interface (GUI) for automated processing of chromatograms, including baseline correction, intelligent peak picking, peak purity checks, deconvolution of overlapping peaks, and compound tracking. The algorithm accuracy was benchmarked using three datasets and compared to the previous MOCCA implementation and published results. The processing is fully automated with the possibility to include calibration and internal standards. The software supports chromatograms with photo-diode array detector (DAD) data from most commercial HPLC systems, and the Python package and GUI implementation are open-source to allow addition of new features and further development.
A new method, named dynamic experiment optimization (DynO), is developed for the current needs of chemical reaction optimization by leveraging for the first time both Bayesian optimization and data-rich dynamic experimentation in flow chemistry. DynO is readily implementable in automated systems and it is augmented with simple stopping criteria to guide non-expert users in fast and reagent-efficient optimization campaigns. The developed algorithms is compared in silico with the algorithm Dragonfly and an optimizer based on random selection, showing remarkable results in Euclidean design spaces superior to Dragonfly. Finally, DynO is validated with an ester hydrolysis reaction on an automated platform showcasing the simplicity of the method.
Collaboration between synthesis laboratories requires procedures that are reproducible despite differences in equipment. Now, a digital standard for automated chemical synthesis reproduces results between distinct laboratory systems almost half a world apart.
Reaction optimization and characterization depend on reliable measures of reaction yield, often measured by high-performance liquid chromatography (HPLC). Peak areas in HPLC chromatograms are correlated to analyte concentrations by way of calibration standards, typically pure samples of known concentration. Preparing the pure material required for calibration runs can be tedious for low-yielding reactions and technically challenging at small reaction scales. Herein, we present a method to quantify the yield of reactions by HPLC without needing to isolate the product(s) by combining a machine learning model for molar extinction coefficient estimation, and both UV-vis absorption and mass spectra. We demonstrate the method for a variety of reactions important in medicinal and process chemistry, including amide couplings, palladium catalyzed cross-couplings, nucleophilic aromatic substitutions, aminations, and heterocycle syntheses. The reactions were all performed using an automated synthesis and isolation platform. Calibration-free methods such as the presented approach are necessary for such automated platforms to be able to discover, characterize, and optimize reactions automatically.
We present an automated droplet reactor platform possessing parallel reactor channels and a scheduling algorithm that orchestrates all of the parallel hardware operations and ensures droplet integrity as well as overall efficiency. We design and incorporate all of the necessary hardware and software to enable the platform to be used to study both thermal and photochemical reactions. We incorporate a Bayesian optimization algorithm into the control software to enable reaction optimization over both categorical and continuous variables. We demonstrate the capabilities of both the preliminary single-channel and parallelized versions of the platform using a series of model thermal and photochemical reactions. We conduct a series of reaction optimization campaigns and demonstrate rapid acquisition of the data necessary to determine reaction kinetics. The platform is flexible in terms of use case: it can be used either to investigate reaction kinetics or to perform reaction optimization over a wide range of chemical domains.
The Community Resource for Innovation in Polymer Technology (CRIPT) data model is designed to address the high complexity in defining a polymer structure and the intricacies involved with characterizing material properties.
The goals of this Perspective are threefold: (1) to inform a broad audience, including machine learning (ML) and artificial intelligence (AI) academics and professionals, about synthetic drug substance process development, (2) to break down the general synthetic drug substance process development task into more tractable subtasks, and (3) to highlight areas in which machine learning and artificial intelligence might be beneficially developed and applied. Application of machine learning and artificial intelligence to chemical synthesis of medicinal compounds has long been discussed and has resulted in the development of a number of computer-aided synthesis planning tools by both academic groups and commercial enterprises. The focus of these efforts has primarily centered on the task of retrosynthetic analysis, as seen from the perspective of a medicinal chemist. This has left significant unrealized opportunities in the application of machine learning and artificial intelligence to aid the process chemist or engineer in commercial drug substance process development.