Chiral Ni-PHOX complexes in combination with photoredox catalysis are shown to be effective to achieve the asymmetric coupling of amino acids with 2-iodopyridines at the C2-position. The regio- and enantioselective protocol developed herein enables direct access to chiral pyrid-2-yl beta-aminoalcohols, which are found in many active pharmaceutical ingredients. This methodology can be extended to other heteroarenes, such as quinolines and azines, on one gram scale. Computational studies supported by experimentation revealed the trend of ligand steric and electronic influences on enantioselectivity. A plausible mechanism was proposed using DFT to further rationalize the observed regio- and enantioselectivity.
The exponential growth of scientific literature presents an increasingly acute challenge across disciplines. Hundreds of thousands of new chemical reactions are reported annually, yet translating them into actionable experiments becomes an obstacle1,2. Recent applications of large language models (LLMs) have shown promise3-6, but systems that reliably work for diverse transformations across de novo compounds have remained elusive. Here we introduce MOSAIC (Multiple Optimized Specialists for AI-assisted Chemical Prediction), a computational framework that enables chemists to make use of the collective knowledge of millions of reaction protocols. MOSAIC is built on the Llama-3.1-8B-Instruct architecture7, training 2,498 specialized chemical experts in Voronoi-clustered spaces. This approach delivers reproducible and executable experimental protocols with confidence metrics for complex syntheses. With an overall 71% success rate, experimental validation demonstrates the realizations of more than 35 new compounds, spanning pharmaceuticals, materials, agrochemicals and cosmetics. Notably, MOSAIC also enables the discovery of new reaction methodologies that are absent from the expert's training, a cornerstone for advancing chemical synthesis. This scalable model of partitioning vast domains into searchable expert regions enables a generalizable strategy for AI-assisted discovery wherever accelerating information growth outpaces efficient knowledge access and application.
Heterocyclic amines are a key structural motif for the synthesis of pharmaceuticals (e.g., antibiotics) as well as pesticides and flavors. In this regard, imine reductases (IREDs) have recently emerged as a highly selective and sustainable alternative for asymmetric reductive amination reactions. Herein, we have applied six IREDs, two of which were newly identified, in the reduction of heterocyclic imines with either a N, S, or O substitution at C-4. Since IREDs are NADPH-dependent enzymes, a commercially available, supported glucose dehydrogenase was added as a cofactor-regenerating system. IREDs were then immobilized on porous microparticles to further improve the efficiency and sustainability of the system. The strategic combination of bioinformatic analysis and immobilization screening resulted in immobilized biocatalysts with 95% retained activity. This enabled the integration of the bienzymatic system into a continuous-flow reactor leading to >90% conversion of 50 mM of the S-heterocyclic amine, 5-methyl-3,6-dihydro-2H-1,4-thiazine, with a residence time of 30 min, and reaching space-time yields up to 14.3 g L-1 h-1. In addition, (S)- or (R)-stereoselectivity of the biocatalytic reduction of the 1,4-disubstituted heterocyclic imines was achieved by using the newly identified IREDs fromGoodfellowiella coeruleoviolaceaandLabilithrix luteola, respectively.
The enabling synthesis of the first route to HER2 inhibitor BI-4142 (1) to deliver a drug substance in kilogram quantity is reported. The synthetic route involves (1) a fit-for-purpose synthesis of pyrimido[5,4-d]pyrimidine 2; (2) a high yielding, scalable synthesis of aniline 3; (3) a safer sodium tungstate-catalyzed sulfide oxidation; (4) SNAr reactions to form C-N bonds; and (5) amidation via Schotten-Baumann conditions. With the speed of delivery prioritized, a purification protocol using silica gel filtration and crystallizations was developed in time to control the quality of API. The overall yield of the delivery route was improved from 22% to 46% over a prior route starting from 2.
A telescoped flow process was implemented for synthesizing a chiral spiroketone, a pivotal building block of several active pharmaceutical ingredients. This process combined a ring closing metathesis step and a hydrogenation step using one single catalyst (Hoveyda-Grubbs 2nd generation catalyst). This innovative approach offers substantial benefits including cost savings, enhanced throughput with adaptable demo-scale devices, real-time reaction monitoring via process analytical technology, elimination of laborious intermediate separation, decreased process mass intensity, and streamlined unit operations. Notably, the developed telescoped flow process utilizes one single catalyst with a significantly reduced loading, resulting in a remarkable 70% saving on process cost and a 60% decrease in PMI compared to the original batch procedure.
In this work we describe the development of a chemistry-based encoding approach utilizing nucleophilicity to perform Bayesian optimization campaigns. A fully automated slug continuous flow platform leveraging a liquid handler to investigate categorical variables is used for the self-optimization of organic reactions. We compared our chemistry-based approach to a chemistry-agnostic label-encoding approach. The use of encoding a physical property allowed the optimization to proceed rapidly and more successfully than existing methods, identifying not only the correct discrete parameter in the system, but also favorable conditions at the same time. Reactions were analyzed using two complementary process analytical technologies (PATs), Fourier-transform infrared spectroscopy (FT-IR) and ultra high performance liquid chromatography (UHPLC). This approach was applied to two different nucleophile-catalyzed amide coupling reactions, for single and multi-objective optimization. A long run was performed as a comparison to the slug flow operation with the liquid handler-based slug flow reactor.
We present and analyze the results of a pharmaceutical industry-wide survey of the strategies and approaches that companies have for implementing photochemical reactions in discovery chemistry, process development, and commercial manufacturing. The survey questions encompass the types of photochemical reactions that pharmaceutical companies pursue and why, the types of reactors (batch and flow) used, how they are characterized for photochemistry, scale-up of photochemical reactions, and how companies prioritize resources for developing photochemical reactions, among other topics. The survey focuses on many of these topics from the perspectives of discovery chemists and process scientists to highlight similarities and differences between their approaches on the specific photochemical transformations and materials of construction used in photochemical reactors. The survey results clearly demonstrate that photochemistry is a viable synthetic strategy for producing active pharmaceutical ingredients (APIs), from discovery all the way to commercialization. While photochemistry is more prevalent for discovery and early stage process development, the survey results indicate that more companies are leveraging photochemistry in successful scale-ups in later stages of process development, often using flow chemistry, and have significant knowledge resulting from these experiences. Nevertheless, there still are gaps to adopting photochemical reactions across stages of development and companies have very limited experience discussing photochemistry with regulatory agencies. Overall, the survey results demonstrate that photochemistry offers considerable benefits for synthesizing APIs and developing API processes, such as greenness and sustainability, shortening synthetic routes, and potential for improved product quality.
The identification of catalysts that promote chemical reactions is a critical challenge in the production of pharmaceuticals. One of the main bottlenecks in this process is the synthesis of vast libraries of precatalysts, although assessing catalyst effectiveness can be rapidly conducted through high-throughput experimentation. The rational design and development of high-performing precatalysts can circumvent this challenge and lead to important advances. In this study, we apply the transformer-based Kernel-Elastic Autoencoder (KAE) equipped with a conditioned latent space, enabling the targeted generation of ligands with desired steric and electronic properties. Our KAE model has facilitated the identification of a monodentate alkynylphosphine, dubbed MachinePhos A, as an effective precatalyst for forming carbon-carbon bond. Its utility was demonstrated experimentally in the Mizoroki-Heck reaction, using a variety of nitrogen-rich arenes pertinent to pharmaceutical applications.
Retrosynthesis, the strategy of devising laboratory pathways by working backwards from the target compound, is crucial yet challenging. Enhancing retrosynthetic efficiency requires overcoming the vast complexity of chemical space, the limited known interconversions between molecules, and the challenges posed by limited experimental datasets. This study introduces generative machine learning methods for retrosynthetic planning. The approach features three innovations: generating reaction templates instead of reactants or synthons to create novel chemical transformations, allowing user selection of specific bonds to change for human-influenced synthesis, and employing a conditional kernel-elastic autoencoder (CKAE) to measure the similarity between generated and known reactions for chemical viability insights. These features form a coherent retrosynthetic framework, validated experimentally by designing a 3-step synthetic pathway for a challenging small molecule, demonstrating a significant improvement over previous 5-9 step approaches. This work highlights the utility and robustness of generative machine learning in addressing complex challenges in chemical synthesis. Enhancing retrosynthetic efficiency requires overcoming the vast complexity of chemical space, the limited known interconversions between molecules, and the challenges posed by limited experimental datasets. Here, the authors introduce generative machine learning methods for retrosynthetic planning that generate reaction templates.
The development of a scalable asymmetric synthesis of KRAS G12C inhibitor building block 1 is described. The all-carbon quaternary stereocenter was installed enantioselectively via Shi epoxidation, followed by a newly discovered regioselective LaCl3·2LiCl-catalyzed epoxide opening. Subsequent organocatalyzed oxidation provided the requisite ketone, which underwent the final assembly of the heterocyclic core, delivering 1 with high chemical and enantiomeric purities in 40% overall yield in only five steps, enabling robust and rapid manufacturing of over 300 kg of 1.
Screening catalysts is an essential step to identify more sustainable and efficient routes for the desired chemical transformations. While several concepts for streamlining high-throughput screenings are already present in chemocatalysis, biocatalysis requires the more challenging dosing of very small amounts of hygroscopic and electrostatic enzyme powders in large numbers to cover the catalyst space of interest. Within this work, the EnzyBeads technology, where enzymes are transiently bound to an inert, free-flowing support to allow dosing of tiny quantities of enzymes, was further developed and coupled with a parallel volumetric dosing approach in a 96-well format. A simplified coating procedure to coat glass beads, an inert carrier that was previously not usable due to inefficient adsorption, with the enzymes of interest was developed, which circumvents the use of resonance acoustic mixing and of the electrostatic polystyrene beads currently required in the status quo procedure. Using a custom-built 3D printable solid dispenser, named the 12X-Dosing Rack, the obtained enzyme-coated glass beads (G-EnzyBeads) can be dispensed in microtiter plates in a quick and reliable fashion to greatly facilitate biocatalysis screenings.
We report the development of oxoammonium-catalyzed oxidation of N-substituted amines via a hydride transfer mechanism. Steric and electronic tuning of catalyst led to complementary sets of conditions that can oxidize a broad scope of carbamates, sulfonamides, ureas, and amides into the corresponding imides. The reaction was further demonstrated on a 100-g scale using a continuous flow setup.
Self-optimizing flow reactors have received significant attention in recent years, with Bayesian optimization (BO) being identified as the most effective method for reaction optimization. However, there are many different approaches using BO algorithms, which is overwhelming for experimentalists. Here, using pharmaceutically relevant amide coupling reactions, we explore "best practices" in three areas, to promote the efficient design of sustainable processes: (1) A high extent of exploration in an optimization algorithm was deemed necessary to ensure a good design space overview. (2) Yield was optimized within a small experimental budget, while minimizing environmental impact, by setting up an objective function with penalties (e.g., for excess reagent usage). (3) An optimization algorithm using an auxiliary data set appeared to behave well for the same substrates using a different coupling reagent, but provided no advantage when using substrates with substantially lower reactivity. We envisage that these general recommendations will aid flow chemists utilizing BO for automated development of sustainable reactions.
Flow processing offers many opportunities to optimize reactions in a rapid and automated manner, yet often requires relatively large quantities of input materials. To combat this, we report the use of a flexible droplet flow reactor, equipped with two analytical instruments, for low-volume optimization experiments. A Buchwald-Hartwig amination toward the drug olanzapine, with 6 independent optimizable variables, was optimized using three different automated approaches: self-optimization, design of experiments and kinetic modeling. These approaches are complementary and provide differing information on the reaction: pareto optimal operating points, response surface models and mechanistic models, respectively. The results were achieved using <10% of the material that would be required for standard flow operation. Finally, a chemometric model was built utilizing automated data handling and three subsequent validation experiments demonstrated good agreement between the droplet flow reactor and a standard (larger scale) flow reactor.
In this study, we introduce an approach for predicting the enantioselectivity of P-chiral monophosphorus ligands from ligand-based descriptors that can be applied to catalytic systems with small experimental datasets without reliance on mechanistic knowledge. Principal component analysis (PCA) is used to map out the chemical space described by steric and electronic descriptors computed for dihydrobenzooxaphosphole (BOP) and dihydrobenzoazaphosphole (BAP) ligands. The PCA map captures trends in the experimentally measured enantioselectivity of four C-C bond-forming reactions and identifies "hotspots" of selective ligands that provide insight into the optimal balance of sterics and electronics for each reaction. Furthermore, the descriptors are used to train a ridge regression model that quantitatively predicts the enantioselectivity of a Pd-catalyzed Negishi cross-coupling reaction. The coefficients of the model provide fundamental chemical understanding and reveal that a pi-stacking interaction with one of the ligands results in an unexpected selectivity inversion. Overall, this integrated approach combines ligand-based descriptors with small experimental datasets to provide qualitative (PCA) and quantitative (ridge regression) enantioselectivity predictions.
Introduction: Biocatalysis, particularly through engineered enzymes, presents a cost-effective, efficient, and eco-friendly approach to compound synthesis. We sought to identify ketoreductases capable of synthesizing optically pure alcohols or ketones, essential chiral building blocks for active pharmaceutical ingredients.Methods: Using BioMatchMaker®, an in silico high-throughput platform that allows the identification of wild-type enzyme sequences for a desired chemical transformation, we identified a bacterial SDR ketoreductase from Thermus caliditerrae, Tcalid SDR, that demonstrates favorable reaction efficiency and desired enantiomeric excess.Results: Here we present two crystal structures of the Tcalid SDR in an apo-form at 1.9 Å and NADP-complexed form at 1.7 Å resolution (9FE6 and 9FEB, respectively). This enzyme forms a homotetramer with each subunit containing an N-terminal Rossmann-fold domain. We use computational analysis combined with site-directed mutagenesis and enzymatic characterization to define the substrate-binding pocket. Furthermore, the enzyme retained favorable reactivity and selectivity after incubation at elevated temperature.Conclusion: The enantioselectivity combined with the thermostability of Tcalid SDR makes this enzyme an attractive engineering starting point for biocatalysis applications.
The efficient copper-catalyzed construction of imidazopyridinones is described. Beginning from iodopyridinyl ureas, Ullmann coupling provides imidazopyridinones in good yield. Manufacturing by this process is described on 473 kg scale; the scope and limitations of substrates is also detailed.
Density functional theory (DFT) is a powerful tool to model transition state (TS) energies to predict selectivity in chemical synthesis. However, a successful multistep synthesis campaign must navigate energetically narrow differences in pathways that create some limits to rapid and unambiguous application of DFT to these problems. While powerful data science techniques may provide a complementary approach to overcome this problem, doing so with the relatively small data sets that are widespread in organic synthesis presents a significant challenge. Herein, we show that a small data set can be labeled with features from DFT TS calculations to train a feed-forward neural network for predicting enantioselectivity of a Negishi cross-coupling reaction with P-chiral hindered phosphines. This approach to modeling enantioselectivity is compared with conventional approaches, including exclusive use of DFT energies and data science approaches, using features from ligands or ground states with neural network architectures.
An entry from the Cambridge Structural Database, the world’s repository for small molecule crystal structures. The entry contains experimental data from a crystal diffraction study. The deposited dataset for this entry is freely available from the CCDC and typically includes 3D coordinates, cell parameters, space group, experimental conditions and quality measures.