N,N-Bidentate ligands are widely employed in Ni-catalyzed cross-electrophile coupling (CEC) reactions; however, it is often difficult to predict a priori which scaffold will provide optimal selectivity and yield for a reaction under development. More generally, for a given Ni-catalyzed reaction, models that provide structure-reactivity and structure-selectivity relationships across different N,N-bidentate ligand scaffolds remain elusive. Here, we report BUNNY, a density functional theory-based descriptor library of approximately 1100 N,N-bidentate ligands designed to support modeling tasks for Ni-catalyzed cross-coupling. Using BUNNY, a screening set of 31 ligands was selected to represent 7 ligand scaffolds, and the value of the screening set was demonstrated by modeling of two Ni-catalyzed CEC case studies. The first case study demonstrates that enantioselectivity can be modeled across different ligand families for two different benzylic electrophiles. The second case study used the screening library to collect ee data for an established Ni-CEC, which was used to guide development of an enantioselective Ni-catalyzed CEC of a different but related electrophile. Differences in descriptors selected by the enantioselectivity models for the two case studies inspired investigation of the radical capture and reductive elimination steps by DFT, which found that either step can be selectivity-determining, depending on the reaction. This work illustrates how BUNNY can be used to model reactivity and selectivity across ligand scaffolds, guiding reaction development and mechanistic understanding.
Empowering large language models (LLMs) with chemical intelligence remains a challenge due to the scarcity of high-quality, domain-specific instruction-response datasets and the misalignment of existing synthetic data generation pipelines with the inherently hierarchical and rule-governed structure of chemical information. To address this, we propose ChemOrch, a framework that synthesizes chemically grounded instruction–response pairs through a two-stage process: task-controlled instruction generation and tool-aware response construction. ChemOrch enables controllable diversity and levels of difficulty for the generated tasks and ensures response precision through tool planning \& distillation, and tool-based self-repair mechanisms. The effectiveness of ChemOrch is evaluated based on: 1) the \textbf{high quality} of generated instruction data, demonstrating superior diversity and strong alignment with chemical constraints; 2) the \textbf{dynamic generation of evaluation tasks} that more effectively reveal LLM weaknesses in chemistry; and 3) the significant \textbf{improvement of LLM chemistry capabilities} when the generated instruction data are used for fine-tuning. Our work thus represents a critical step toward scalable and verifiable chemical intelligence in LLMs. The code is available at \url{https://anonymous.4open.science/r/ChemOrch-854A}.
Reaction virtual screening and discovery are fundamental challenges in chemistry and materials science, where traditional graph neural networks (GNNs) struggle to model multi-reactant interactions. In this work, we propose ChemHGNN, a hypergraph neural network (HGNN) framework that effectively captures high-order relationships in reaction networks. Unlike GNNs, which require constructing complete graphs for multi-reactant reactions, ChemHGNN naturally models multi-reactant reactions through hyperedges, enabling more expressive reaction representations. To address key challenges, such as combinatorial explosion, model collapse, and chemically invalid negative samples, we introduce a reaction center-aware negative sampling strategy (RCNS) and a hierarchical embedding approach combining molecule, reaction and hypergraph level features. Experiments on the USPTO dataset demonstrate that ChemHGNN significantly outperforms HGNN and GNN baselines, particularly in large-scale settings, while maintaining interpretability and chemical plausibility. Our work establishes HGNNs as a superior alternative to GNNs for reaction virtual screening and discovery, offering a chemically informed framework for accelerating reaction discovery.
Empowering large language models (LLMs) with chemical intelligence remains a challenge due to the scarcity of high-quality, domain-specific instruction-response datasets and the misalignment of existing synthetic data generation pipelines with the inherently hierarchical and rule-governed structure of chemical information. To address this, we propose ChemOrch, a framework that synthesizes chemically grounded instruction-response pairs through a two-stage process: task-controlled instruction generation and tool-aware response construction. ChemOrch enables controllable diversity and levels of difficulty for the generated tasks, and ensures response precision through tool planning and distillation, and tool-based self-repair mechanisms. The effectiveness of ChemOrch is evaluated based on: 1) the high quality of generated instruction data, demonstrating superior diversity and strong alignment with chemical constraints; 2) the reliable generation of evaluation tasks that more effectively reveal LLM weaknesses in chemistry; and 3) the significant improvement of LLM chemistry capabilities when the generated instruction data are used for fine-tuning. Our work thus represents a critical step toward scalable and verifiable chemical intelligence in LLMs.
The development of machine learning models to predict the regioselectivity of C(sp3)-H functionalization reactions is reported. A data set for dioxirane oxidations was curated from the literature and used to generate a model to predict the regioselectivity of C-H oxidation. To assess whether smaller, intentionally designed data sets could provide accuracy on complex targets, a series of acquisition functions were developed to select the most informative molecules for the specific target. Active learning-based acquisition functions that leverage predicted reactivity and model uncertainty were found to outperform those based on molecular and site similarity alone. The use of acquisition functions for data set elaboration significantly reduced the number of data points needed to perform accurate prediction, and it was found that smaller, machine-designed data sets can give accurate predictions when larger, randomly selected data sets fail. Finally, the workflow was experimentally validated on five complex substrates and shown to be applicable to predicting the regioselectivity of arene C-H radical borylation. These studies provide a quantitative alternative to the intuitive extrapolation from "model substrates" that is frequently used to estimate reactivity on complex molecules.
Palladium-catalyzed cross-coupling reactions, particularly the Suzuki-Miyaura coupling, are efficient tools for constructing C-C bonds due to their exceptional versatility and efficiency. Recently, nitroarenes have been explored as new electrophilic substrates in palladium-catalyzed denitrative Suzuki-Miyaura coupling, offering an alternative to traditionally used organic halides or triflates. The oxidative addition of nitro derivatives onto palladium catalysts remains challenging and often requires harsh conditions and expensive catalytic systems. Nevertheless, we recently demonstrated that nitro-perylenediimide derivatives can effectively engage in various cross-couplings with unsophisticated Pd(PPh3)4 as a catalytic system. The mechanistic study of the oxidative addition step for this particular class of nitro derivatives revealed an unprecedented single electron transfer event, which is supported by a comprehensive range of analyses including NMR, X-ray diffraction, HRMS and EPR experiments, complemented with cyclic voltammetry, and theoretical calculations.
Flavins and their alloxazine isomers are key chemical scaffolds for bioinspired electron transfer strategies. Their properties can be fine-tuned by functional groups, which must be introduced at an early stage of the synthesis as their aromatic ring is inert towards post-functionalization. We show that the introduction of a remote metal-binding redox site on alloxazine and flavin activates their aromatic ring towards direct C−H functionalization. Mechanistic studies are consistent with a synthetic sequence involving ground-state single electron transfer (SET) with an electrophilic source followed by radical-radical coupling. This unprecedented reactivity opens new opportunities in molecular editing of flavins by direct aromatic post-functionalization and the utility of the method is demonstrated with the site-selective C6 functionalization of alloxazine and flavin with a CF 3 group, Br or Cl, that can be further elaborated into OH and aryl for chemical diversification.
Models can codify our understanding of chemical reactivity and serve a useful purpose in the development of new synthetic processes via, for example, evaluating hypothetical reaction conditions or in silico substrate tolerance. Perhaps the most determining factor is the composition of the training data and whether it is sufficient to train a model that can make accurate predictions over the full domain of interest. Here, we discuss the design of reaction datasets in ways that are conducive to data-driven modeling, emphasizing the idea that training set diversity and model generalizability rely on the choice of molecular or reaction representation. We additionally discuss the experimental constraints associated with generating common types of chemistry datasets and how these considerations should influence dataset design and model building.
Potential inversion refers to the situation where a protein cofactor or a synthetic molecule can be oxidized or reduced twice in a cooperative manner; that is, the second electron transfer is easier than the first. This property is very important regarding the catalytic mechanism of enzymes that bifurcate electrons and the properties of bidirectional redox molecular catalysts that function in either direction of the reaction with no overpotential. Cyclic voltammetry is the most common technique for characterizing the thermodynamics and kinetics of electron transfer to or from these molecules. However, a gap in the literature is the absence of analytical predictions to help interpret the values of the voltammetric peak potentials when potential inversion occurs; the cyclic voltammograms are therefore often analyzed by simulating the data, with no discussion of the possibility of overfitting and often no estimation of the error on the determined parameters. Here we formulate the theory for the voltammetry of freely diffusing or surface-confined two-electron redox species in the experimentally relevant irreversible limit where the peak separation depends on the scan rate. We explain why the model is intrinsically underdetermined, and we illustrate this conclusion by analysis of the voltammetry of a nickel complex with redox-active iminosemiquinone ligands. Being able to characterize the thermodynamics of two-electron electron-transfer reactions will be crucial for designing more efficient catalysts.
Synthetic yield prediction using machine learning is intensively studied. While previous work focused on an ideal use case, High-Throughput Experiment datasets, predicting yields using literature data remains elusive. We built a large literature- based dataset of more than a thousand reactions, focusing on the activation of carbon-oxygen bonds of phenol derivatives under nickel catalysis. Detailed reaction conditions and associated yields were manually curated and stored in an open- access database. We assessed the performances of state-of-the-art machine learning models on this dataset, and explored their ability to realize predictions on novel publications, coupling partners and substrates. Our work shows that on well- designed yield prediction tasks, machine learning can have practical applications, and provides a unique public database for further improvements of these methods adapted to literature chemical data.
Mechanisms combining organic radicals and metallic intermediates hold strong potential in homogeneous catalysis. Such activation modes require careful optimization of two interconnected processes: one for the generation of radicals and one for their productive integration towards the final product. We report that a bioinspired polymetallic nickel complex can combine ligand- and metal-centered reactivities to perform fast hydrosilylation of alkenes under mild conditions through an unusual dual radical- and metal-based mechanism. This earth-abundant polymetallic complex incorporating a catechol-alloxazine motif as redox-active ligand operates at low catalyst loading (0.25 mol%) and generates silyl radicals and a nickel-hydride intermediate through a hydrogen atom transfer (HAT) step. Evidence of an isomerization sequence enabling terminal hydrosilylation of internal alkenes points towards the involvement of the nickel-hydride species in chain walking. This single catalyst promotes a hybrid pathway by combining synergistically ligand and metal participation in both inner- and outer- sphere processes.
Synthetic yield prediction using machine learning is intensively studied. Previous work focused on two categories of datasets: High-Throughput Experimentation data, as an ideal case study and datasets extracted from proprietary databases, which are known to have a strong reporting bias towards high yields. However, predicting yields using published reaction data remains elusive. To fill the gap, we built a dataset on nickel-catalyzed cross-couplings extracted from organic reaction publications, including scope and optimization information. We demonstrate the importance of including optimization data as a source of failed experiments and emphasize how publication constraints shape the exploration of the chemical space by the synthetic community. While machine learning models still fail to perform out-of-sample predictions, this work shows that adding chemical knowledge enables fair predictions in a low-data regime. Eventually, we hope that this unique public database will foster further improvements of machine learning methods for reaction yield prediction in a more realistic context.
An entry from the Cambridge Structural Database, the world’s repository for small molecule crystal structures. The entry contains experimental data from a crystal diffraction study. The deposited dataset for this entry is freely available from the CCDC and typically includes 3D coordinates, cell parameters, space group, experimental conditions and quality measures.
Synthetic yield prediction using machine learning is intensively studied. Previous work has focused on two categories of data sets: high-throughput experimentation data, as an ideal case study, and data sets extracted from proprietary databases, which are known to have a strong reporting bias toward high yields. However, predicting yields using published reaction data remains elusive. To fill the gap, we built a data set on nickel-catalyzed cross-couplings extracted from organic reaction publications, including scope and optimization information. We demonstrate the importance of including optimization data as a source of failed experiments and emphasize how publication constraints shape the exploration of the chemical space by the synthetic community. While machine learning models still fail to perform out-of-sample predictions, this work shows that adding chemical knowledge enables fair predictions in a low-data regime. Eventually, we hope that this unique public database will foster further improvements of machine learning methods for reaction yield prediction in a more realistic context.
We report herein an unprecedented combination of light and P(III)/P(V) redox cycling for the efficient deoxygenation of aromatic amine N-oxides. Moreover, we discovered that a large variety of aliphatic amine N-oxides can easily be deoxygenated by using only phenylsilane. These practically simple approaches proceed well under metal-free conditions, tolerate many functionalities and are highly chemoselective. Combined experimental and computational studies enabled a deep understanding of factors controlling the reactivity of both aromatic and aliphatic amine N-oxides.
This contains the dataset of (Cot)2LnK2(thf)4_Yb98%Tm2%
A new class of redox-active ligands merging catechol and alloxazine structures is reported. A trimetallic triangular complex is formed upon complexation to nickel.