One-dimensional NMR spectroscopy is one of the most widely used techniques for the characterization of organic compounds and natural products. For molecules with up to 36 non-hydrogen atoms, the number of possible structures has been estimated to range from 1020-1060. The task of determining the structure (formula and connectivity) of a molecule of this size using only its one-dimensional 1H and/or 13C NMR spectrum, i.e., de novo structure generation, thus appears completely intractable. Here, we show how it is possible to achieve this task for systems with up to 40 non-hydrogen atoms across the full elemental coverage typically encountered in organic chemistry (C, N, O, H, P, S, Si, B, and the halogens) using a deep learning framework, thus covering a vast portion of the drug-like chemical space. Leveraging insights from natural language processing, we show that our transformer-based architecture predicts the correct molecule with 60.4% accuracy within the first 15 predictions using only the 1H and 13C NMR spectra, thus overcoming the combinatorial growth of the chemical space while also being extensible to experimental data via fine-tuning.
A biaryl peroxide natural product is reassigned to an industrial antioxidant following inspection of the original data, calculation of 13C NMR chemical shifts, and comparison to an authentic sample.
The proposed structure for myrrhain A, containing an unlikely 1,2-di-t-butylbenzofuran substitution pattern, was reassigned as a known plasticizer by direct spectroscopic comparison and Computer Assisted Structure Elucidation (CASE). Concerningly, plasticizers and their precursors, degradation and/or oxidation products, arising from environmental or laboratory contamination, continue to be misconstrued as natural products.
The structure of the doubly anti-Bredt tropone natural product crotonguaienone G has been shown to be incorrect, and the isolated compound is shown to be identical to the known natural product pernambucone. The misassignment can be traced to erroneous 2D NMR data.
By a combination of computational methods and comparison of spectroscopic data, the benzooxonin structure proposed for setosol is shown to be incorrect. The correct structure is that of a known biaryl ether natural product.
Biodiscovery efforts in Indonesia have aimed to explore the understudied chemical diversity of its rich lichen flora, seeking to find new products endowed with significant biological properties. The chemical screening of a Teloschistes flavicans extract led to selection of this species for further investigation. LC/MS and 1H NMR-based dereplication pinpointed six chlorodepsidones from the thallus of a sample of this lichen. This led to the streamlined isolation and the subsequent structure elucidation of the three new compounds norflavicansone 1, flavicansone 2, and isocaloploicin 3, along with the known chlorodepsidones 4-6, stictic acid 7, aurantiamide acetate 8, and parietin 9. The challenging structure elucidation of these proton-deficient metabolites benefited from a state-of-the-art workflow involving a synergistic combination of Computer-Assisted Structure Elucidation (CASE) and Density Functional Theory (DFT) calculations of the top-ranked candidates. This investigation also led to the revision of flavicansone's structure, previously described from this species. The three new molecules that are being reported here are remarkable in that they represent hybrid depsidones in which one of the aromatic rings is derived from orsellinic acid and the other is derived from β-orcinol, a rare structural feature for lichen depsidones.
the proposed structure of arneroma B has been revised from a cyclopentadienone to a 2,4-disubstituted furan. The reassignment has been confirmed by total synthesis of the revised structure.
The proposed structure for the natural product penicitone, which contained a chemically improbable acid chloride functional group, was reassigned to a more probable structure using a combination of chemical knowledge, computer-assisted structure elucidation, and DFT methods.
Using computational methods and chemical intuition, the proposed structure of janthinolide A is shown to be incorrect. It is further shown that the material described as janthinolide A is highly likely to be janthinolide C.
Unusual polyenols that defied chemical principles were reassigned as the nucleosides, adenosine and uridine, using a combination of chemical intuition underpinned by Computer Assisted Structure Elucidation (CASE) and DFT methods.
The UHPLC–HRMS analysis of Cortinarius ominosus basidiomata extract revealed that this mushroom accumulated elevated yields of an unreported specialized metabolite. The molecular formula of this unknown compound, C17H10O8, indicated that a challenging structure elucidation lay ahead, owing to its critically low H/C atom ratio. The structure of this new isolate, namely ominoxanthone (1), could not be solved from the interpretation of the usual set of 1D/2D NMR data that conveyed too limited information to afford a single, unambiguous structure. To remedy this, a Computer-Assisted Structure Elucidation (CASE) workflow was used to rank the different possible structure candidates consistent with our scarce spectroscopic data. DFT-based chemical shift calculations on a limited set of top-ranked structures further ascertained the determined structure for ominoxanthone. Although the determined scaffold of ominoxanthone is unprecedented as a natural product, a plausible biosynthetic scenario involving a precursor known from cortinariaceous sources and classical biogenetic reactions could be proposed.
Natural products remain one of the major sources of coveted, biologically active compounds. Each isolated compound undergoes biological testing, and its structure is usually established using a set of spectroscopic techniques (NMR, MS, UV-IR, ECD, VCD, etc.). However, the number of erroneously determined structures remains noticeable. Structure revisions are very costly, as they usually require extensive use of spectroscopic data, computational chemistry, and total synthesis. The cost is particularly high when a biologically active compound is resynthesized and the product is inactive because its structure is wrong and remains unknown. In this paper, we propose using Computer-Assisted Structure Elucidation (CASE) and Density Functional Theory (DFT) methods as tools for preventive verification of the originally proposed structure, and elucidation of the correct structure if the original structure is deemed to be incorrect. We examined twelve real cases in which structure revisions of natural products were performed using total synthesis, and we showed that in each of these cases, time-consuming total synthesis could have been avoided if CASE and DFT had been applied. In all described cases, the correct structures were established within minutes of using the originally published NMR and MS data, which were sometimes incomplete or had typos.
Natural products continue to be reported at an astonishing rate from a wide range of multidisciplinary research activities in the pursuit of understanding the chemistry of biodiversity. However, the elucidation of chemical structure in the modern era is heavily reliant on the analysis and interpretation of multiple spectroscopic outputs, and in most cases this activity is by no means trivial. Structural errors continue to be described given the inherent complexity of natural products. Computer-Assisted Structure Elucidation (CASE) continues to provide improved resolving power in this regard, but for enhanced accuracy quantum chemical spectrum prediction methodology is paramount. Reported herein are a range of counterfactual natural products, identified through chemical principal screening, which have been reassigned using a combination of chemical intuition, chemical synthesis, CASE and DU8+ spectrum prediction.
The first methods associated with the Computer-Assisted Structure Elucidation (CASE) of small molecules were published over fifty years ago when spectroscopy and computer science were both in their infancy. The incredible leaps in both areas of technology could not have been envisaged at that time, but both have enabled CASE expert systems to achieve performance levels that in their present state can outperform many scientists in terms of speed to solution. The computer-assisted analysis of enormous matrices of data exemplified 1D and 2D high-resolution NMR spectroscopy datasets can easily solve what just a few years ago would have been deemed to be complex structures. While not a panacea, the application of such tools can provide support to even the most skilled spectroscopist. By this point the structures of a great number of molecular skeletons, including hundreds of complex natural products, have been elucidated using such programs. At this juncture, the expert system ACD/Structure Elucidator is likely the most advanced CASE system available and, being a commercial software product, is installed and used in many organizations. This article will provide an overview of the research and development required to pursue the lofty goals set almost two decades ago to facilitate highly automated approaches to solving complex structures from analytical spectroscopy data, using NMR as the primary data-type.
Computer-assisted structure elucidation (CASE) is the class of expert systems that derives molecular structures primarily from one-dimensional and two-dimensional nuclear magnetic resonance data. Contemporary CASE systems, including Advanced Chemistry Development/Structure Elucidator (ACD/SE), consider cross-peaks in heteronuclear multiple bond coherence (HMBC) and correlation spectroscopy (COSY) spectra as two- or three-bond correlations by default. However, four and more bond correlations (nonstandard correlations [NSCs]) could be present in these spectra too. The indiscriminate addition of NSCs to the CASE computations is prohibitively expensive. To address this problem, the ACD/SE program performs a logical analysis of observed correlations and determines the minimum number of NSCs. Guided by this information, a more efficient fuzzy structure generation (FSG) algorithm is subsequently applied. Until now, the FSG algorithm was utilized without any verification of the reliability of found NSCs. Here, we report a verification method for NSCs based on the relationship between NSCs and J-couplings computed with high accuracy density functional theory (DFT) methods. We used the example of strychnine to show that 41 (32%) of 8-Hz HMBC cross-peaks were NSCs and were consistent with (4-6)J(CH) couplings greater than 0.3 Hz. This cutoff value was largely confirmed by the analysis of NSCs in 11 real-world natural products elucidated by ACD/SE. Additionally, utilizing the example of the CASE study of cleospinol A, we showed that the DFT-computed J-couplings of NSCs can distinctively differentiate the correct structure among six proposed isomers. The proposed approach of NSC verification should further improve the robustness of CASE analysis and can help reveal potential problems with reported experimental data.
The first efforts for the development of methods for Computer-Assisted Structure Elucidation (CASE) were published more than 50 years ago. CASE expert systems based on one-dimensional (1D) and two-dimensional (2D) Nuclear Magnetic Resonance (NMR) data have matured considerably by now. The structures of a great number of complex natural products have been elucidated and/or revised using such programs. In this article, we discuss the most likely directions in which CASE will evolve. We act on the premise that a synergistic interaction exists between CASE, new NMR experiments, and methods of computational chemistry, which are continuously being improved. The new developments in NMR experiments (long-range correlation experiments, pure-shift methods, coupling constants measurement and prediction, residual dipolar couplings [RDCs]), and residual chemical shift anisotropies [RCSAs], evolution of density functional theory (DFT), and machine learning algorithms will have an influence on CASE systems and vice versa. This is true also for new techniques for chemical analysis (Atomic Force Microscopy [AFM], "crystalline sponge" X-ray analysis, and micro-Electron Diffraction [micro-ED]), which will be used in combination with expert systems. We foresee that CASE will be utilized widely and become a routine tool for NMR spectroscopists and analysts in academic and industrial laboratories. We believe that the "golden age" of CASE is still in the future.
CASE (computer-assisted structure elucidation) first appeared in the late 1960s but really gained traction in the 1990s as more information-rich 2D NMR experiments were developed. In this article, we discuss the strategies of CASE for small organic molecules in solution. Cognitive grounds and the principal CASE flow-diagram, as well as the main obstacles impeding structure elucidation (presence of 'nonstandard' correlations and ambiguity of 2D NMR data, deficit of hydrogen, etc.) are discussed and methods to overcome these challenges are suggested. The methods are illustrated by examples of solving challenging problems. It has been shown that CASE can be used to avoid pitfalls during structure elucidation and to determine the most efficient combinations of 2D NMR experiments. Methodology of DR-based NMR spectrum prediction in synergistic combination with CASE is explained. The last advances in 3D structure elucidation and stereochemistry determination using RDCs and RCSAs are considered. In conclusions, perspectives of CASE development are discussed.
Even though NMR has found countless applications in the field of small molecule characterization, there is no standard file format available for the NMR data relevant to structure characterization of small molecules. A new format is therefore introduced to associate the NMR parameters extracted from 1D and 2D spectra of organic compounds to the proposed chemical structure. These NMR parameters, which we shall call NMReDATA (for nuclear magnetic resonance extracted data ), include chemical shift values, signal integrals, intensities, multiplicities, scalar coupling constants, lists of 2D correlations, relaxation times, and diffusion rates. The file format is an extension of the existing Structure Data Format, which is compatible with the commonly used MOL format. The association of an NMReDATA file with the raw and spectral data from which it originates constitutes an NMR record. This format is easily readable by humans and computers and provides a simple and efficient way for disseminating results of structural chemistry investigations, allowing automatic verification of published results, and for assisting the constitution of highly needed open‐source structural databases.
Computer‐assisted structure elucidation (CASE) is composed of two steps: (a) generation of all possible structural isomers for a given molecular formula and 2D NMR data (COSY, HSQC, and HMBC) and (b) selection of the correct isomer based on empirical chemical shift predictions. This method has been very successful in solving structural problems of small organic molecules and natural products. However, CASE applications are generally limited to structural isomer problems and can sometimes be inconclusive due to insufficient accuracy of empirical shift predictions. Here, we report a synergistic combination of a CASE algorithm and density functional theory calculations that broadens the range of amenable structural problems to encompass proton‐deficient molecules, molecules with heavy elements (e.g., halogens), conformationally flexible molecules, and configurational isomers.
Christoph Steinbeck合作论文数EMBL Outstation - Hinxton,
European Bioinformatics Institute,
Wellcome Trust Genome Campus3