The macrodomain of severe acute respiratory syndrome coronavirus 2 nonstructural protein 3 is required for viral pathogenesis and is an emerging antiviral target. We previously performed an x-ray crystallography–based fragment screen and found submicromolar inhibitors by fragment linking. However, these compounds had poor membrane permeability and liabilities that complicated optimization. Here, we developed a shape-based virtual screening pipeline—FrankenROCS. We screened the Enamine high-throughput collection of 2.1 million compounds, selecting 39 compounds for testing, with the most potent binding with a 130 μM median inhibitory concentration (IC 50 ). We then paired FrankenROCS with an active learning algorithm (Thompson sampling) to efficiently search the Enamine REAL database of 22 billion molecules, testing 32 compounds with the most potent binding with a 220 μM IC 50 . Further optimization led to analogs with IC 50 values better than 10 μM. This lead series has improved membrane permeability and is poised for optimization. FrankenROCS is a scalable method for fragment linking to exploit synthesis-on-demand libraries.
We present a simple method for assigning accurate confidence levels to molecular property predictions from regression models. These confidence levels are easy to interpret and useful for making decisions in drug discovery programs. We demonstrate their performance using time-split validation with assay data from the Relay Therapeutics internal database.
The macrodomain contained in the SARS-CoV-2 non-structural protein 3 (NSP3) is required for viral pathogenesis and lethality. Inhibitors that block the macrodomain could be a new therapeutic strategy for viral suppression. We previously performed a large-scale X-ray crystallography-based fragment screen and discovered a sub-micromolar inhibitor by fragment linking. However, this carboxylic acid-containing lead had poor membrane permeability and other liabilities that made optimization difficult. Here, we developed a shape-based virtual screening pipeline - FrankenROCS - to identify new macrodomain inhibitors using fragment X-ray crystal structures. We used FrankenROCS to exhaustively screen the Enamine high-throughput screening (HTS) collection of 2.1 million compounds and selected 39 compounds for testing, with the most potent compound having an IC50 value equal to 130 μM. We then paired FrankenROCS with an active learning algorithm (Thompson sampling) to efficiently search the Enamine REAL database of 22 billion molecules, testing 32 compounds with the most potent having an IC50 equal to 220 μM. Further optimization led to analogs with IC50 values better than 10 μM, with X-ray crystal structures revealing diverse binding modes despite conserved chemical features. These analogs represent a new lead series with improved membrane permeability that is poised for optimization. In addition, the collection of 137 X-ray crystal structures with associated binding data will serve as a resource for the development of structure-based drug discovery methods. FrankenROCS may be a scalable method for fragment linking to exploit ever-growing synthesis-on-demand libraries.
Protein class-focused drug discovery has a long and successful history in pharmaceutical research, yet most members of druggable protein families remain unliganded, often for practical reasons. Here we combined experiment and computation to enable discovery of ligands for WD40 repeat (WDR) proteins, one of the largest human protein families. This resource includes expression clones, purification protocols, and a comprehensive assessment of the druggability for hundreds of WDR proteins. We solved 21 high resolution crystal structures, and have made available a suite of biophysical, biochemical, and cellular assays to facilitate the discovery and characterization of small molecule ligands. To this end, we use the resource in a hit-finding pilot involving DNA-encoded library (DEL) selection followed by machine learning (ML). This led to the discovery of first-in-class, drug-like ligands for 9 of 20 targets. This result demonstrates the broad ligandability of WDRs. This extensive resource of reagents and knowledge will enable further discovery of chemical tools and potential therapeutics for this important class of proteins.### Competing Interest StatementBLS, BG, JSD, JZ, JWC, MvR, PR, SK, and TK are employees and shareholders of Relay Therapeutics.* BLI : Biolayer Interferometry DEL : DNA Encoded Library DLID : drug-like density DSF : Differential Scanning Fluorimetry FP : Fluorescence Polarization GCNN : Graph Convolutional Neural Network HDX : Hydrogen-Deuterium eXchange HTS : High-Throughput Screening ML : Machine Learning NanoBRET : NanoLuciferase Bioluminescence Resonance Energy Transfer PPI : Protein-Protein Interaction PROTAC : Proteolysis Targeting Chimera SPR : Surface Plasmon Resonance Tm : Protein melting temperature WDR : Tryptophan-Aspartate Repeat
Over the last five years, virtual screening of ultralarge synthesis on-demand libraries has emerged as a powerful tool for hit identification in drug discovery programs. As these libraries have grown to tens of billions of molecules, we have reached a point where it is no longer cost-effective to screen every molecule virtually. To address these challenges, several groups have developed heuristic search methods to rapidly identify the best molecules on a virtual screen. This article describes the application of Thompson sampling (TS), an active learning approach that streamlines the virtual screening of large combinatorial libraries by performing a probabilistic search in the reagent space, thereby never requiring the full enumeration of the library. TS is a general technique that can be applied to various virtual screening modalities, including 2D and 3D similarity search, docking, and application of machine-learning models. In an illustrative example, we show that TS can identify more than half of the top 100 molecules from a docking-based virtual screen of 335 million molecules by evaluating 1% of the data set.
Magnetic field changes as earthquake precursors have been the subject of numerous studies and some controversy. Infrequent large earthquakes and sparse magnetometer coverage along fault zones complicate statistical analysis. We present an analysis of ground‐based magnetic time‐series measurements before 19 earthquakes ≥M4.5 in California drawing from over 330,000 site‐days of measurement spanning a decade. To perform a fair existential test for electromagnetic antecedents we applied a pre‐specified statistical analysis with two key ideas. First, we combine signals from nearby (≤40 km) sites via spectral cross‐power, and then look for large spikes in frequency domain (0.016–25 Hz). The former is only possible with a dense set of sites running over a long period of time. In this statistical case‐control study we used the machine learning concept of rigorously separated train and test sets of earthquakes which were generated via a rule‐based query of the USGS earthquake catalog. Before each declustered earthquake, we constructed one period 24–72 hr before (the “precursor” or “p‐period”) and a series of seven equally‐sized preceding periods (“quiescent” or “q‐periods”). We distilled the data in each period to a frequency‐dependent feature—the 98th percentile of spectral cross power. We trained a model based on Linear Discriminant Analysis and applied the discriminator to the test set revealing a modest effect in the days leading up to an earthquake. While the observed effect size is not directly useful for earthquake prediction (long a scientific goal), it suggests a relationship which should be further investigated for a physical link.
Systematic development of accurate density functionals has been a decades-long challenge for scientists. Despite emerging applications of machine learning (ML) in approximating functionals, the resulting ML functionals usually contain more than tens of thousands of parameters, leading to a huge gap in the formulation with the conventional human-designed symbolic functionals. We propose a new framework, Symbolic Functional Evolutionary Search (SyFES), that automatically constructs accurate functionals in the symbolic form, which is more explainable to humans, cheaper to evaluate, and easier to integrate to existing codes than other ML functionals. We first show that, without prior knowledge, SyFES reconstructed a known functional from scratch. We then demonstrate that evolving from an existing functional ωB97M-V, SyFES found a new functional, GAS22 (Google Accelerated Science 22), that performs better for most of the molecular types in the test set of Main Group Chemistry Database (MGCDB84). Our framework opens a new direction in leveraging computing power for the systematic development of symbolic density functionals.
One application area of computational methods in drug discovery is the automated design of small molecules. Despite the large number of publications describing methods and their application in both retrospective and prospective studies, there is a lack of agreement on terminology and key attributes to distinguish these various systems. We introduce Automated Chemical Design (ACD) Levels to clearly define the level of autonomy along the axes of ideation and decision making. To fully illustrate this framework, we provide literature exemplars and place some notable methods and applications into the levels. The ACD framework provides a common language for describing automated small molecule design systems and enables medicinal chemists to better understand and evaluate such systems.
Computational approaches in drug discovery and development hold great promise, with artificial intelligence methods undergoing widespread contemporary use, but the experimental validation of these new approaches is frequently inadequate. We are initiating Critical Assessment of Computational Hit-finding Experiments (CACHE) as a public benchmarking project that aims to accelerate the development of small molecule hit-finding algorithms by competitive assessment. Compounds will be identified by participants using a wide range of computational methods for dozens of protein targets selected for different types of prediction scenarios, as well as for their potential biological or pharmaceutical relevance. Community-generated predictions will be tested centrally and rigorously in an experimental hub(s), and all data, including the chemical structures of experimentally tested compounds, will be made publicly available without restrictions. The ability of a range of computational approaches to find novel compounds will be evaluated, compared, and published. The overarching goal of CACHE is to accelerate the development of computational chemistry methods by providing rapid and unbiased feedback to those developing methods, with an ancillary and valuable benefit of identifying new compound-protein binding pairs for biologically interesting targets. The initiative builds on the power of crowd sourcing and expands the open science paradigm for drug discovery.
Symbolic techniques based on Satisfiability Modulo Theory (SMT) solvers have been proposed for analyzing and verifying neural network properties, but their usage has been fairly limited owing to their poor scalability with larger networks. In this work, we propose a technique for combining gradient-based methods with symbolic techniques to scale such analyses and demonstrate its application for model explanation. In particular, we apply this technique to identify minimal regions in an input that are most relevant for a neural network's prediction. Our approach uses gradient information (based on Integrated Gradients) to focus on a subset of neurons in the first layer, which allows our technique to scale to large networks. The corresponding SMT constraints encode the minimal input mask discovery problem such that after masking the input, the activations of the selected neurons are still above a threshold. After solving for the minimal masks, our approach scores the mask regions to generate a relative ordering of the features within the mask. This produces a saliency map which explains "where a model is looking" when making a prediction. We evaluate our technique on three datasets - MNIST, ImageNet, and Beer Reviews, and demonstrate both quantitatively and qualitatively that the regions generated by our approach are sparser and achieve higher saliency scores compared to the gradient-based methods alone. Code and examples are at - https://github.com/google-research/google-research/tree/master/smug_saliency
The quest to identify materials with tailored properties is increasingly expanding into high-order composition spaces, with a corresponding combinatorial explosion in the number of candidate materials. A key challenge is to discover regions in composition space where materials have novel properties. Traditional predictive models for material properties are not accurate enough to guide the search. Herein, we use high-throughput measurements of optical properties to identify novel regions in three-cation metal oxide composition spaces by identifying compositions whose optical trends cannot be explained by simple phase mixtures. We screen 376,752 distinct compositions from 108 three-cation oxide systems based on the cation elements Mg, Fe, Co, Ni, Cu, Y, In, Sn, Ce, and Ta. Data models for candidate phase diagrams and three-cation compositions with emergent optical properties guide the discovery of materials with complex phase-dependent properties, as demonstrated by the discovery of a Co-Ta-Sn substitutional alloy oxide with tunable transparency, catalytic activity, and stability in strong acid electrolytes. These results required close coupling of data validation to experiment design to generate a reliable end-to-end high-throughput workflow for accelerating scientific discovery.
Including prior knowledge is important for effective machine learning models in physics and is usually achieved by explicitly adding loss terms or constraints on model architectures. Prior knowledge embedded in the physics computation itself rarely draws attention. We show that solving the Kohn-Sham equations when training neural networks for the exchange-correlation functional provides an implicit regularization that greatly improves generalization. Two separations suffice for learning the entire one-dimensional H_{2} dissociation curve within chemical accuracy, including the strongly correlated region. Our models also generalize to unseen types of molecules and overcome self-interaction error.
Autonomous experimentation (AE) accelerates research by combining automation and machine learning to perform experiments intelligently and rapidly in a sequential fashion. While AE systems are most needed to study properties that cannot be predicted analytically or computationally, even imperfect predictions can in principle be useful. Here, we investigate whether imperfect data from simulation can accelerate AE using a case study on the mechanics of additively manufactured structures. Initially, we study resilience, a property that is well-predicted by finite element analysis (FEA), and find that FEA can be used to build a Bayesian prior and experimental data can be integrated using discrepancy modeling to reduce the number of needed experiments ten-fold. Next, we study toughness, a property not well-predicted by FEA and find that FEA can still improve learning by transforming experimental data and guiding experiment selection. These results highlight multiple ways that simulation can improve AE through transfer learning.
DNA-encoded small molecule libraries (DELs) have enabled discovery of novel inhibitors for many distinct protein targets of therapeutic value through screening of libraries with up to billions of unique small molecules. We demonstrate a new approach applying machine learning to DEL selection data by identifying active molecules from a large commercial collection and a virtual library of easily synthesizable compounds. We train models using only DEL selection data and apply automated or automatable filters with chemist review restricted to the removal of molecules with potential for instability or reactivity. We validate this approach with a large prospective study (nearly 2000 compounds tested) across three diverse protein targets: sEH (a hydrolase), ER{\alpha} (a nuclear receptor), and c-KIT (a kinase). The approach is effective, with an overall hit rate of {\sim}30% at 30 {\textmu}M and discovery of potent compounds (IC50 <10 nM) for every target. The model makes useful predictions even for molecules dissimilar to the original DEL and the compounds identified are diverse, predominantly drug-like, and different from known ligands. Collectively, the quality and quantity of DEL selection data; the power of modern machine learning methods; and access to large, inexpensive, commercially-available libraries creates a powerful new approach for hit finding.
The quantum approximate optimization algorithm (QAOA) is a standard method for combinatorial optimization with a gate-based quantum computer. The QAOA consists of a particular ansatz for the quantum circuit architecture, together with a prescription for choosing the variational parameters of the circuit. We propose modifications to both. First, we define the Gibbs objective function and show that it is superior to the energy expectation value for use as an objective function in tuning the variational parameters. Second, we describe an ansatz architecture search (AAS) algorithm for searching the discrete space of quantum circuit architectures near the QAOA to find a better ansatz. Applying these modifications for a complete graph Ising model results in a 244.7% median relative improvement in the probability of finding a low-energy state while using 33.3% fewer two-qubit gates. For Ising models on a 2d grid we similarly find 44.4% median improvement in the probability with a 20.8% reduction in the number of two-qubit gates. This opens a new research field of quantum circuit architecture design for quantum optimization algorithms.
Symbolic regression is a type of discrete optimization problem that involves searching expressions that fit given data points. In many cases, other mathematical constraints about the unknown expression not only provide more information beyond just values at some inputs, but also effectively constrain the search space. We identify the asymptotic constraints of leading polynomial powers as the function approaches zero and infinity as useful constraints and create a system to use them for symbolic regression. The first part of the system is a conditional production rule generating neural network which preferentially generates production rules to construct expressions with the desired leading powers, producing novel expressions outside the training domain. The second part, which we call Neural-Guided Monte Carlo Tree Search, uses the network during a search to find an expression that conforms to a set of data points and desired leading powers. Lastly, we provide an extensive experimental validation on thousands of target expressions showing the efficacy of our system compared to exiting methods for finding unknown functions outside of the training set.
While additive manufacturing (AM) has facilitated the production of complex structures, it has also highlighted the immense challenge inherent in identifying the optimum AM structure for a given application. Numerical methods are important tools for optimization, but experiment remains the gold standard for studying nonlinear, but critical, mechanical properties such as toughness. To address the vastness of AM design space and the need for experiment, we develop a Bayesian experimental autonomous researcher (BEAR) that combines Bayesian optimization and high-throughput automated experimentation. In addition to rapidly performing experiments, the BEAR leverages iterative experimentation by selecting experiments based on all available results. Using the BEAR, we explore the toughness of a parametric family of structures and observe an almost 60-fold reduction in the number of experiments needed to identify high-performing structures relative to a grid-based search. These results show the value of machine learning in experimental fields where data are sparse.
The aim of this paper is to build a computer based clinical decision support tool using a semi-supervised framework, the Fisher Information Network (FIN), for visualization of a set of mammographic images. The FIN organizes the images into a similarity network from which, for any new image, reference images that are closely related can be identified. This enables clinicians to review not just the reference images but also ancillary information e.g. about response to therapy. The Fisher information metric defines a Riemannian space where distances reflect similarity with respect to a given probability distribution. This metric is informed about generative properties of data, and hence assesses the importance of directions in space of parameters. It automatically performs feature relevance detection. This approach focusses on the interpretability of the model from the standpoint of the clinical user. Model predictions were validated using the prevalence of classes in each of the clusters identified by the FIN.
Tristan Bereau, Kathy Breen, Xavier Brumwell, Bing Brunton, Steven Brunton, Jared Callaham, Onur Çaylak, Maghesree Chakraborty, Kathleen Champion, Nicholas Charron, Yu-Chia Chen, Cecilia Clementi, Gianni De Fabritiis, Alex Goeßmann, Rhys Goodall, Francois Gygi, Richard G. Hennig, Moritz Hoffmann, Brooke Husic, Lukas Kades, Tom Kennedy, Stefan Klus, Jonas Köhler, Dominik Lemm, Andreas Mardt, Marina Meila, Kim A. Nicoli, Frank Noé, Luca Pasquali, Patrick Riley, Adam Rupe, Matthias Rupp, Lars Ruthotto, Anastasiya Salova, Kirill Shmilovich, Jordan Snyder, Matthew Spellings, Marc Stieffenhofer, Luca Venturi, Jiang Wang, Andrew D. White, Stephen R. Xie
scientists from myriad fields rush to perform algorithmic analyses, Google’s Patrick Riley calls for clear standards in research and reporting.