Abstract We introduce the Serna Bio GenAI platform, a generative chemistry and multiparametric optimization platform for the design of RNA-targeting small molecules. Targeting RNA with small molecules has proven historically challenging but offers notable potential upsides, including access to unique mechanisms of action and the ability to target otherwise untargetable genes. We consider a major challenge here to be designing chemistry specific to RNA-targeting. Molecular design is a valuable application of AI in drug discovery, but many publicly available models use training data focused on protein-targeting - the modality best historically explored in drug discovery. We showcase the difference and value in building a specifically RNA-targeting platform, comparing its performance to state-of-the-art public chemical generators and experimentally validating its chemical designs in comparison to chemistry designed by a human expert.
The possibility of using RNA-targeting small molecules to treat diseases is gaining traction as the next frontier of drug discovery and development. The chemical characteristics of small molecules that bind to RNA are still relatively poorly understood, particularly in comparison to protein-targeting small molecules. To fill this gap, we have generated an unprecedented amount of RNA-small molecule binding data, and used it to derive physicochemical rules of thumb that could be used to define areas of chemical space enriched for RNA binders - the Small molecules Targeting RNA (STaR) rules of thumb. These rules have been applied to publicly available RNA-small molecule datasets and found to be largely generalizable. Furthermore, a number of patented RNA-targeting compounds and FDA-approved compounds also pass these rules, as well as key RNA binding approved drug case studies including Risdiplam. We anticipate this work will significantly accelerate the exploration of the RNA-targeted chemical space, towards unlocking RNA’s potential as a small molecule drug target. ![Figure][1] ### Competing Interest Statement The authors declare the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: T.E.H.A., J.L.M., M.B., C.J.B. and R.T.K are current or former employees of Serna Bio and may hold stock or other financial interests in Serna Bio. * ASA : Accessible surface area ASMS : Affinity selection-mass spectrometry CC-DDR : cell cycle and DNA damage repair cLogP : calculated LogP DRTL : Duke RNA targeted library PBF over MW : Plane of best fit over molecular weight R-BIND : RNA-targeted bioactive ligand database Relative PSA : Relative polar surface area ROBIN : Repository of binders to nucleic acids STaR : Small molecule targeting RNA TPSA : Topological polar surface area [1]: pending:yes
AbstractSmall molecule targeting of RNA has emerged as a new frontier in medicinal chemistry, but compared to the protein targeting literature our understanding of chemical matter that binds to RNA is limited. In this study, we reportedRepositoryOfBInders toNucleic acids (ROBIN), a new library of nucleic acid binders identified by small molecule microarray (SMM) screening. The complete results of 36 individual nucleic acid SMM screens against a library of 24 572 small molecules were reported (including a total of 1 627 072 interactions assayed). A set of 2 003 RNA‐binding small molecules was identified, representing the largest fully public, experimentally derived library of its kind to date. Machine learning was used to develop highly predictive and interpretable models to characterize RNA‐binding molecules. This work demonstrates that machine learning algorithms applied to experimentally derived sets of RNA binders are a powerful method to inform RNA‐targeted chemical space.
Key Clinical MessageCerebritis may present with encephalopathy, seizures, localizing neurologic deficits, or death. Topical betadine has a high iodine content and application to open scalp wounds may cause sterile cerebritis.
Next-generation risk assessment (NGRA) involves the combination of in vitro and in silico models for more human-relevant, ethical, and sustainable human chemical safety assessment. NGRA requires a quantitative mechanistic understanding of the effects of chemicals across human biology (be they molecular, cellular, organ level or higher) coupled with a quantitative understanding of the uncertainty in any experimentally measured or predicted values. These values with their uncertainties can then be considered as a probability distribution, which can then be compared to exposure estimates to establish the presence or absence of a margin of safety. We have constructed Bayesian learning neural networks to provide such quantitative predictions and uncertainties for 20 pharmacologically important human molecular initiating events. These models produce high quality quantitative estimates (p(IC50), p(EC50), p(Ki), p(Kd)) of biochemical activity at a molecular initiating event (MIE) with average mean absolute errors (in Log units) of 0.625 +/- 0.048 in test data and 0.941 +/- 0.215 in external validation data. The key advantage of these models is their ability to also produce standard deviations and credible intervals (CIs) to quantify the uncertainty in these predictions, which we show to be able to distinguish between molecules close to the training data in chemical structure, those less similar to the training data, and decoy compounds drawn from the wider ChEMBL database. These uncertainty values mean that when a prediction is made a user can understand the certainty of the prediction, similar to a quantitative applicability domain, aiding prediction usefulness in NGRA. The ability for in silico methods to produce quantitative predictions with these kinds of probability distributions will be vital to their further use in NGRA, and here clear first steps have been taken.
In silico (computational) methods continue to evolve as part of a robust 21st century public health strategy in risk assessment, relevant to all sectors of chemical safety including preclinical drug discovery, industrial chemicals testing, food and cosmetics. Alongside in vitro methods as components of intelligent testing and pathway driven strategies, in silico models provide the potential for more human relevant solutions to the use of animals in safety testing and biomedical research. These are often termed 'New Approach Methodologies' (NAMs). Some NAMs incorporate the use of 'big data', for example the information provided from high throughput or high content in vitro screening assays or 'omics' technologies. Big data has increasing relevance to predictive toxicology but must be appropriately defined, particularly with regard to 'quality vs quantity'. The purpose of this article is to provide a commentary on the progress of in silico human-based research methods within the context of NAMs, as well as discussion of the emerging use of big data with relevance to safety assessment. The current status of in silico methods is discussed, with input from researchers in the field. Scientific and legislative drivers for change are also considered, along with next steps to address challenges in funding and recognition, to achieve regulatory acceptance and uptake within the research community. To provide some wider context, the use of in silico methods alongside other relevant approaches (e.g., human-based in vitro) is also discussed.
Presentation to the Society of Toxicology annual meeting March 2021, March 12–26, 2021, Virtual/Abstract submission
BACKGROUND:Humans are exposed to tens of thousands of chemical substances that need to be assessed for their potential toxicity. Acute systemic toxicity testing serves as the basis for regulatory hazard classification, labeling, and risk management. However, it is cost- and time-prohibitive to evaluate all new and existing chemicals using traditional rodent acute toxicity tests. In silico models built using existing data facilitate rapid acute toxicity predictions without using animals. OBJECTIVES:The U.S. Interagency Coordinating Committee on the Validation of Alternative Methods (ICCVAM) Acute Toxicity Workgroup organized an international collaboration to develop in silico models for predicting acute oral toxicity based on five different end points: Lethal Dose 50 (LD50 value, U.S. Environmental Protection Agency hazard (four) categories, Globally Harmonized System for Classification and Labeling hazard (five) categories, very toxic chemicals [LD50 (LD50≤50mg/kg)], and nontoxic chemicals (LD50>2,000mg/kg). METHODS:An acute oral toxicity data inventory for 11,992 chemicals was compiled, split into training and evaluation sets, and made available to 35 participating international research groups that submitted a total of 139 predictive models. Predictions that fell within the applicability domains of the submitted models were evaluated using external validation sets. These were then combined into consensus models to leverage strengths of individual approaches. RESULTS:The resulting consensus predictions, which leverage the collective strengths of each individual model, form the Collaborative Acute Toxicity Modeling Suite (CATMoS). CATMoS demonstrated high performance in terms of accuracy and robustness when compared with in vivo results. DISCUSSION:CATMoS is being evaluated by regulatory agencies for its utility and applicability as a potential replacement for in vivo rat acute oral toxicity studies. CATMoS predictions for more than 800,000 chemicals have been made available via the National Toxicology Program's Integrated Chemical Environment tools and data sets (ice.ntp.niehs.nih.gov). The models are also implemented in a free, standalone, open-source tool, OPERA, which allows predictions of new and untested chemicals to be made. https://doi.org/10.1289/EHP8495.
Having a measure of confidence in computational predictions of biological activity from in silico tools is vital when making predictions for new chemicals, for example, in chemical risk assessment. Where predictions of biological activity are used as an indicator of a potential hazard, false-negative predictions are the most concerning prediction; however, assigning confidence in inactive predictions is particularly challenging. How can one confidently identify the absence of activating features? In this study, we present methods for assigning confidence to both active and inactive predictions from structural alerts for protein-binding molecular initiating events (MIEs). Structural alerts were derived through an iterative statistical method. Confidence in the activity predictions is assigned by measuring the Tanimoto similarity between Morgan fingerprints of chemicals in the test set to relevant chemicals in the training set, and suitable cutoff values have been defined to give different confidence categories. To avoid a potential compound series bias in the test set and hence overestimate the performance of the method, we measured the biological activity of 27 compounds with 24 proteins, which gave us an additional 648 experimental measurements; many of the measurements are currently nonexistent in the literature and databases. This data set was complemented with newly measured biological activities published in ChEMBL25 and formed a combined independent validation data set. Applying the confidence categories to the computational predictions for the new data leads to the identification of chemicals for which one should be confident of either an inactive or active prediction, allowing model predictions to be used responsibly.
There is a growing recognition that application of mechanistic approaches to understand cross-species shared molecular targets and pathway conservation in the context of hazard characterization, provide significant opportunities in risk assessment (RA) for both human health and environmental safety. Specifically, it has been recognized that a more comprehensive and reliable understanding of similarities and differences in biological pathways across a variety of species will better enable cross-species extrapolation of potential adverse toxicological effects. Ultimately, this would also advance the generation and use of mechanistic data for both human health and environmental RA. A workshop brought together representatives from industry, academia and government to discuss how to improve the use of existing data, and to generate new NAMs data to derive better mechanistic understanding between humans and environmentally-relevant species, ultimately resulting in holistic chemical safety decisions. Thanks to a thorough dialogue among all participants, key challenges, current gaps and research needs were identified, and potential solutions proposed. This discussion highlighted the common objective to progress toward more predictive, mechanistically based, data-driven and animal-free chemical safety assessments. Overall, the participants recognized that there is no single approach which would provide all the answers for bridging the gap between mechanism-based human health and environmental RA, but acknowledged we now have the incentive, tools and data availability to address this concept, maximizing the potential for improvements in both human health and environmental RA.
Molecular initiating events (MIEs) are key events in adverse outcome pathways that link molecular chemistry to target biology. As they are based on chemistry, these interactions are excellent targets for computational chemistry approaches to in silico modeling. In this work, we aim to link ligand chemical structures to MIEs for androgen receptor (AR) and glucocorticoid receptor (GR) binding using ToxCast data. This has been done using an automated computational algorithm to perform maximal common substructure searches on chemical binders for each target from the ToxCast dataset. The models developed show a high level of accuracy, correctly assigning 87.20% of AR binders and 96.81% of GR binders in a 25% test set using holdout cross-validation. The 2D structural alerts developed can be used as in silico models to predict these MIEs and as guidance for in vitro ToxCast assays to confirm hits. These models can target such experimental work, reducing the number of assays to be performed to gain required toxicological insight. Development of these models has also allowed some structural alerts to be identified as predictors for agonist or antagonist behavior at the receptor target. This work represents a first step in using computational methods to guide and target experimental approaches.
The role of computers in science has changed dramatically because of the increase in computational power, accessible platforms for data storage and use, and the development of artificial intelligence and machine learning. This chapter addresses a number of important questions regarding the role of computers in science and presents some relevant examples.
Disruption of mitochondrial function selectively targets tumour cells that are dependent on oxidative phosphorylation. However, due to their high energy demands, cardiac cells are disproportionately targeted by mitochondrial toxins resulting in a loss of cardiac function. An analysis of the effects of mubritinib on cardiac cells showed that this drug did not inhibit HER2 as reported, but directly inhibits mitochondrial respiratory complex I, reducing cardiac-cell beat rate, with prolonged exposure resulting in cell death. We used a library of chemical variants of mubritinib and showed that modifying the 1H-1,2,3-triazole altered complex I inhibition, identifying the heterocyclic 1,3-nitrogen motif as the toxicophore. The same toxicophore is present in a second anti-cancer therapeutic carboxyamidotriazole (CAI) and we demonstrate that CAI also functions through complex I inhibition, mediated by the toxicophore. Complex I inhibition is directly linked to anti-cancer cell activity, with toxicophore modification ablating the desired effects of these compounds on cancer cell proliferation and apoptosis.
Deep learning neural networks, constructed for the prediction of chemical binding at 79 pharmacologically important human biological targets, show extremely high performance on test data (accuracy 92.2 ± 4.2%, MCC 0.814 ± 0.093 and ROC-AUC 0.96 ± 0.04). A new molecular similarity measure, Neural Network Activation Similarity, has been developed, based on signal propagation through the network. This is complementary to standard Tanimoto similarity, and the combined use increases confidence in the computer's prediction of activity for new chemicals by providing a greater understanding of the underlying justification. The in silico prediction of these human molecular initiating events is central to the future of chemical safety risk assessment and improves the efficiency of safety decision making.
In recent times, machine learning has become increasingly prominent in predictive toxicology as it has shifted from in vivo studies toward in silico studies. Currently, in vitro methods together with other computational methods such as quantitative structure-activity relationship modeling and absorption, distribution, metabolism, and excretion calculations are being used. An overview of machine learning and its applications in predictive toxicology is presented here, including support vector machines (SVMs), random forest (RF) and decision trees (DTs), neural networks, regression models, naïve Bayes, k-nearest neighbors, and ensemble learning. The recent successes of these machine learning methods in predictive toxicology are summarized, and a comparison of some models used in predictive toxicology is presented. In predictive toxicology, SVMs, RF, and DTs are the dominant machine learning methods due to the characteristics of the data available. Lastly, this review describes the current challenges facing the use of machine learning in predictive toxicology and offers insights into the possible areas of improvement in the field.
In the last decade, adverse outcome pathways have been introduced in the fields of toxicology and risk assessment of chemicals as pragmatic tools with broad application potential. While their use in the pharmaceutical and cosmetics sectors has been well documented, their application in the food area remains largely unexplored. In this respect, an expert group of the International Life Sciences Institute Europe has recently explored the use of adverse outcome pathways in the safety evaluation of food additives. A key activity was the organization of a workshop, gathering delegates from the regulatory, industrial and academic areas, to discuss the potentials and challenges related to the application of adverse outcome pathways in the safety assessment of food additives. The present paper describes the outcome of this workshop followed by a number of critical considerations and perspectives defined by the International Life Sciences Institute Europe expert group.
A key question in machine learning is how learning is accomplished. Humans learn through experience in our lives, and it would be extremely beneficial to have machines that can do the same thing. Therefore, we should try to develop them, especially as machines are far better at some tasks than humans, such as large-scale mathematical calculations and high-capacity data storage. But how to find a way to make a machine learn? This is an entirely different problem, which this chapter addresses.
Claudia Rivetti , Timothy E. H. Allen , James B. Brown , Emma Butler , Paul L. Carmichael 5 , John K. Colbourne , Matthew Dent , Francesco Falciani , Lina Gunnarsson , Steve Gutsell 6 , Joshua A. Harrill , Geoff Hodges , Paul Jennings , Richard Judson , Aude Kienzler , 7 Luigi Margiotta-Casaluci , Iris Muller , Stewart F. Owen , Cecilie Rendal , Paul J. Russell 8 , Sharon Scott , Fiona Sewell , Imran Shah , Ian Sorrel , Mark R. Viant , Carl 9 Westmoreland , Andrew White , Bruno Campos 1* 10