We show that the latent vectors, broken down to atom and fragment contributions, produced by a graph convolutional network (GCN) encode chemically meaningful information. Latent vectors for atoms, further characterized by their neighbors, and fragments are largely independent of the rest of the molecule. This observation extends to transformations, making the organization of the latent space comparable to the capturing of syntax and semantics in word embeddings such as word2vec. Applying this observation to bioisostere identification, we find that bioisosteric fragment pairs are also closer in latent space than random pairs. Building on this result, we developed a method to generate bioisosteric suggestions.
We report a silver nanoparticle (AgNP) sensor array to detect SARS-CoV-2 viral particles utilizing five peptide receptors to produce a colorimetric response. We show that the interaction of the virus particles with the peptides interferes with the shift in plasmonic absorbance and hinders the formation of snowflake-like fractal assemblies between the AgNPs and the bridging peptides. The limit of detection in water of the sensor array was 500,000 viral copies/mL. We demonstrate the capabilities of the sensor array to distinguish between SARS-CoV-2 (Beta, Delta, and Omicron variants) from the influenza virus at the 99% confidence level through linear discriminant analysis. The sensor array was also able to discriminate three out of five negative exhaled breath condensate samples from five positive SARS-CoV-2 EBC samples.
We present a standardized metadata template for assays used in pharmaceutical drug discovery research, according to the FAIR principles. We also describe the use of an automated tool for annotating assays from a variety of sources, including PubChem, commercial assay providers, and the peer-reviewed literature, to this metadata template. Adoption of a standardized metadata template will allow drug discovery scientists to better understand and compare the increasing amounts of assay data becoming available, and will facilitate the use of artificial intelligence tools and other computational methods for analysis and prediction. Since bioassays drive advances in biomedical research, improvements in assay metadata can improve productivity in discovery of new therapeutics, platform technologies, and assay methods.
RATIONALE:Preventive health is a core part of primary care clinical practice and it is critical for both disease prevention and reducing the consequences of chronic disease. In primary care, the 5As framework is often used to guide behaviour change consultations for smoking, nutrition, alcohol use and physical activity.AIMS AND OBJECTIVES:Our objective was to analyze the emphasis placed on each 5As term in commonly used guidelines in Australian general practice and compare this to behaviour change terms/concepts essential to effective consultations.METHOD:A content analysis was undertaken to explore frequency of 5A terms and key behaviour change concepts/terms chapter-by-chapter across the three most commonly used guidelines in Australian general practice.RESULTS:The prevalence of each 5As term differed in all three guidelines, with 'Arrange' being mentioned the least often. Behaviour change concepts and terms, such as patient-centredness, listening, trust and tailoring, were infrequently used and were often confined to a separate chapter of the guidelines.CONCLUSION:The language and content of the guidelines contrast with known effective components of behaviour change consultations. Future revisions could reconsider emphasis of 5As terms to avoid paternalistic approaches, improve shared language across guidelines and incorporate behavioural science principles to enhance preventative care delivery.
AIMS To map the existing body of heart failure (HF) telehealth interventions for vulnerable populations, and to conduct an intersectionality-based analysis utilizing a structured checklist. DESIGN A scoping review and intersectionality-based analysis. DATA SOURCES The search was conducted in March 2022 in the following databases: MEDLINE, CINAHL, Scopus and the Cochrane Central Register of Controlled Trials, ProQuest Dissertations and Theses Global. REVIEW METHODS First, the titles and abstracts were screened, and then the entire articles were screened against the inclusion criteria. Two of the investigators screened the articles independently in Covidence. The studies included and excluded at various stages of screening were depicted through a PRISMA flow diagram. The quality of the included studies was assessed based on the mixed methods appraisal tool (MMAT). Each study was read thoroughly and the intersectionality-based checklist by Ghasemi et al. (2021) was applied, whereby a yes/no response was marked for each question on the checklist and the relevant supporting data were extracted. RESULTS A total of 22 studies were included in this review. About 42.2% of the responses indicated that studies incorporated the principles of intersectionality at the 'problem identification' stage, followed by 42.9% and 29.44% responses indicating incorporation of these principles at the 'design and implementation' and 'evaluation' stages respectively. CONCLUSIONS The findings suggest that the research around HF telehealth interventions for vulnerable populations is not adequately grounded in appropriate theoretical underpinning. The principles of intersectionality have been applied mostly to the problem identification and the intervention development and implementation stages, and not so much at the evaluation stage. Future research must fill the identified gaps in this area of research. NO PATIENT OR PUBLIC CONTRIBUTION Since this was a scoping, there was no patient contribution to this work; however, based on this study's findings, we are undertaking patient-centred studies with patient contribution.
The prevalence of “Long COVID”, including among vaccinated patients, is just one of the conundrums that indicate how much remains unknown about the lung’s response to viral infection, particularly to SARS-CoV-2 for which the lung is the point of entry. Therefore, we used an in vitro human lung system to enable a prospective, unbiased, sequential single cell level analysis of pulmonary cell responses following infection by multiple strains of SARS-CoV-2. By starting with human induced pluripotent stem cells (hiPSCs) and emulating lung organogenesis, three-dimensional lung organoids were generated and infected in which several unexpected but pertinent insights emerged. First, SARS-CoV-2 tropism is much broader than previously believed: most lung cell types can be infected, if not through a canonical receptor-mediated route (e.g., via ACE2) then via a non-canonical “backdoor” endocytosis/micropinocytosis route. Such entry can be abrogated by FDA-approved endocytosis blockers, suggesting novel adjunctive therapies. Regardless of route-of-entry, the virus triggers a heretofore unrecognized lung epithelial cell-intrinsic autonomous innate immune response involving interferons and cytokine/chemokine production in the absence of hematopoietic cells or their derivatives. The virus can spread rapidly throughout human lung organoid cell cultures resulting in mitochondrial apoptosis mediated by the pro-survival protein Bcl-xL. This host cytopathic response to the virus may help explain persistent inflammatory signatures in a dysfunctional pulmonary environment of long COVID. The host response to the virus is, in part, dependent on the presence of pulmonary Surfactant Protein-B (SP-B), which plays an unanticipated role in signal transduction, viral resistance, dampens systemic inflammatory cytokine production, and minimizes the induction of apoptosis.
Substances of unknown or variable composition, complex reaction products, or biological materials (UVCBs) are over 70 000 "complex" chemical mixtures produced and used at significant levels worldwide. Due to their unknown or variable composition, applying chemical assessments originally developed for individual compounds to UVCBs is challenging, which impedes sound management of these substances. Across the analytical sciences, toxicology, cheminformatics, and regulatory practice, new approaches addressing specific aspects of UVCB assessment are being developed, albeit in a fragmented manner. This review attempts to convey the "big picture" of the state of the art in dealing with UVCBs by holistically examining UVCB characterization and chemical identity representation, as well as hazard, exposure, and risk assessment. Overall, information gaps on chemical identities underpin the fundamental challenges concerning UVCBs, and better reporting and substance characterization efforts are needed to support subsequent chemical assessments. To this end, an information level scheme for improved UVCB data collection and management within databases is proposed. The development of UVCB testing shows early progress, in line with three main methods: whole substance, known constituents, and fraction profiling. For toxicity assessment, one option is a whole-mixture testing approach. If the identities of (many) constituents are known, grouping, read across, and mixture toxicity modeling represent complementary approaches to overcome data gaps in toxicity assessment. This review highlights continued needs for concerted efforts from all stakeholders to ensure proper assessment and sound management of UVCBs.
We review the disease, biology and biochemistry of kinetoplastids, as well as the new drugs and drug candidates that have entered the clinic in the last decade. We also describe examples of the pre-clinical exploration of small molecules against various protein targets, (e.g., cysteine proteases, the proteasome and tubulin), as well as cutting-edge molecular and computational strategies, and technologies being brought to bear to discover and develop new anti-trypanosomal drugs. For comprehensive descriptions of the disease, biology and drug therapies prior to 2011, the reader is encouraged to review the chapter by P.M. Woster that appeared in 2010 in the seventh edition of Burger’s Medicinal Chemistry, Drug Discovery, and Development, with the title Antiprotozoal/Antiparasitic Agents.
Tuberculosis is a global health dilemma. In 2016, the WHO reported 10.4 million incidences and 1.7 million deaths. The need to develop new treatments for those infected with Mycobacterium tuberculosis (Mtb) has led to many large-scale phenotypic screens and many thousands of new active compounds identified in vitro. However, with limited funding, efforts to discover new active molecules against Mtb needs to be more efficient. Several computational machine learning approaches have been shown to have good enrichment and hit rates. We have curated small molecule Mtb data and developed new models with a total of 18,886 molecules with activity cutoffs of 10 μM, 1 μM, and 100 nM. These data sets were used to evaluate different machine learning methods (including deep learning) and metrics and to generate predictions for additional molecules published in 2017. One Mtb model, a combined in vitro and in vivo data Bayesian model at a 100 nM activity yielded the following metrics for 5-fold cross validation: accuracy = 0.88, precision = 0.22, recall = 0.91, specificity = 0.88, kappa = 0.31, and MCC = 0.41. We have also curated an evaluation set (n = 153 compounds) published in 2017, and when used to test our model, it showed the comparable statistics (accuracy = 0.83, precision = 0.27, recall = 1.00, specificity = 0.81, kappa = 0.36, and MCC = 0.47). We have also compared these models with additional machine learning algorithms showing Bayesian machine learning models constructed with literature Mtb data generated by different laboratories generally were equivalent to or outperformed deep neural networks with external test sets. Finally, we have also compared our training and test sets to show they were suitably diverse and different in order to represent useful evaluation sets. Such Mtb machine learning models could help prioritize compounds for testing in vitro and in vivo.
ADVERTISEMENT RETURN TO ISSUEPREVAddition and Correct...Addition and CorrectionNEXTORIGINAL ARTICLEThis notice is a correctionCorrection to "Comparing and Validating Machine Learning Models for Mycobacterium tuberculosis Drug Discovery"Thomas LaneThomas LaneMore by Thomas Lane, Daniel P. RussoDaniel P. RussoMore by Daniel P. Russo, Kimberley M. ZornKimberley M. ZornMore by Kimberley M. Zorn, Alex M. ClarkAlex M. ClarkMore by Alex M. Clark, Alexandru KorotcovAlexandru KorotcovMore by Alexandru Korotcov, Valery TkachenkoValery TkachenkoMore by Valery Tkachenko, Robert C. ReynoldsRobert C. ReynoldsMore by Robert C. Reynolds, Alexander L. PerrymanAlexander L. PerrymanMore by Alexander L. Perryman, Joel S. FreundlichJoel S. FreundlichMore by Joel S. Freundlich, and Sean Ekins*Sean EkinsMore by Sean Ekinshttps://orcid.org/0000-0002-5691-5790Cite this: Mol. Pharmaceutics 2021, 18, 7, 2833Publication Date (Web):June 17, 2021Publication History Published online17 June 2021Published inissue 5 July 2021https://pubs.acs.org/doi/10.1021/acs.molpharmaceut.1c00428https://doi.org/10.1021/acs.molpharmaceut.1c00428correctionACS PublicationsCopyright © 2021 American Chemical Society. This publication is available under these Terms of Use. Request reuse permissions This publication is free to access through this site. Learn MoreArticle Views843Altmetric-Citations-LEARN ABOUT THESE METRICSArticle Views are the COUNTER-compliant sum of full text article downloads since November 2008 (both PDF and HTML) across all institutions and individuals. These metrics are regularly updated to reflect usage leading up to the last few days.Citations are the number of other articles citing this article, calculated by Crossref and updated daily. Find more information about Crossref citation counts.The Altmetric Attention Score is a quantitative measure of the attention that a research article has received online. Clicking on the donut icon will load a page at altmetric.com with additional details about the score and the social media presence for the given article. Find more information on the Altmetric Attention Score and how the score is calculated. Share Add toView InAdd Full Text with ReferenceAdd Description ExportRISCitationCitation and abstractCitation and referencesMore Options Share onFacebookTwitterWechatLinked InRedditEmail PDF (415 KB) Get e-AlertscloseSUBJECTS:Bacteria,Drug discovery,Graduate education,Machine learning,Therapeutics Get e-Alerts
BACKGROUND:Humans are exposed to tens of thousands of chemical substances that need to be assessed for their potential toxicity. Acute systemic toxicity testing serves as the basis for regulatory hazard classification, labeling, and risk management. However, it is cost- and time-prohibitive to evaluate all new and existing chemicals using traditional rodent acute toxicity tests. In silico models built using existing data facilitate rapid acute toxicity predictions without using animals. OBJECTIVES:The U.S. Interagency Coordinating Committee on the Validation of Alternative Methods (ICCVAM) Acute Toxicity Workgroup organized an international collaboration to develop in silico models for predicting acute oral toxicity based on five different end points: Lethal Dose 50 (LD50 value, U.S. Environmental Protection Agency hazard (four) categories, Globally Harmonized System for Classification and Labeling hazard (five) categories, very toxic chemicals [LD50 (LD50≤50mg/kg)], and nontoxic chemicals (LD50>2,000mg/kg). METHODS:An acute oral toxicity data inventory for 11,992 chemicals was compiled, split into training and evaluation sets, and made available to 35 participating international research groups that submitted a total of 139 predictive models. Predictions that fell within the applicability domains of the submitted models were evaluated using external validation sets. These were then combined into consensus models to leverage strengths of individual approaches. RESULTS:The resulting consensus predictions, which leverage the collective strengths of each individual model, form the Collaborative Acute Toxicity Modeling Suite (CATMoS). CATMoS demonstrated high performance in terms of accuracy and robustness when compared with in vivo results. DISCUSSION:CATMoS is being evaluated by regulatory agencies for its utility and applicability as a potential replacement for in vivo rat acute oral toxicity studies. CATMoS predictions for more than 800,000 chemicals have been made available via the National Toxicology Program's Integrated Chemical Environment tools and data sets (ice.ntp.niehs.nih.gov). The models are also implemented in a free, standalone, open-source tool, OPERA, which allows predictions of new and untested chemicals to be made. https://doi.org/10.1289/EHP8495.
Chemical mixtures have recently come to the attention of open standards and data structures for capturing machine-readable descriptions for informatics uses. At the present time, essentially all transmission of information about mixtures is done using short text descriptions that are readable only by trained scientists, and there are no accessible repositories of marked-up mixture data. We have designed a machine learning tool that can interpret mixture descriptions and upgrade them to the high-level Mixfile format, which can in turn be used to generate Mixtures InChI notation. The interpretation achieves a high success rate and can be used at scale to markup large catalogs and inventories, with some expert checking to catch edge cases. The training data that was accumulated during the project is made openly available, along with previously released mixture editing tools and utilities.
We have previously described the first Bayesian machine learning models from FDA-approved drug screens, for identifying compounds active against the Ebola virus (EBOV). These models led to the identification of three active molecules in vitro: tilorone, pyronaridine, and quinacrine. A follow-up study demonstrated that one of these compounds, tilorone, has 100% in vivo efficacy in mice infected with mouse-adapted EBOV at 30 mg/kg/day intraperitoneal. This suggested that we can learn from the published data on EBOV inhibition and use it to select new compounds for testing that are active in vivo. We used these previously built Bayesian machine learning EBOV models alongside our chemical insights for the selection of 12 molecules, absent from the training set, to test for in vitro EBOV inhibition. Nine molecules were directly selected using the model, and eight of these molecules possessed a promising in vitro activity (EC50 < 15 μM). Three further compounds were selected for an in vitro evaluation because they were antimalarials, and compounds of this class like pyronaridine and quinacrine have previously been shown to inhibit EBOV. We identified the antimalarial drug arterolane (IC50 = 4.53 μM) and the anticancer clinical candidate lucanthone (IC50 = 3.27 μM) as novel compounds that have EBOV inhibitory activity in HeLa cells and generally lack cytotoxicity. This work provides further validation for using machine learning and medicinal chemistry expertize to prioritize compounds for testing in vitro prior to more costly in vivo tests. These studies provide further corroboration of this strategy and suggest that it can likely be applied to other pathogens in the future.
We have recently stressed the need to scale and “industrialize” rare disease drug discovery. Finding information on compounds relevant to rare diseases across the hundreds of available databases is a complex challenge, even for experts and the linkage between targets and data is often non-existent. Disease-associated targets are only identified after a deep dive into the literature for a specific disease as disease annotation is usually not included alongside publicly accessible structure activity relationship data. This would be valuable so that we could use public datasets for specific targets to build machine learning models through our Assay Central software. This could then help us to identify small molecules that might have value as inhibitors or chaperones for testing with collaborators working on those diseases and move development forward more quickly. To date, many of the lysosomal storage diseases (LSDs) have no FDA approved treatments. As a first step we have mined PubMed and other databases such as ChEMBL to identify and understand disease targets associated with LSDs and available datasets. We will demonstrate selected targets and public datasets we have identified to date for LSDs and show how they can be used to generate machine learning models (Bayesian, Random Forest and Deep learning etc) that can suggest new molecules to test. We ultimately plan to create linkages between these models and disease targets for the 7000 rare and neglected diseases so that we can create hierarchies of protein families and disease categories. This will then enable the selection of specific machine learning models in Assay Central which we can use with our collaborators. Efforts such as this could help in dramatically scaling research and development efforts for rare diseases.
The human immunodeficiency virus (HIV) causes over a million deaths every year and has a huge economic impact in many countries. The first class of drugs approved were nucleoside reverse transcriptase inhibitors. A newer generation of reverse transcriptase inhibitors have become susceptible to drug resistant strains of HIV, and hence, alternatives are urgently needed. We have recently pioneered the use of Bayesian machine learning to generate models with public data to identify new compounds for testing against different disease targets. The current study has used the NIAID ChemDB HIV, Opportunistic Infection and Tuberculosis Therapeutics Database for machine learning studies. We curated and cleaned data from HIV-1 wild-type cell-based and reverse transcriptase (RT) DNA polymerase inhibition assays. Compounds from this database with ≤1 μM HIV-1 RT DNA polymerase activity inhibition and cell-based HIV-1 inhibition are correlated (Pearson r = 0.44, n = 1137, p < 0.0001). Models were trained using multiple machine learning approaches (Bernoulli Naive Bayes, AdaBoost Decision Tree, Random Forest, support vector classification, k-Nearest Neighbors, and deep neural networks as well as consensus approaches) and then their predictive abilities were compared. Our comparison of different machine learning methods demonstrated that support vector classification, deep learning, and a consensus were generally comparable and not significantly different from each other using 5-fold cross validation and using 24 training and test set combinations. This study demonstrates findings in line with our previous studies for various targets that training and testing with multiple data sets does not demonstrate a significant difference between support vector machine and deep neural networks.
A variety of machine learning methods such as naive Bayesian, support vector machines and more recently deep neural networks are demonstrating their utility for drug discovery and development. These leverage the generally bigger datasets created from high-throughput screening data and allow prediction of bioactivities for targets and molecular properties with increased levels of accuracy. We have only just begun to exploit the potential of these techniques but they may already be fundamentally changing the research process for identifying new molecules and/or repurposing old drugs. The integrated application of such machine learning models for end-to-end (E2E) application is broadly relevant and has considerable implications for developing future therapies and their targeting.
We describe a file format that is designed to represent mixtures of compounds in a way that is fully machine readable. This Mixfile format is intended to fill the same role for substances that are composed of multiple components as the venerable Molfile does for specifying individual structures. This much needed datastructure is intended to replace current practices for communicating information about mixtures, which usually relies on human-readable text descriptions, drawing several species within a single molecular diagram, or mutually incompatible ad hoc solutions. We describe an open source software application for editing mixture files, which can also be used as web-ready tools for manipulating the file format. We also present a corpus of mixture examples, which we have extracted from collections of text-based descriptions. Furthermore, we present an early look at the proposed IUPAC Mixtures InChI specification, instances of which can be automatically generated using the Mixfile format as a precursor.