Selecting the best model of sequence evolution for a multiple-sequence-alignment (MSA) constitutes the first step of phylogenetic tree reconstruction. Common approaches for inferring nucleotide models typically apply maximum likelihood (ML) methods, with discrimination between models determined by one of several information criteria. This requires tree reconstruction and optimisation which can be computationally expensive. We demonstrate that neural networks can be used to perform model selection, without the need to reconstruct trees, optimise parameters, or calculate likelihoods. We introduce ModelRevelator, a model selection tool underpinned by two deep neural networks. The first neural network, NNmodelfind, recommends one of six commonly used models of sequence evolution, ranging in complexity from Jukes and Cantor to General Time Reversible. The second, NNalphafind, recommends whether or not a Γ-distributed rate heterogeneous model should be incorporated, and if so, provides an estimate of the shape parameter, ɑ. Users can simply input an MSA into ModelRevelator, and swiftly receive output recommending the evolutionary model, inclusive of the presence or absence of rate heterogeneity, and an estimate of ɑ. We show that ModelRevelator performs comparably with likelihood-based methods and the recently published machine learning method ModelTeller over a wide range of parameter settings, with significant potential savings in computational effort. Further, we show that this performance is not restricted to the alignments on which the networks were trained, but is maintained even on unseen empirical data. We expect that ModelRevelator will provide a valuable alternative for phylogeneticists, especially where traditional methods of model selection are computationally prohibitive.
The emergence of novel SARS coronavirus 2 (SARS-CoV-2) in 2019 has triggered an ongoing global pandemic of severe pneumonia-like disease designated as coronavirus disease 2019 (COVID-19). To date, more than 2.1 million confirmed cases and 139,500 deaths have been reported worldwide, and there are currently no medical countermeasures available to prevent or treat the disease. As the development of a vaccine could require at least 12-18 months, and the typical timeline from hit finding to drug registration of an antiviral is >10 years, repositioning of known drugs can significantly accelerate the development and deployment of therapies for COVID-19. To identify therapeutics that can be repurposed as SARS-CoV-2 antivirals, we profiled a library of known drugs encompassing approximately 12,000 clinical-stage or FDA-approved small molecules. Here, we report the identification of 30 known drugs that inhibit viral replication. Of these, six were characterized for cellular dose-activity relationships, and showed effective concentrations likely to be commensurate with therapeutic doses in patients. These include the PIKfyve kinase inhibitor Apilimod, cysteine protease inhibitors MDL-28170, Z LVG CHN2, VBY-825, and ONO 5334, and the CCR1 antagonist MLN-3897. Since many of these molecules have advanced into the clinic, the known pharmacological and human safety profiles of these compounds will accelerate their preclinical and clinical evaluation for COVID-19 treatment.
Maximum likelihood and maximum parsimony are two key methods for phylogenetic tree reconstruction. Under certain conditions, each of these two methods can perform more or less efficiently, resulting in unresolved or disputed phylogenies. We show that a neural network can distinguish between four-taxon alignments that were evolved under conditions susceptible to either long-branch attraction or long-branch repulsion. When likelihood and parsimony methods are discordant, the neural network can provide insight as to which tree reconstruction method is best suited to the alignment. When applied to the contentious case of Strepsiptera evolution, our method shows robust support for the current scientific view, that is, it places Strepsiptera with beetles, distant from flies.
Wikidata is a community-maintained knowledge base that has been assembled from repositories in the fields of genomics, proteomics, genetic variants, pathways, chemical compounds, and diseases, and that adheres to the FAIR principles of findability, accessibility, interoperability and reusability. Here we describe the breadth and depth of the biomedical knowledge contained within Wikidata, and discuss the open-source tools we have built to add information to Wikidata and to synchronize it with source databases. We also demonstrate several use cases for Wikidata, including the crowdsourced curation of biomedical ontologies, phenotype-based diagnosis of disease, and drug repurposing.
The emergence of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) in 2019 has triggered an ongoing global pandemic of the severe pneumonia-like disease coronavirus disease 2019 (COVID-19) 1 . The development of a vaccine is likely to take at least 12–18 months, and the typical timeline for approval of a new antiviral therapeutic agent can exceed 10 years. Thus, repurposing of known drugs could substantially accelerate the deployment of new therapies for COVID-19. Here we profiled a library of drugs encompassing approximately 12,000 clinical-stage or Food and Drug Administration (FDA)-approved small molecules to identify candidate therapeutic drugs for COVID-19. We report the identification of 100 molecules that inhibit viral replication of SARS-CoV-2, including 21 drugs that exhibit dose–response relationships. Of these, thirteen were found to harbour effective concentrations commensurate with probable achievable therapeutic doses in patients, including the PIKfyve kinase inhibitor apilimod 2 – 4 and the cysteine protease inhibitors MDL-28170, Z LVG CHN2, VBY-825 and ONO 5334. Notably, MDL-28170, ONO 5334 and apilimod were found to antagonize viral replication in human pneumocyte-like cells derived from induced pluripotent stem cells, and apilimod also demonstrated antiviral efficacy in a primary human lung explant model. Since most of the molecules identified in this study have already advanced into the clinic, their known pharmacological and human safety profiles will enable accelerated preclinical and clinical evaluation of these drugs for the treatment of COVID-19.
To develop a map of cell-cell communication mediated by extracellular RNA (exRNA), the NIH Extracellular RNA Communication Consortium created the exRNA Atlas resource (https://exrna-atlas.org). The Atlas version 4P1 hosts 5,309 exRNA-seq and exRNA qPCR profiles from 19 studies and a suite of analysis and visualization tools. To analyze variation between profiles, we apply computational deconvolution. The analysis leads to a model with six exRNA cargo types (CT1, CT2, CT3A, CT3B, CT3C, CT4), each detectable in multiple biofluids (serum, plasma, CSF, saliva, urine). Five of the cargo types associate with known vesicular and non-vesicular (lipoprotein and ribonucleoprotein) exRNA carriers. To validate utility of this model, we re-analyze an exercise response study by deconvolution to identify physiologically relevant response pathways that were not detected previously. To enable wide application of this model, as part of the exRNA Atlas resource, we provide tools for deconvolution and analysis of user-provided case-control studies.
The chemical diversity and known safety profiles of drugs previously tested in humans make them a valuable set of compounds to explore potential therapeutic utility in indications outside those originally targeted, especially neglected tropical diseases. This practice of "drug repurposing" has become commonplace in academic and other nonprofit drug-discovery efforts, with the appeal that significantly less time and resources are required to advance a candidate into the clinic. Here, we report a comprehensive open-access, drug repositioning screening set of 12,000 compounds (termed ReFRAME; Repurposing, Focused Rescue, and Accelerated Medchem) that was assembled by combining three widely used commercial drug competitive intelligence databases (Clarivate Integrity, GVK Excelra GoStar, and Citeline Pharmaprojects), together with extensive patent mining of small molecules that have been dosed in humans. To date, 12,000 compounds (∼80% of compounds identified from data mining) have been purchased or synthesized and subsequently plated for screening. To exemplify its utility, this collection was screened against Cryptosporidium spp., a major cause of childhood diarrhea in the developing world, and two active compounds previously tested in humans for other therapeutic indications were identified. Both compounds, VB-201 and a structurally related analog of ASP-7962, were subsequently shown to be efficacious in animal models of Cryptosporidium infection at clinically relevant doses, based on available human doses. In addition, an open-access data portal (https://reframedb.org) has been developed to share ReFRAME screen hits to encourage additional follow-up and maximize the impact of the ReFRAME screening collection.
Wikidata, a project of the Wikimedia Foundation, is an openly editable, semantic web-compatible framework for knowledge management. Wikidata has a large and active community that contributes to, maintains, and improves the quality of the data in Wikidata as well as deciding how the data itself should be represented. Our team has been populating Wikidata with a foundational semantic network linking genes, proteins, drugs, and diseases. Upon this foundation, we hope to stimulate the growth of this knowledge graph that can be used to build new knowledge-based applications that drive new discoveries. A cornerstone of the Wikidata knowledge graph is its built-in model for tracking the evidence underlying claims. For any claim in the graph (e.g. gene A regulates gene B), it is possible to provide evidence supporting or refuting that claim. The manner in which these evidence statements can be constructed is left open for the community to decide. Therefore, the patterns for representing the semantics of the associated evidence and the provenance trails linking back to the original sources of information must be defined and consistently used, such that this information is easily accessible by the end user or by software that exposes this information to end users.
With the advancement of genome sequencing technologies, new genomes are being sequenced daily. While these sequences are deposited in publicly available data warehouses, their functional and genomic annotations (beyond genes which are predicted automatically) mostly reside in the text of primary publications. Professional curators are hard at work extracting those annotations from the literature for the most studied organisms and depositing them in structured databases. However, the resources don’t exist to fund the comprehensive curation of the thousands of newly sequenced organisms in this manner. Here, we describe WikiGenomes ( wikigenomes.org ), a web application that facilitates the consumption and curation of genomic data by the entire scientific community. WikiGenomes is based on Wikidata, an openly editable knowledge graph with the goal of aggregating published knowledge into a free and open database. WikiGenomes empowers the individual genomic researcher to contribute their expertise to the curation effort and integrates the knowledge into Wikidata, enabling it to be accessed by anyone without restriction.
Ten steps for integrating CIViC DB into Wikidata. What you have to do to get your data into the melting pot that is Wikidata. Make it free, open and accessible to everybody.
—Wikidata is a world readable and writable knowledge base maintained by the Wikimedia Foundation. It offers the opportunity to collaboratively construct a fully open access knowledge graph spanning biology, medicine, and all other domains of knowledge. To meet this potential, social and technical challenges must be overcome most of which are familiar to the biocuration community. These include community ontology building, high precision information extraction, provenance, and license management. By working together with Wikidata now, we can help shape it into a trustworthy, unencumbered central node in the Semantic Web of biomedical data.
The last 20 years of advancement in sequencing technologies have led to sequencing thousands of microbial genomes, creating mountains of genetic data. While efficiency in generating the data improves almost daily, applying meaningful relationships between taxonomic and genetic entities on this scale requires a structured and integrative approach. Currently, knowledge is distributed across a fragmented landscape of resources from government-funded institutions such as National Center for Biotechnology Information (NCBI) and UniProt to topic-focused databases like the ODB3 database of prokaryotic operons, to the supplemental table of a primary publication. A major drawback to large scale, expert-curated databases is the expense of maintaining and extending them over time. No entity apart from a major institution with stable long-term funding can consider this, and their scope is limited considering the magnitude of microbial data being generated daily. Wikidata is an openly editable, semantic web compatible framework for knowledge representation. It is a project of the Wikimedia Foundation and offers knowledge integration capabilities ideally suited to the challenge of representing the exploding body of information about microbial genomics. We are developing a microbial specific data model, based on Wikidata's semantic web compatibility, which represents bacterial species, strains and the gene and gene products that define them. Currently, we have loaded 43,694 gene and 37,966 protein items for 21 species of bacteria, including the human pathogenic bacteriaChlamydia trachomatis.Using this pathogen as an example, we explore complex interactions between the pathogen, its host, associated genes, other microbes, disease and drugs using the Wikidata SPARQL endpoint. In our next phase of development, we will add another 99 bacterial genomes and their gene and gene products, totaling ∼900,000 additional entities. This aggregation of knowledge will be a platform for community-driven collaboration, allowing the networking of microbial genetic data through the sharing of knowledge by both the data and domain expert.
IMPORTANCEDespite the unquestioned relationship of UV radiation (UVR) exposure and melanoma development, UVR-independent development of melanoma has only recently been described in mice. These findings in mice highlight the importance of the genetic background of the host and could be relevant for preventive measures in humans.OBJECTIVETo study the role of the melanocortin-1 receptor (MC1R) and melanoma risk independently from UVR in a clinical setting.DESIGN, SETTING, AND PARTICIPANTSHospital-based case-control study, including genetic testing, questionnaires, and physical data (Molecular Markers of Melanoma Study data set) including 991 melanoma patients (cases) and 800 controls.MAIN OUTCOMES AND MEASURESAssociation of MC1R variants and melanoma risk independent from sun exposure variables.RESULTSThe 1791 participants included 991 with a diagnosis of melanoma and 800 control patients (mean [SD] age, 59.2 [15.6] years; 50.5% male). Compared with wild-type carriers, carriers of MC1R variants were at higher melanoma risk after statistically adjusting for previous UVR exposure (represented by prior sunburns and signs of actinic skin damage identified by dermatologists), age, and sex compared with wild-type carriers (≥2 variants, OR, 2.13 [95% CI, 1.66-2.75], P < .001; P for trend <.001). After adjustment for sex, age, sunburns in the past, and signs of actinic skin damage, the associations remained significant (OR, 1.65 [95% CI, 1.02-2.67] for R/R, OR, 2.63 [95% CI, 1.82-3.81] for R/r; OR, 1.83 [95% CI, 1.36-2.48] for R/0; and OR, 1.50 [95% CI, 1.01-2.21] for r/r, with P values ranging from <.001 to .04 when adjusted for facial actinic skin damage; OR, 2.36 [95% CI, 1.62-3.43] for R/r; and OR, 1.47 [95% CI, 1.08-1.99] for R/0 with P values ranging from <.001 to .01 when adjusted for dorsal actinic skin damage; and OR, 2.54 [95% CI, 1.76-3.67] for R/r, OR, 1.75 [95% CI, 1.30-2.36] for R/0; and OR, 1.50 [95% CI, 1.02-2.20] for r/r with P values ranging from <.001 to .04 when adjusted for actinic skin damage on the hands).CONCLUSIONS AND RELEVANCECarriers of MC1R variants were at increased melanoma risk independent of their sun exposure. Further studies are required to elucidate the causes of melanoma development in these individuals.