Epidemiology studies evaluate associations between the metabolome and disease risk. Urine is a common biospecimen used for such studies due to its wide availability and non-invasive collection. Evaluating the robustness of urinary metabolomic profiles under varying preanalytical conditions is thus of interest. Here we evaluate the impact of sample handling conditions on urine metabolome profiles relative to the gold standard condition (no preservative, no refrigeration storage, single freeze thaw). Conditions tested included the use of borate or chlorhexidine preservatives, various storage and freeze/thaw cycles. We demonstrate that sample handling conditions impact metabolite levels, with borate showing the largest impact with 125 of 1,048 altered metabolites (adjusted P < 0.05). When simulating a case-control study with expected inconsistencies in sample handling, we predicted the occurrence of false positive altered metabolites to be low (< 11). Predicted false positives increased substantially (³63) when cases were simulated to undergo alternate handling. Finally, we demonstrate that sample handling impacts on the urinary metabolome were markedly smaller than those in serum. While changes in urine metabolites incurred by sample handling are generally small, we recommend implementing consistent handling conditions and evaluating robustness of metabolite measurements for those showing significant associations with disease outcomes.
Understanding the underlying etiologies of rare diseases may facilitate research across multiple conditions, enabling basket trail design and drug repurposing. In this study, we aligned clusters of rare diseases with Orphanet classifications to represent their shared etiologies and establish a foundation for further investigation on underly biological mechanism discovery. By utilizing the linearized Orphanet categories, we connected 35 clusters of rare diseases into 18 classifications. Significant associations were found between the categories "Rare Developmental Defects During Embryogenesis" and "Rare Inborn Errors of Metabolism" and the clusters in this study, suggesting that many rare diseases originating in the prenatal period or related to metabolism may present a substantial opportunity for success in future investigation.
Efficiently circumventing the blood-brain barrier (BBB) poses a major hurdle in the development of drugs that target the central nervous system. Although there are several methods to determine BBB permeability of small molecules, the Parallel Artificial Membrane Permeability Assay (PAMPA) is one of the most common assays in drug discovery due to its robust and high-throughput nature. Drug discovery is a long and costly venture, thus, any advances to streamline this process are beneficial. In this study, ∼2,000 compounds from over 60 NCATS projects were screened in the PAMPA-BBB assay to develop a quantitative structure-activity relationship model to predict BBB permeability of small molecules. After analyzing both state-of-the-art and latest machine learning methods, we found that random forest based on RDKit descriptors as additional features provided the best training balanced accuracy (0.70 ± 0.015) and a message-passing variant of graph convolutional neural network that uses RDKit descriptors provided the highest balanced accuracy (0.72) on a prospective validation set. Finally, we correlated in vitro PAMPA-BBB data with in vivo brain permeation data in rodents to observe a categorical correlation of 77%, suggesting that models developed using data from PAMPA-BBB can forecast in vivo brain permeability. Given that majority of prior research has relied on in vitro or in vivo data for assessing BBB permeability, our model, developed using the largest PAMPA-BBB dataset to date, offers an orthogonal means to estimate BBB permeability of small molecules. We deposited a subset of our data into PubChem bioassay database (AID: 1845228) and deployed the best performing model on the NCATS Open Data ADME portal (https://opendata.ncats.nih.gov/adme/). These initiatives were undertaken with the aim of providing valuable resources for the drug discovery community.
The Knowledge Management Center (KMC) for the Illuminating the Druggable Genome (IDG) project aims to aggregate, update, and articulate protein-centric data knowledge for the entire human proteome, with emphasis on the understudied proteins from the three IDG protein families. KMC collates and analyzes data from over 70 resources to compile the Target Central Resource Database (TCRD), which is the web-based informatics platform (Pharos). These data include experimental, computational, and text-mined information on protein structures, compound interactions, and disease and phenotype associations. Based on this knowledge, proteins are classified into different Target Development Levels (TDLs) for identification of understudied targets. Additional work by the KMC focuses on enriching target knowledge and producing DrugCentral and other data visualization tools for expanding investigation of understudied targets.
Objectives: Pharmacokinetic issues were the leading cause of drug attrition, accounting for approximately 40% of all cases before the turn of the century. To this end, several high-throughput in vitro assays like microsomal stability have been developed to evaluate the pharmacokinetic profiles of compounds in the early stages of drug discovery. At NCATS, a single-point rat liver microsomal (RLM) stability assay is used as a Tier I assay, while human liver microsomal (HLM) stability is employed as a Tier II assay. We experimentally screened and collected data on over 30,000 compounds for RLM stability and over 7000 compounds for HLM stability. Although HLM stability screening provides valuable insights, the increasing number of hits generated, along with the time- and resource-intensive nature of the assay, highlights the need for alternative strategies. One promising approach is leveraging in silico models trained on these experimental datasets. Methods: We describe the development of an HLM stability prediction model using our in-house HLM stability dataset. Results: Employing both classical machine learning methods and advanced techniques, such as neural networks, we achieved model accuracies exceeding 80%. Moreover, we validated our model using external test sets and found that our models are comparable to some of the best models in literature. Additionally, the strong correlation observed between our RLM and HLM data was further reinforced by the fact that our HLM model performance improved when using RLM stability predictions as an input descriptor. Conclusions: The best model along with a subset of our dataset (PubChem AID: 1963597) has been made publicly accessible on the ADME@NCATS website for the benefit of the greater drug discovery community. To the best of our knowledge, it is the largest open-source model of its kind and the first to leverage cross-species data.
Although multi-omics integration relevant to patient outcome is typically characterized by an analyte interactome, current multi-omic integration methods either (1) model outcome without directly including associations between analytes, (2) model the interactome without directly evaluating the saliency of the model in the context of outcome, or (3) model outcome in a high-dimensional parameter space not suitable for small sample sizes (which are common in multi-omics studies). We introduce Graph Ensemble Neural Network (GENN), a methodology that learns the interactome most predictive of outcome in a low-dimensional parameter space built on complementary attributes for all possible analyte associations (metafeatures). We show that GENN is robust to noise in measurements using a theoretical model, outperforms the predictive performance of existing methods when evaluated on Tegafur drug response in NCI-60 cancer cell line data, and uncovers potentially novel multi-omic mechanisms driving total serum IgE levels in pediatric asthma and patient survival in glioblastomas. ### Competing Interest Statement STW receives royalties from UpToDate Inc. and is on the Board of Histolix Inc. JL-S is a scientific advisor to TruDiagnostic Inc, Precion Inc, and Ahara Corp and is on the Metabolomics Society Board.
Understanding the molecular underpinnings of disease severity and progression in human studies is necessary to develop metabolism-related preventative strategies for severe COVID-19. Metabolites and metabolic pathways that predispose individuals to severe disease are not well understood. In this study, we generated comprehensive plasma metabolomic profiles in >550 patients from the Longitudinal EMR and Omics COVID-19 Cohort. Samples were collected before (n = 441), during (n = 86), and after (n = 82) COVID-19 diagnosis, representing 555 distinct patients, most of which had single timepoints. Regression models adjusted for demographics, risk factors, and comorbidities, were used to determine metabolites associated with predisposition to and/or persistent effects of COVID-19 severity, and metabolite changes that were transient/lingering over the disease course. Sphingolipids/phospholipids were negatively associated with severity and exhibited lingering elevations after disease, while modified nucleotides were positively associated with severity and had lingering decreases after disease. Cytidine and uridine metabolites, which were positively and negatively associated with COVID-19 severity, respectively, were acutely elevated, reflecting the particular importance of pyrimidine metabolism in active COVID-19. This is the first large metabolomics study using COVID-19 plasma samples before, during, and/or after disease. Our results lay the groundwork for identifying putative biomarkers and preventive strategies for severe COVID-19.
Prebiotic galactooligosaccharides (GOS) reduce anxiety-like behaviors in mice and humans. However, the biological pathways behind these behavioral changes are not well understood. To begin to study these pathways, we utilized C57BL/6 mice that were fed a standard diet with or without GOS supplementation for 3 weeks prior to testing on the open field. After behavioral testing, colonic contents and serum were collected for bacteriome (16S rRNA gene sequencing, colonic contents only) and metabolome (UPLC-MS, colonic contents and serum data) analyses. As expected, GOS significantly reduced anxiety-like behavior (i.e., increased time in the center) and decreased cytokine gene expression ( Tnfa and Ccl2) ) in the prefrontal cortex. Notably, time in the center of the open field was significantly correlated with serum methyl-indole-3-acetic acid (methyl-IAA). This metabolite is a methylated form of indole-3-acetic acid (IAA) that is derived from bacterial metabolism of tryptophan. Sequencing analyses showed that GOS significantly increased Lachnospiraceae UCG006 and Akkermansia; ; these taxa are known to metabolize both GOS and tryptophan. To determine the extent to which methyl-IAA can affect anxiety-like behavior, mice were intraperitoneally injected with methyl-IAA. Mice given methyl-IAA had a reduction in anxiety-like behavior in the open field, along with lower Tnfa in the prefrontal cortex. Methyl-IAA was also found to reduce TNF-alpha (as well as CCL2) production by LPS-stimulated BV2 microglia. Together, these data support a novel pathway through which GOS reduces anxiety-like behaviors in mice and suggests that the bacterial metabolite methyl-IAA reduces microglial cytokine and chemokine production, which in turn reduces anxiety-like behavior.
Drug repurposing is a strategy for identifying new uses of approved or investigational drugs that are outside the scope of the original medical indication. Even though many repurposed drugs have been found serendipitously in the past, the increasing availability of large volumes of biomedical data has enabled more systemic, data-driven approaches for drug candidate identification. At National Center of Advancing Translational Sciences (NCATS), we invent new methods to generate new data and information publicly available to spur innovation and scientific discovery. In this study, we aimed to explore and demonstrate biomedical data generated and collected via two NCATS research programs, the Toxicology in the 21st Century program (Tox21) and the Biomedical Data Translator (Translator) for the application of drug repurposing. These two programs provide complementary types of biomedical data from uncovering underlying biological mechanisms with bioassay screening data from Tox21 for chemical clustering, to enrich clustered chemicals with scientific evidence mined from the Translator towards drug repurposing. 129 chemical clusters have been generated and three of them have been further investigated for drug repurposing candidate identification, which is detailed as case studies.
Drug development in rare diseases is challenging due to the limited availability of subjects with the diseases and recruiting from a small patient population. The high cost and low success rate of clinical trials motivate deliberate analysis of existing clinical trials to understand status of clinical development of orphan drugs and discover new insight for new trial. In this project, we aim to develop a user centered Rare disease based Clinical Trial Knowledge Graph (RCTKG) to integrate publicly available clinical trial data with rare diseases from the Genetic and Rare Disease (GARD) program in a semantic and standardized form for public use. To better serve and represent the interests of rare disease users, user stories were defined for three types of users, patients, healthcare providers and informaticians, to guide the RCTKG design in supporting the GARD program at NCATS/NIH and the broad clinical/research community in rare diseases.
Serum total immunoglobulin E levels (total IgE) capture the state of the immune system in relation to allergic sensitization. High levels are associated with airway obstruction and poor clinical outcomes in pediatric asthma. Inconsistent patient response to anti-IgE therapies motivates discovery of molecular mechanisms underlying serum IgE level differences in children with asthma. To uncover these mechanisms using complementary metabolomic and transcriptomic data, abundance levels of 529 named metabolites and expression levels of 22,772 genes were measured among children with asthma in the Childhood Asthma Management Program (CAMP, N=564) and the Genetic Epidemiology of Asthma in Costa Rica Study (GACRS, N=309) via the TOPMed initiative. Gene-metabolite associations dependent on IgE were identified within each cohort using multivariate linear models and were interpreted in a biochemical context using network topology, pathway and chemical enrichment, and representation within reactions. A total of 1,617 total IgE-dependent gene-metabolite associations from GACRS and 29,885 from CAMP met significance cutoffs. Of these, glycine and guanidinoacetic acid (GAA) were associated with the most genes in both cohorts, and the associations represented reactions central to glycine, serine, and threonine metabolism and arginine and proline metabolism. Pathway and chemical enrichment analysis further highlighted additional related pathways of interest. The results of this study suggest that GAA may modulate total IgE levels in two independent pediatric asthma cohorts with different characteristics, supporting the use of L-Arginine as a potential therapeutic for asthma exacerbation. Other potentially new targetable pathways are also uncovered.
OBJECTIVE:Identifying sets of rare diseases with shared aspects of etiology and pathophysiology may enable drug repurposing. Toward that aim, we utilized an integrative knowledge graph to construct clusters of rare diseases. MATERIALS AND METHODS:Data on 3242 rare diseases were extracted from the National Center for Advancing Translational Science Genetic and Rare Diseases Information center internal data resources. The rare disease data enriched with additional biomedical data, including gene and phenotype ontologies, biological pathway data, and small molecule-target activity data, to create a knowledge graph (KG). Node embeddings were trained and clustered. We validated the disease clusters through semantic similarity and feature enrichment analysis. RESULTS:Thirty-seven disease clusters were created with a mean size of 87 diseases. We validate the clusters quantitatively via semantic similarity based on the Orphanet Rare Disease Ontology. In addition, the clusters were analyzed for enrichment of associated genes, revealing that the enriched genes within clusters are highly related. DISCUSSION:We demonstrate that node embeddings are an effective method for clustering diseases within a heterogenous KG. Semantically similar diseases and relevant enriched genes have been uncovered within the clusters. Connections between disease clusters and drugs are enumerated for follow-up efforts. CONCLUSION:We lay out a method for clustering rare diseases using graph node embeddings. We develop an easy-to-maintain pipeline that can be updated when new data on rare diseases emerges. The embeddings themselves can be paired with other representation learning methods for other data types, such as drugs, to address other predictive modeling problems.
Abstract The Illuminating the Druggable Genome (IDG) project aims to improve our understanding of understudied proteins and our ability to study them in the context of disease biology by perturbing them with small molecules, biologics, or other therapeutic modalities. Two main products from the IDG effort are the Target Central Resource Database (TCRD) (http://juniper.health.unm.edu/tcrd/), which curates and aggregates information, and Pharos (https://pharos.nih.gov/), a web interface for users to extract and visualize data from TCRD. Since the 2021 release, TCRD/Pharos has focused on developing visualization and analysis tools that help reveal higher-level patterns in the underlying data. The current iterations of TCRD and Pharos enable users to perform enrichment calculations based on subsets of targets, diseases, or ligands and to create interactive heat maps and UpSet charts of many types of annotations. Using several examples, we show how to address disease biology and drug discovery questions through enrichment calculations and UpSet charts.
Chapter 8Full Access Pharos and TCRD : Informatics Tools for Illuminating Dark Targets Keith J. Kelleher, Keith J. Kelleher National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorTimothy K. Sheils, Timothy K. Sheils National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorStephen L. Mathias, Stephen L. Mathias University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USASearch for more papers by this authorDac-Trung Nguyen, Dac-Trung Nguyen National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorVishal Siramshetty, Vishal Siramshetty National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorAjay Pillai, Ajay Pillai National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorJeremy J. Yang, Jeremy J. Yang University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USASearch for more papers by this authorCristian G. Bologa, Cristian G. Bologa University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USASearch for more papers by this authorJeremy S. Edwards, Jeremy S. Edwards University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USASearch for more papers by this authorTudor I. Oprea, Tudor I. Oprea University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USA Expert Systems Inc., 12760 High Bluff Dr Ste 370, San Diego, 92130 CA, USASearch for more papers by this authorEwy Mathé, Ewy Mathé National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this author Keith J. Kelleher, Keith J. Kelleher National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorTimothy K. Sheils, Timothy K. Sheils National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorStephen L. Mathias, Stephen L. Mathias University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USASearch for more papers by this authorDac-Trung Nguyen, Dac-Trung Nguyen National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorVishal Siramshetty, Vishal Siramshetty National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorAjay Pillai, Ajay Pillai National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this authorJeremy J. Yang, Jeremy J. Yang University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USASearch for more papers by this authorCristian G. Bologa, Cristian G. Bologa University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USASearch for more papers by this authorJeremy S. Edwards, Jeremy S. Edwards University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USASearch for more papers by this authorTudor I. Oprea, Tudor I. Oprea University of New Mexico Health Sciences Center, Department of Internal Medicine, Translational Informatics Division, MSC09-5025, 1 University of New Mexico, Albuquerque, 87131 NM, USA Expert Systems Inc., 12760 High Bluff Dr Ste 370, San Diego, 92130 CA, USASearch for more papers by this authorEwy Mathé, Ewy Mathé National Center for Advancing Translational Science, 9800 Medical Center Drive, Rockville, 20850 MD, USASearch for more papers by this author Book Editor(s):Antoine Daina, Antoine Daina Swiss Institute of Bioinformatics, Unil, Quartier Sorge, Bâtiment Amphipôle, Lausanne, 1015 SwitzerlandSearch for more papers by this authorMichael Przewosny, Michael Przewosny Borngasse 43, Aachen, 52064 GermanySearch for more papers by this authorVincent Zoete, Vincent Zoete University of Lausanne, Route de la Corniche 9A, Epalinges, 1015 SwitzerlandSearch for more papers by this author First published: 20 October 2023 https://doi.org/10.1002/9783527830497.ch8Book Series:Methods and Principles in Medicinal Chemistry AboutPDFPDFView & download chapterView & download chapterView & download full book ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onEmailFacebookTwitterLinkedInRedditWechat Summary The Illuminating the Druggable Genome (IDG) program, a National Institutes of Health Common Fund Program, started in 2013 to improve our understanding of understudied protein families. Outcomes of this program include the Target Central Resource Database (TCRD), a relational database that aggregates and harmonizes a vast array of knowledge sources for drug target data, and Pharos, an interactive web-based tool that enables users to interact with and analyze TCRD data. This chapter provides an overview of the data sources that comprise TCRD and highlights various methods and specific use cases by which users can interact with the TCRD through Pharos. TCRD aggregates annotations and information on 339,220 ligands, 13,704 diseases, and 20,412 targets, which are collected and aggregated from 79 sources. One key feature of TCRD is the definition of four target development levels (TDLs): Tdark, Tclin, Tbio, and Tchem, which describe the level and types of information known for each target in the resource. Beyond displaying the rich TCRD data, Pharos produces interactive visualizations that enable users to view high-level trends in the data (heat maps) and explore common combinations of values in the dataset (UpSet plots). Users can also perform enrichment calculations for a population that can tell users which values are overrepresented, such as which associated diseases are overrepresented in a list of targets or which target classes are overrepresented in a list of active ligands. Users can search and browse data using various types of inputs, including but not limited to target sequences, biological pathways, ligand structures, names, common IDs of targets, drugs, or diseases. Within Pharos, different functionalities can be used to address real-world scientific problems, as shown in use-case scenarios. These include generating hypotheses for the role of a "dark" target (with virtually no known information) and predicting potential targets and target families for a novel chemical compound. References Edwards , A.M. , Isserlin , R. , Bader , G.D. et al. ( 2011 ). Too many roads not taken . Nature 163 – 165 . https://doi.org/10.1038/470163a . Oprea , T.I. , Bologa , C.G. , Brunak , S. et al. ( 2018 ). Unexplored therapeutic opportunities in the human genome . Nature Reviews. Drug Discovery 17 : 317 – 332 . Stoeger , T. , Gerlach , M. , Morimoto , R.I. , and Nunes Amaral , L.A. ( 2018 ). Large-scale investigation of the reasons why potentially important genes are ignored . PLoS Biology 16 : e2006643 . Illuminating the druggable genome . 9 Jul 2013 [cited 29 Aug 2022]. https://commonfund.nih.gov/idg . Lin , Y. , Mehta , S. , Küçük-McGinty , H. et al. ( 2017 ). Drug target ontology to classify and integrate drug discovery data . Journal of Biomedical Semantics 8 : 50 . Nguyen , D.-T. , Mathias , S. , Bologa , C. et al. ( 2017 ). Pharos: collating protein information to shed light on the druggable genome . Nucleic Acids Research 45 : D995 – D1002 . Pletscher-Frankild , S. , Pallejà , A. , Tsafou , K. et al. ( 2015 ). DISEASES: text mining and data integration of disease–gene associations . Methods 74 : 83 – 89 . Gaulton , A. , Hersey , A. , Nowotka , M. et al. ( 2017 ). The ChEMBL database in 2017 . Nucleic Acids Research D945 – D954 . https://doi.org/10.1093/nar/gkw1074 . Ursu , O. , Holmes , J. , Bologa , C.G. et al. ( 2019 ). DrugCentral 2018: an update . Nucleic Acids Research 47 : D963 – D970 . Sheils , T.K. , Mathias , S.L. , Kelleher , K.J. et al. ( 2021 ). TCRD and Pharos 2021: mining the human proteome for disease biology . Nucleic Acids Research 49 : D1334 – D1346 . Carvalho-Silva , D. , Pierleoni , A. , Pignatelli , M. et al. ( 2019 ). Open targets platform: new developments and updates two years on . Nucleic Acids Research 47 : D1056 – D1065 . Safran , M. , Rosen , N. , Twik , M. et al. ( 2021 ). The GeneCards suite . In: Practical Guide to Life Science Databases (ed. I. Abugessaisa and T. Kasukawa ), 27 – 56 . Singapore : Springer Nature Singapore . UniProt Consortium ( 2021 ). UniProt: the universal protein knowledgebase in 2021 . Nucleic Acids Research 49 : D480 – D489 . Davis , A.P. , Grondin , C.J. , Johnson , R.J. et al. ( 2021 ). Comparative Toxicogenomics Database (CTD): update 2021 . Nucleic Acids Research 49 : D1138 – D1143 . Piñero , J. , Ramírez-Anguita , J.M. , Saüch-Pitar ch , J. et al. ( 2020 ). The DisGeNET knowledge platform for disease genomics: 2019 update . Nucleic Acids Research 48 : D845 – D855 . Jia , J. , An , Z. , Ming , Y. et al. ( 2018 ). eRAM: encyclopedia of rare disease annotations for precision medicine . Nucleic Acids Research 46 : D937 – D943 . Papatheodorou , I. , Moreno , P. , Manning , J. et al. ( 2020 ). Expression Atlas update: from tissues to single cells . Nucleic Acids Research 48 : D77 – D83 . Mungall , C.J. , McMurry , J.A. , Köhler , S. et al. ( 2017 ). The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species . Nucleic Acids Research 45 : D712 – D722 . UniProt Consortium ( 2019 ). UniProt: a worldwide hub of protein knowledge . Nucleic Acids Research 47 : D506 – D515 . Nastou , K.C. , Nasi , G.I. , Tsiolaki , P.L. et al. ( 2019 ). AmyCo: the amyloidoses collection . Amyloid 26 : 112 – 117 . Schriml , L.M. , Mitraka , E. , Munro , J. et al. ( 2018 ). Update: classification, content and workflow expansion . Nucleic Acids Research 2019 : D955 – D962 . https://doi.org/10.1093/nar/gky1032 . McKusick , V.A. ( 1998 ). Mendelian Inheritance in Man: A Catalog of Human Genes and Genetic Disorders . JHU Press . Bodenreider , O. ( 2004 ). The Unified Medical Language System (UMLS): integrating biomedical terminology . Nucleic Acids Research 32 : D267 – D270 . Medical subject headings – home page. 2020 [cited 26 Aug 2022]. https://www.nlm.nih.gov/mesh/meshhome.html . Weinreich , S.S. , Mangon , R. , Sikkens , J.J. et al. ( 2008 ). Orphanet: a European database for rare diseases . Nederlands Tijdschrift voor Geneeskunde 152 : 518 – 519 . Harding , S.D. , Armstrong , J.F. , Faccenda , E. et al. The IUPHAR/BPS guide to PHARMACOLOGY in 2022: curating pharmacology for COVID-19, malaria and antibacterials . Nucleic Acids Research 2022 : D1282 – D1294 . https://doi.org/10.1093/nar/gkab1010 . Bajusz , D. , Rácz , A. , and Héberger , K. ( 2015 ). Why is Tanimoto index an appropriate choice for fingerprint-based similarity calculations? Journal of Cheminformatics 7 : 20 . Altschul , S.F. , Gish , W. , Miller , W. et al. ( 1990 ). Basic local alignment search tool . Journal of Molecular Biology 215 : 403 – 410 . Levandowsky , M. and Winter , D. ( 1971 ). Distance between sets . Nature 234 : 34 – 35 . Zakharov , A.V. , Zhao , T. , Nguyen , D.-T. et al. ( 2019 ). Novel consensus architecture to improve performance of large-scale multitask deep learning QSAR models . Journal of Chemical Information and Modeling 59 : 4613 – 4624 . Fisher , R.A. ( 1992 ). Statistical methods for research workers . In: Breakthroughs in Statistics: Methodology and Distribution (ed. S. Kotz and N.L. Johnson ), 66 – 70 . New York, NY : Springer New York . Yekutieli , D. and Benjamini , Y. ( 1999 ). Resampling-based false discovery rate controlling multiple test procedures for correlated test statistics . Journal of Statistical Planning and Inference 82 : 171 – 196 . Rouillard , A.D. , Gundersen , G.W. , Fern andez , N.F. et al. ( 2016 ). The harmonizome: a collection of processed datasets gathered to serve and mine knowledge about genes and proteins . Database 2016 : https://doi.org/10.1093/database/baw100 . Thomas , P.D. , Campbell , M.J. , Kejariwal , A. et al. ( 2003 ). PANTHER: a library of protein families and subfamilies indexed by function . Genome Research 13 : 2129 – 2141 . Carithers , L.J. , Ardlie , K. , Barcus , M. et al. ( 2015 ). A novel approach to high-quality postmortem tissue procurement: the GTEx project . Biopreservation and Biobanking 13 : 311 – 319 . Thul , P.J. and Lindskog , C. ( 2018 ). The human protein atlas: a spatial map of the human proteome . Protein Science 27 : 233 – 244 . Kim , M.-S. , Pinto , S.M. , Getnet , D. et al. ( 2014 ). A draft map of the human proteome . Nature 509 : 575 – 581 . Palasca , O. , Santos , A. , Stolte , C. et al. ( 2018 ). TISSUES 2.0: an integrative web resource on mammalian tissue expression . Database 2018 : https://doi.org/10.1093/database/bay003 . Mungall , C.J. , Torniai , C. , Gkoutos , G.V. et al. ( 2012 ). Uberon, an integrative multi-species anatomy ontology . Genome Biology 13 : R5 . Watkins X, Garcia LJ, Pundir S, Martin MJ, UniProt Consortium ( 2017 ). ProtVista: visualization of protein sequence annotations . Bioinformatics 33 : 2040 – 2041 . Berman , H.M. , Westbrook , J. , Feng , Z. et al. ( 2000 ). The Protein Data Bank . Nucleic Acids Research 28 : 235 – 242 . Jumper , J. , Evans , R. , Pritzel , A. et al. ( 2021 ). Highly accurate protein structure prediction with AlphaFold . Nature 596 : 583 – 589 . Varadi , M. , Anyango , S. , Deshpande , M. et al. ( 2022 ). AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models . Nucleic Acids Research 50 : D439 – D444 . Szklarczyk , D. , Gable , A.L. , Nastou , K.C. et al. ( 2020 ). The STRING database in 2021: customizable protein–protein networks, and functional characterization of user-uploaded gene/measurement sets . Nucleic Acids Research 49 : D605 – D612 . Fabregat , A. , Jupe , S. , Matthews , L. et al. ( 2018 ). The Reactome Pathway Knowledgebase . Nucleic Acids Research 46 : D649 – D655 . Huttlin , E.L. , Bruckner , R.J. , Navarrete-Perea , J. et al. ( 2021 ). Dual proteome-scale networks reveal cell-specific remodeling of the human interactome . Cell 184 : 3022 – 3040.e28 . Kanehisa , M. and Goto , S. ( 2000 ). KEGG: kyoto encyclopedia of genes and genomes . Nucleic Acids Research 28 : 27 – 30 . Cerami , E.G. , Gross , B.E. , Demir , E. et al. ( 2011 ). Pathway Commons, a web resource for biological pathway data . Nucleic Acids Research 39 : D685 – D690 . Martens , M. , Ammar , A. , Riutta , A. et al. ( 2021 ). WikiPathways: connecting communities . Nucleic Acids Research 49 : D613 – D621 . Ashburner , M. , Ball , C.A. , Blake , J.A. et al. ( 2000 ). Gene ontology: tool for the unification of biology . The Gene Ontology Consortium. Nature Genetics 25 : 25 – 29 . Gene Ontology Consortium ( 2021 ). The Ge ne Ontology resource: enriching a GOld mine . Nucleic Acids Research 49 : D325 – D334 . Cannon , D.C. , Yang , J.J. , Mathias , S.L. et al. ( 2017 ). TIN-X: target importance and novelty explorer . Bioinformatics 33 : 2601 – 2603 . Buniello , A. , MacArthur , J.A.L. , Cerezo , M. et al. ( 2019 ). The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019 . Nucleic Acids Research 47 : D1005 – D1012 . Yang , J.J. , Grissa , D. , Lambert , C.G. et al. ( 2021 ). TIGA: target illumination GWAS analytics . Bioinformatics , https://doi.org/10.1093/bioinformatics/btab427 . Wei , C.-H. , Kao , H.-Y. , and Lu , Z. ( 2013 ). PubTator: a web-based text mining tool for assisting biocuration . Nucleic Acids Research 41 : W518 – W522 . Lex , A. , Gehlenborg , N. , Strobelt , H. et al. ( 2014 ). UpSet: visualization of intersecting sets . IEEE Transactions on Visualization and Computer Graphics 20 : 1983 – 1992 . Tawa , G.J. , Braisted , J. , Gerhold , D. et al. ( 2021 ). Transcriptomic profiling in canines and humans reveals cancer specific gene modules and biological mechanisms common to both species . PLoS Computational Biology 17 : e1009450 . Marazzi , L. , Shah , M. , Balakrishnan , S. et al. ( 2022 ). NETISCE: a network-based tool for cell fate reprogramming . npj Systems Biology and Applications 8 ( 1 ): 21 . https://doi.org/10.1038/s41540-022-00231-y . Korrapati , S. , Taukulis , I. , Olszewski , R. et al. ( 2019 ). Single cell and single nucleus RNA-Seq reveal cellular heterogeneity and homeostatic regulatory networks in adult mouse stria vascularis . Frontiers in Molecular Neuroscience 12 : 316 . https://doi.org/10.3389/fnmol.2019.00316. PMID: 31920542; PMCID: PMC6933021. Federico , A. , Pavel , A. , Moebus , L. et al. ( 2021 ). The integration of large-scale public data and network analysis uncovers molecular characteristics of . bioRxiv https://doi.org/10.1101/2021.05.10.443441 . Gosal , G. , Kochut , K.J. , and Kannan , N. ( 2011 ). ProKinO: an ontology for integrative analysis of protein kinases in cancer . PLoS One 6 : e28782 . Berginski , M.E. , Moret , N. , Liu , C. et al. ( 2021 ). The Dark Kinase Knowledgebase: an online compendium of knowledge and experimental results of understudied kinases . Nucleic Acids Research 49 : D529 – D535 . Ravanmehr , V. , Blau , H. , Cappelletti , L. et al. ( 2021 ). Supervised learning with word embeddings derived from PubMed captures latent knowledge about protein kinases and cancer . NAR Genomics and Bioinformatics 3 : lqab113 . Keiser , M.J. , Roth , B.L. , Armbruster , B.N. et al. ( 2007 ). Relating protein pharmacology by ligand chemistry . Nature Biotechnology 25 : 197 – 206 . Open Access Databases and Datasets for Drug Discovery ReferencesRelatedInformation
The computational metabolomics field brings together computer scientists, bioinformaticians, chemists, clinicians, and biologists to maximize the impact of metabolomics across a wide array of scientific and medical disciplines. The field continues to expand as modern instrumentation produces datasets with increasing complexity, resolution, and sensitivity. These datasets must be processed, annotated, modeled, and interpreted to enable biological insight. Techniques for visualization, integration (within or between omics), and interpretation of metabolomics data have evolved along with innovation in the databases and knowledge resources required to aid understanding. In this review, we highlight recent advances in the field and reflect on opportunities and innovations in response to the most pressing challenges. This review was compiled from discussions from the 2022 Dagstuhl seminar entitled “Computational Metabolomics: From Spectra to Knowledge”.
Supplementary Data from Association of Inflammation-Related and microRNA Gene Expression with Cancer-Specific Mortality of Colon Adenocarcinoma
BACKGROUND:Glioblastoma (GBM) is the most aggressive and common malignant primary brain tumor; however, treatment remains a significant challenge. This study aims to identify drug repurposing or repositioning candidates for GBM by developing an integrative rare disease profile network containing heterogeneous types of biomedical data. METHODS:We developed a Glioblastoma-based Biomedical Profile Network (GBPN) by extracting and integrating biomedical information pertinent to GBM-related diseases from the NCATS GARD Knowledge Graph (NGKG). We further clustered the GBPN based on modularity classes which resulted in multiple focused subgraphs, named mc_GBPN. We then identified high-influence nodes by performing network analysis over the mc_GBPN and validated those nodes that could be potential drug repurposing or repositioning candidates for GBM. RESULTS:We developed the GBPN with 1,466 nodes and 107,423 edges and consequently the mc_GBPN with forty-one modularity classes. A list of the ten most influential nodes were identified from the mc_GBPN. These notably include Riluzole, stem cell therapy, cannabidiol, and VK-0214, with proven evidence for treating GBM. CONCLUSION:Our GBM-targeted network analysis allowed us to effectively identify potential candidates for drug repurposing or repositioning. Further validation will be conducted by using other different types of biomedical and clinical data and biological experiments. The findings could lead to less invasive treatments for glioblastoma while significantly reducing research costs by shortening the drug development timeline. Furthermore, this workflow can be extended to other disease areas.