Serum total immunoglobulin E levels (total IgE) capture the state of the immune system in relation to allergic sensitization. High levels are associated with airway obstruction and poor clinical outcomes in pediatric asthma. Inconsistent patient response to anti-IgE therapies motivates discovery of molecular mechanisms underlying serum IgE level differences in children with asthma. To uncover these mechanisms using complementary metabolomic and transcriptomic data, abundance levels of 529 named metabolites and expression levels of 22,772 genes were measured among children with asthma in the Childhood Asthma Management Program (CAMP, N=564) and the Genetic Epidemiology of Asthma in Costa Rica Study (GACRS, N=309) via the TOPMed initiative. Gene-metabolite associations dependent on IgE were identified within each cohort using multivariate linear models and were interpreted in a biochemical context using network topology, pathway and chemical enrichment, and representation within reactions. A total of 1,617 total IgE-dependent gene-metabolite associations from GACRS and 29,885 from CAMP met significance cutoffs. Of these, glycine and guanidinoacetic acid (GAA) were associated with the most genes in both cohorts, and the associations represented reactions central to glycine, serine, and threonine metabolism and arginine and proline metabolism. Pathway and chemical enrichment analysis further highlighted additional related pathways of interest. The results of this study suggest that GAA may modulate total IgE levels in two independent pediatric asthma cohorts with different characteristics, supporting the use of L-Arginine as a potential therapeutic for asthma exacerbation. Other potentially new targetable pathways are also uncovered.
Motivation: Functional interpretation of high-throughput metabolomic and transcriptomic results is a crucial step in generating insight from experimental data. However, pathway and functional information for genes and metabolites are distributed among many siloed resources, limiting the scope of analyses that rely on a single knowledge source. Results: RaMP-DB 2.0 is a web interface, relational database, API and R package designed for straightforward and comprehensive functional interpretation of metabolomic and multi-omic data. RaMP-DB 2.0 has been upgraded with an expanded breadth and depth of functional and chemical annotations (ClassyFire, LIPID MAPS, SMILES, InChIs, etc.), with new data types related to metabolites and lipids incorporated. To streamline entity resolution across multiple source databases, we have implemented a new semi-automated process, thereby lessening the burden of harmonization and supporting more frequent updates. The associated RaMP-DB 2.0 R package now supports queries on pathways, common reactions (e.g. metabolite-enzyme relationship), chemical functional ontologies, chemical classes and chemical structures, as well as enrichment analyses on pathways (multi-omic) and chemical classes. Lastly, the RaMP-DB web interface has been completely redesigned using the Angular framework.
Motivation:IntLIM uncovers phenotype-dependent linear associations between two types of analytes (e.g. genes and metabolites) in a multi-omic dataset, which may reflect chemically or biologically relevant relationships.Results:The new IntLIM R package includes newly added support for generalized data types, covariate correction, continuous phenotypic measurements, model validation and unit testing. IntLIM analysis uncovered biologically relevant gene-metabolite associations in two separate datasets, and the run time is improved over baseline R functions by multiple orders of magnitude.Availability and implementation:IntLIM is available as an R package with a detailed vignette (https://github.com/ncats/IntLIM) and as an R Shiny app (see Supplementary Figs S1-S6) (https://intlim.ncats.io/).Supplementary information:Supplementary data are available at Bioinformatics Advances online.
ABSTRACTRaMP-DB 2.0 is a web interface, API, relational database and R package designed for straightforward and comprehensive functional interpretation of metabolomic and multi-omic data. Since its first release in 2018, RaMP-DB 2.0 has been upgraded with an expanded breadth and depth of functional and chemical annotation. Content from the source databases (Reactome, HMDB, and Wikipathways) has been updated, and new data types related to metabolite annotations have been incorporated. Structural information incorporated in RaMP-DB 2.0 includes SMILES strings, InChIs, InChIKeys. Chemical classes have been sourced from ClassyFire and LIPID MAPS. Accordingly, the RaMP-DB 2.0 R package has been updated and supports queries on pathways, common reactions, ontologies, chemical classes, and chemical structures. Additionally, RaMP-DB 2.0 now supports enrichment analyses on pathways and chemical classes. Our process for integrating annotations across resources has also been upgraded to lessen the burden of harmonization, thereby supporting more frequent updates. The code used to build all components of RaMP-DB 2.0 is freely available on GitHub athttps://github.com/ncats/ramp-dbandhttps://github.com/ncats/RaMP-Backend.
Adult-type diffuse gliomas have been classified according to histopathological characteristics only. Based on molecular profiles, the World Health Organization (WHO) classification defines three distinct biologic and prognostic types of diffuse gliomas: (i) IDH wild type (IDHwt), (ii) IDH mutant, 1p19q intact (IDHmut-non-codel), and (iii) IDH mutant, 1p19q codeleted (IDHmut-codel). Gliomas with missing molecular information are classified as “Not otherwise specified” (NOS). A histopathological signature that distinguishes IDH status in gliomas using hematoxylin & eosin (H&E)-stained whole slide images (WSIs) has yet to be developed to be used as proxy for molecular alterations. We employed WSIs with corresponding molecular data from The Cancer Genome Atlas (TCGA) and reclassified the gliomas according to the 2021 WHO guidelines. Weakly supervised deep learning approaches, i.e. the Uncertainty-Aware CNN (UA-CNN), and accompanying workflows developed by our group were successful in the segmentation and classification of other cancers. We used the weak labels of IDHwt or IDHmut for given WSIs and utilized an informative sampling algorithm to identify the most relevant tiles that are predictive of their WSI-level label. Next, a second UA-CNN was trained on the relevant subset of tiles determined from the screening stage. For inferencing, given a WSI, the trained UA-CNN model yielded both a predicted label and uncertainty measure for each tile. When reassembled into a WSI, an aggregated uncertainty classification map is formed, which will allow neuropathologists to more rapidly identify regions of interest. While molecular testing is indispensable for current glioma classification, our tool may help reduce the number of NOS diagnosis, mainly in smaller services, by identifying image features that can differentiate between IDHwt and IDHmut. We expect to develop and deploy a method to increase the amount of information extracted from the H&E stain, and help prioritize downstream molecular analyses mandatory for the layered diagnosis of gliomas.
As researchers are increasingly able to collect data on a large scale from multiple clinical and omics modalities, multi-omics integration is becoming a critical component of metabolomics research. This introduces a need for increased understanding by the metabolomics researcher of computational and statistical analysis methods relevant to multi-omics studies. In this review, we discuss common types of analyses performed in multi-omics studies and the computational and statistical methods that can be used for each type of analysis. We pinpoint the caveats and considerations for analysis methods, including required parameters, sample size and data distribution requirements, sources of a priori knowledge, and techniques for the evaluation of model accuracy. Finally, for the types of analyses discussed, we provide examples of the applications of corresponding methods to clinical and basic research. We intend that our review may be used as a guide for metabolomics researchers to choose effective techniques for multi-omics analyses relevant to their field of study.
The National Center for Advancing Translational Sciences (NCATS) has developed an online open science data portal for its COVID-19 drug repurposing campaign - named OpenData - with the goal of making data across a range of SARS-CoV-2 related assays available in real-time. The assays developed cover a wide spectrum of the SARS-CoV-2 life cycle, including both viral and human (host) targets. In total, over 10,000 compounds are being tested in full concentration-response ranges from across multiple annotated small molecule libraries, including approved drug, repurposing candidates and experimental therapeutics designed to modulate a wide range of cellular targets. The goal is to support research scientists, clinical investigators and public health officials through open data sharing and analysis tools to expedite the development of SARS-CoV-2 interventions, and to prioritize promising compounds and repurposed drugs for further development in treating COVID-19.
Background: Proteomic measurements, which closely reflect phenotypes, provide insights into gene expression regulations and mechanisms underlying altered phenotypes. Further, integration of data on proteome and transcriptome levels can validate gene signatures associated with a phenotype. However, proteomic data is not as abundant as genomic data, and it is thus beneficial to use genomic features to predict protein abundances when matching proteomic samples or measurements within samples are lacking. Results: We evaluate and compare four data-driven models for prediction of proteomic data from mRNA measured in breast and ovarian cancers using the 2017 DREAM Proteogenomics Challenge data. Our results show that Bayesian network, random forests, LASSO, and fuzzy logic approaches can predict protein abundance levels with median ground truth-predicted correlation values between 0.2 and 0.5. However, the most accurately predicted proteins differ considerably between approaches. Conclusions: In addition to benchmarking aforementioned machine learning approaches for predicting protein levels from transcript levels, we discuss challenges and potential solutions in state-of-the-art proteogenomic analyses.
Down Syndrome is a common disorder which causes intellectual disability among other symptoms. To date, no treatment exists for the learning difficulties associated with Down Syndrome. However, the pharmaceutical drug memantine has been shown to improve learning ability in a Down Syndrome model of mice (Ts65Dn) exposed to Context Fear Conditioning (CFC), an existing technique used in determining the extent of learning capability of mice. While the effect of memantine on learning capability in Ts65Dn mice is significant, the biological mechanism responsible for restoration of learning capability by memantine is poorly understood. One possible way to characterize this mechanism is by analyzing the neural protein profile data of normal and Down Syndrome mice with and without memantine treatment. In this work, we use a series of linear support vector machines to model the differential expression of 77 proteins obtained from the nuclear cortex of normal and Ts65Dn mice, with and without memantine treatment and with and without CFC stimulation. We use feature selection by weight threshold to select those proteins which play a significant role in characterizing each model. Per our findings, these subsets of proteins can be used to build more accurate classification models of the data than those subsets chosen using unsupervised learning or statistical analyses in previous studies. We recommend that the subsets of proteins selected using our proposed method be utilized in further biological study aiming to understand the effects of memantine on learning restoration.