An improved understanding of the human lung necessitates advanced systems models informed by an ever-increasing repertoire of molecular omics, cellular imaging, and pathological datasets. To centralize and standardize information across broad lung research efforts, we expanded the LungMAP.net website into a new gateway portal. This portal connects a broad spectrum of research networks, bulk and single-cell multiomics data, and a diverse collection of image data that span mammalian lung development and disease. The data are standardized across species and technologies using harmonized data and metadata models that leverage recent advances, including those from the Human Cell Atlas, diverse ontologies, and the LungMAP CellCards initiative. To cultivate future discoveries, we have aggregated a diverse collection of single-cell atlases for multiple species (human, rhesus, and mouse) to enable consistent queries across technologies, cohorts, age, disease, and drug treatment. These atlases are provided as independent and integrated queryable datasets, with an emphasis on dynamic visualization, figure generation, reanalysis, cell-type curation, and automated reference-based classification of user-provided single-cell genomics datasets (Azimuth). As this resource grows, we intend to increase the breadth of available interactive interfaces, supported data types, data portals and datasets from LungMAP, and external research efforts.
Numerous studies have provided single-cell transcriptome profiles of host responses to SARS-CoV-2 infection. Critically lacking however is a datamine that allows users to compare and explore cell profiles to gain insights and develop new hypotheses. To accomplish this, we harmonized datasets from COVID-19 and other control condition blood, bronchoalveolar lavage, and tissue samples, and derived a compendium of gene signature modules per cell type, subtype, clinical condition, and compartment. We demonstrate approaches to probe these via a new interactive web portal (http://toppcell.cchmc.org/ COVID-19). As examples, we develop three hypotheses: (1) a multicellular signaling cascade among alternatively differentiated monocyte-derived macrophages whose tasks include T cell recruitment and activation; (2) novel platelet subtypes with drastically modulated expression of genes responsible for adhesion, coagulation and thrombosis; and (3) a multilineage cell activator network able to drive extrafollicular B maturation via an ensemble of genes strongly associated with risk for developing post-viral autoimmunity.
Heterozygous NRXN1 deletions constitute the most prevalent currently known single-gene mutation associated with schizophrenia, and additionally predispose to multiple other neurodevelopmental disorders. Engineered heterozygous NRXN1 deletions impaired neurotransmitter release in human neurons, suggesting a synaptic pathophysiological mechanism. Utilizing this observation for drug discovery, however, requires confidence in its robustness and validity. Here, we describe a multicenter effort to test the generality of this pivotal observation, using independent analyses at two laboratories of patient-derived and newly engineered human neurons with heterozygous NRXN1 deletions. Using neurons transdifferentiated from induced pluripotent stem cells that were derived from schizophrenia patients carrying heterozygous NRXN1 deletions, we observed the same synaptic impairment as in engineered NRXN1-deficient neurons. This impairment manifested as a large decrease in spontaneous synaptic events, in evoked synaptic responses, and in synaptic paired-pulse depression. Nrxn1-deficient mouse neurons generated from embryonic stem cells by the same method as human neurons did not exhibit impaired neurotransmitter release, suggesting a human-specific phenotype. Human NRXN1 deletions produced a reproducible increase in the levels of CASK, an intracellular NRXN1-binding protein, and were associated with characteristic gene-expression changes. Thus, heterozygous NRXN1 deletions robustly impair synaptic function in human neurons regardless of genetic background, enabling future drug discovery efforts.
The human lung plays vital roles in respiration, host defense, and basic physiology. Recent technological advancements such as single-cell RNA sequencing and genetic lineage tracing have revealed novel cell types and enriched functional properties of existing cell types in lung. The time has come to take a new census. Initiated by members of the NHLBI-funded LungMAP Consortium and aided by experts in the lung biology community, we synthesized current data into a comprehensive and practical cellular census of the lung. Identities of cell types in the normal lung are captured in individual cell cards with delineation of function, markers, developmental lineages, heterogeneity, regenerative potential, disease links, and key experimental tools. This publication will serve as the starting point of a live, up-to-date guide for lung research at https://www.lungmap.net/cell-cards/. We hope that Lung CellCards will promote the community-wide effort to establish, maintain, and restore respiratory health.
Using a Systems Biology approach, we integrated genomic, transcriptomic, proteomic, and molecular structure information to provide a holistic understanding of the COVID-19 pandemic. The expression data analysis of the Renin Angiotensin System indicates mild nasal, oral or throat infections are likely and that the gastrointestinal tissues are a common primary target of SARS-CoV-2. Extreme symptoms in the lower respiratory system likely result from a secondary-infection possibly by a comorbidity-driven upregulation of ACE2 in the lung. The remarkable differences in expression of other RAS elements, the elimination of macrophages and the activation of cytokines in COVID-19 bronchoalveolar samples suggest that a functional immune deficiency is a critical outcome of COVID-19. We posit that using a non-respiratory system as a major pathway of infection is likely determining the unprecedented global spread of this coronavirus. One Sentence Summary A Systems Approach Indicates Non-respiratory Pathways of Infection as Key for the COVID-19 Pandemic
Background: The magnitude and severity of the COVID-19 pandemic cannot be overstated. Although the mortality rate is less than SARS and MERS, the global outbreak has already resulted in orders of magnitude more deaths. In order to tackle the complexities of this disease, a Systems Biology approach can provide insights into the biology of the virus and mechanisms of disease. Methods: Using a Systems Biology approach, we have integrated genomic, transcriptomic, proteomic, and molecular evolution data layers to understand its impact on host cells. We overlay these analyses with high-resolution structural models and atomistic molecular dynamics simulations conducted on the Summit supercomputer at the Oak Ridge National Laboratory. Findings: Transcriptomic and proteomic data indicate little to no expression of ACE2 in lung tissue. Molecular modeling simulations support ACE2 as the receptor for SARS-CoV-2, but ACE may also act as a receptor for the virus and may be important for entry of SARS-CoV-1. Gene expression data from bronchoalveolar lavage samples from COVID-19 patients identify upregulation of renin, angiotensin, and the angiotensin 1-7 receptor MAS as well as a cellular landscape consistent with large-scale dissolution of lung parenchyma tissues, likely comprised of all lung epithelial cell types as well as lymphatic endothelial cells, but an absence of cells, such as macrophages, normally essential for host defense. Interpretation: Our analyses indicate that the commonly accepted view that SARS-CoV-2 enters host cells via ACE2 expressed in the lung is unlikely because ACE2 is undetectable there. Instead, given the greater target space of ACE2-positive nasal, oral, and gastrointestinal tissues, a more likely scenario suggests initial infection in those tissues is followed by a secondary infection via migration through the lymphatic system and bloodstream to the lung microvasculature. The elimination of macrophages and complete lack of activated cytokine signature in COVID-19 lung samples suggest that a major component of SARS-CoV-29s virulence is its net effect of causing a functional immune deficiency syndrome. Our structural analysis of the SARS-CoV-2 proteome suggests involvement of the highly conserved nsp5 protein as part of a major mechanism that suppresses the nuclear factor transcription factor kappa B (NF-κB) pathway, eliminating the host cell9s interferon-based antiviral response.
Identifying functionally significant microRNAs (miRs) and their correspondingly most important messenger RNA targets (mRNAs) in specific biological contexts is a critical task to improve our understanding of molecular mechanisms underlying organismal development, physiology and disease. However, current miR-mRNA target prediction platforms rank miR targets based on estimated strength of physical interactions and lack the ability to rank interactants as a function of their potential to impact a given biological system. To address this, we have developed ToppMiR (http://toppmir.cchmc.org), a web-based analytical workbench that allows miRs and mRNAs to be co-analyzed via biologically centered approaches in which gene function associated annotations are used to train a machine learning-based analysis engine. ToppMiR learns about biological contexts based on gene associated information from expression data or from a user-specified set of genes that relate to context-relevant knowledge or hypotheses. Within the biological framework established by the genes in the training set, its associated information content is then used to calculate a features association matrix composed of biological functions, protein interactions and other features. This scoring matrix is then used to jointly rank both the test/candidate miRs and mRNAs. Results of these analyses are provided as downloadable tables or network file formats usable in Cytoscape.
The ongoing expansion of our knowledge about the molecular basis of structure and function of diverse biological entities and processes presents us with both huge opportunities and challenges to improve our understanding of disease. To accomplish this, it is critical that we assemble and use all areas of knowledge and quantitative data in the analysis of normal and disease-affected samples and individuals. The field of bioinformatics and computational biology seeks to build data and knowledge on mining resources, tools, and analysis approaches that leverage the spectrum of biological facts, observations, and measurements, believing that by doing this, we can improve our understanding within all individual areas of pathobiology. An additional hope is both to gain our ability to identify and understand the specific causes, modifiers, and consequences of disease and to develop approaches to reverse engineer these in such a way as to recognize best treatments and preventions. Because the kinds of organized data, databases, and analysis tools now available to facilitate deeper understanding of disease biology may be unfamiliar to biologists, a primary goal of this article is to illustrate how to connect the logic and capabilities of data-driven bioinformatics applications and to mine and combine data and knowledge in ways that can lead to novel testable hypotheses. As part of this, we also present an overview of computational approaches used to prioritize candidate causal gene variations and mutations that can lead to significant chronic disease and how to learn new things from what we know thus far. To further elucidate the power of these computational approaches, through a case study of fat malabsorption, we demonstrate that heterogeneous but related data from diverse resources can be integrated and presented in an intuitive manner that in turn aids in further the understanding of the molecular basis of disease further.
Although a number of computational approaches have been developed to integrate data from multiple sources for the purpose of predicting or prioritizing candidate disease genes, relatively few of them focus on identifying or ranking drug targets. To address this deficit, we have developed an approach to specifically identify and prioritize disease and drug candidate genes. In this chapter, we demonstrate the applicability of integrative systems-biology-based approaches to identify potential drug targets and candidate genes by employing information extracted from public databases. We illustrate the method in detail using examples of two neurodegenerative diseases (Alzheimer's and Parkinson's) and one neuropsychiatric disease (Schizophrenia).
ToppCluster is a web server application that leverages a powerful enrichment analysis and underlying data environment for comparative analyses of multiple gene lists. It generates heatmaps or connectivity networks that reveal functional features shared or specific to multiple gene lists. ToppCluster uses hypergeometric tests to obtain list-specific feature enrichment P-values for currently 17 categories of annotations of human-ortholog genes, and provides user-selectable cutoffs and multiple testing correction methods to control false discovery. Each nameable gene list represents a column input to a resulting matrix whose rows are overrepresented features, and individual cells per-list P-values and corresponding genes per feature. ToppCluster provides users with choices of tabular outputs, hierarchical clustering and heatmap generation, or the ability to interactively select features from the functional enrichment matrix to be transformed into XGMML or GEXF network format documents for use in Cytoscape or Gephi applications, respectively. Here, as example, we demonstrate the ability of ToppCluster to enable identification of list-specific phenotypic and regulatory element features (both cis-elements and 3′UTR microRNA binding sites) among tissue-specific gene lists. ToppCluster’s functionalities enable the identification of specialized biological functions and regulatory networks and systems biology-based dissection of biological states. ToppCluster can be accessed freely at http://toppcluster.cchmc.org.
ToppGene Suite (http://toppgene.cchmc.org; this web site is free and open to all users and does not require a login to access) is a one-stop portal for (i) gene list functional enrichment, (ii) candidate gene prioritization using either functional annotations or network analysis and (iii) identification and prioritization of novel disease candidate genes in the interactome. Functional annotation-based disease candidate gene prioritization uses a fuzzy-based similarity measure to compute the similarity between any two genes based on semantic annotations. The similarity scores from individual features are combined into an overall score using statistical meta-analysis. A P-value of each annotation of a test gene is derived by random sampling of the whole genome. The protein–protein interaction network (PPIN)-based disease candidate gene prioritization uses social and Web networks analysis algorithms (extended versions of the PageRank and HITS algorithms, and the K-Step Markov method). We demonstrate the utility of ToppGene Suite using 20 recently reported GWAS-based gene–disease associations (including novel disease genes) representing five diseases. ToppGene ranked 19 of 20 (95%) candidate genes within the top 20%, while ToppNet ranked 12 of 16 (75%) candidate genes among the top 20%.