The heart, which is the first organ to develop, is highly dependent on its form to function1,2. However, how diverse cardiac cell types spatially coordinate to create the complex morphological structures that are crucial for heart function remains unclear. Here we integrated single-cell RNA-sequencing with high-resolution multiplexed error-robust fluorescence in situ hybridization to resolve the identity of the cardiac cell types that develop the human heart. This approach also provided a spatial mapping of individual cells that enables illumination of their organization into cellular communities that form distinct cardiac structures. We discovered that many of these cardiac cell types further specified into subpopulations exclusive to specific communities, which support their specialization according to the cellular ecosystem and anatomical region. In particular, ventricular cardiomyocyte subpopulations displayed an unexpected complex laminar organization across the ventricular wall and formed, with other cell subpopulations, several cellular communities. Interrogating cell-cell interactions within these communities using in vivo conditional genetic mouse models and in vitro human pluripotent stem cell systems revealed multicellular signalling pathways that orchestrate the spatial organization of cardiac cell subpopulations during ventricular wall morphogenesis. These detailed findings into the cellular social interactions and specialization of cardiac cell types constructing and remodelling the human heart offer new insights into structural heart diseases and the engineering of complex multicellular tissues for human heart repair.
Many disease-causing genetic variants converge on common biological functions and pathways. Precisely how to incorporate pathway knowledge in genetic association studies is not yet clear, however. Previous approaches employ a two-step approach, in which a regular association test is first performed to identify variants associated with the disease phenotype, followed by a test for functional enrichment within the genes implicated by those variants. Here we introduce a concise one-step approach, Hierarchical Genetic Analysis (Higana), which directly computes phenotype associations against each function in the large hierarchy of biological functions documented by the Gene Ontology. Using this approach, we identify risk genes and functions for Chronic Obstructive Pulmonary Disease (COPD), highlighting microtubule transport, muscle adaptation, and nicotine receptor signaling pathways. Microtubule transport has not been previously linked to COPD, as it integrates genetic variants spread over numerous genes. All associations validate strongly in a second COPD cohort.
Circadian disruptions impact nearly all people with Alzheimer's disease (AD), emphasizing both their potential role in pathology and the critical need to investigate the therapeutic potential of circadian-modulating interventions. Here, we show that time-restricted feeding (TRF) without caloric restriction improved key disease components including behavioral timing, disease pathology, hippocampal transcription, and memory in two transgenic (TG) mouse models of AD. We found that TRF had the remarkable capability of simultaneously reducing amyloid deposition, increasing Aβ42 clearance, improving sleep and memory, and normalizing daily transcription patterns of multiple genes, including those associated with AD and neuroinflammation. Thus, our study unveils for the first time the pleiotropic nature of timed feeding on AD, which has far-reaching effects beyond metabolism, ameliorating neurodegeneration and the misalignment of circadian rhythmicity. Since TRF can substantially modify disease trajectory, this intervention has immediate translational potential, addressing the urgent demand for accessible approaches to reduce or halt AD progression.
Rationale: Extraembryonic tissues, including the yolk sac and placenta, and the heart within the embryo, work to provide crucial nutrients to the embryo. The association of congenital heart defects with extraembryonic tissue defects further supports the potential developmental relationship between the heart and extraembryonic tissues. Although the development of early cardiac lineages has been well-studied, the developmental relationship between cardiac lineages, including epicardium, and extraembryonic mesoderm remains to be defined. Objective: To explore the developmental relationships between cardiac and extraembryonic lineages. Methods and Results: Through high-resolution single-cell and genetic lineage/clonal analyses, we show an unsuspected clonal relationship between extraembryonic mesoderm and cardiac lineages. Single-cell transcriptomics and trajectory analyses uncovered 2 mesodermal progenitor sources contributing to left ventricular cardiomyocytes, 1 embryonic and the other with an extraembryonic gene expression signature. Additional lineage-tracing studies revealed that the extraembryonic-related progenitors reside at the embryonic/extraembryonic interface in gastrulating embryos and produce distinct cell types forming the pericardium, septum transversum, epicardium, dorsolateral regions of the left ventricle and atrioventricular canal myocardium, and extraembryonic mesoderm. Clonal analyses demonstrated that these progenitors are multipotent, giving rise to not only cardiomyocytes and serosal mesothelial cell types but also, remarkably, extraembryonic mesoderm. Conclusions: Overall, our results reveal the location of previously unknown multipotent cardiovascular progenitors at the embryonic/extraembryonic interface and define the earliest embryonic origins of serosal mesothelial lineages, including the epicardium, which contributes fibroblasts and vascular support cells to the heart. The shared lineage relationship between embryonic cardiovascular lineages and extraembryonic mesoderm revealed by our studies underscores an underappreciated blurring of boundaries between embryonic and extraembryonic mesoderm. Our findings suggest unexpected underpinnings of the association between congenital heart disease and placental insufficiency anomalies and the potential utility of extraembryonic cells for generating cardiovascular cell types for heart repair.
Complex organs are composed of a multitude of specialized cell types which assemble to form functional biological structures. How these cell types are created and organized remains to be elucidated for many organs including the heart, the first organ to form during embryogenesis. Here, we show the ontogeny of mammalian mesoderm at high-resolution single cell and genetic lineage/clonal analyses, which revealed an unexpected complexity of the contribution and multi-potentiality of mesodermal progenitors to cardiac lineages creating distinct cell types forming specific regions of the heart. Single-cell transcriptomics of Mesp1 lineage-traced cells during embryogenesis and corresponding trajectory analyses uncovered unanticipated developmental relationships between these progenitors and lineages including two mesodermal progenitor sources contributing to the first heart field (FHF), an intraembryonic and a previously uncharacterized extraembryonic-related source, that produce distinct cardiac lineages creating the left ventricle. Lineage-tracing studies revealed that these extraembryonic-related FHF progenitors reside at the extraembryonic-intraembryonic interface in gastrulating embryos and generate cardiac cell types that form the epicardium and the dorsolateral regions of the left ventricle and atrioventricular canal myocardium. Clonal analyses further showed that these progenitors are multi-potent, creating not only cardiomyocytes and epicardial cell types but also extraembryonic mesoderm. Overall, these results reveal unsuspected multiregional origins of the heart fields, and provide new insights into the relationship between intraembryonic cardiac lineages and extraembryonic tissues and the associations between congenital heart disease and placental insufficiency anomalies.
in 2021, working in the lab of Dr Kristy Red-Horse.For his thesis, he studied the coronary vasculature during embryonic development and in
We present an accessible, fast, and customizable network propagation system for pathway boosting and interpretation of genome-wide association studies. This system-NAGA (Network Assisted Genomic Association)-taps the NDEx biological network resource to gain access to thousands of protein networks and select those most relevant and performative for a specific association study. The method works efficiently, completing genome-wide analysis in under 5 minutes on a modern laptop computer. We show that NAGA recovers many known disease genes from analysis of schizophrenia genetic data, and it substantially boosts associations with previously unappreciated genes such as amyloid beta precursor. On this and seven other gene-disease association tasks, NAGA outperforms conventional approaches in recovery of known disease genes and replicability of results. Protein interactions associated with disease are visualized and annotated in Cytoscape, which, in addition to standard programmatic interfaces, allows for downstream analysis.
Biological networks can substantially boost power to identify disease genes in genome-wide association studies. To explore different network GWAS methods, we challenged students of a UC San Diego graduate level bioinformatics course, Network Biology and Biomedicine, to explore and improve such algorithms during a four-week-long classroom competition. Here, we report the many creative solutions and share our experiences in conducting classroom crowd science as both a research and pedagogical tool.
Precision cancer medicine promises to tailor clinical decisions to patients using genomic information. Indeed, successes of drugs targeting genetic alterations in tumors, such as imatinib that targets BCR–ABL in chronic myelogenous leukemia, have demonstrated the power of this approach. However, biological systems are complex, and patients may differ not only by the specific genetic alterations in their tumor, but also by more subtle interactions among such alterations. Systems biology and more specifically, network analysis, provides a framework for advancing precision medicine beyond clinical actionability of individual mutations. Here we discuss applications of network analysis to study tumor biology, early methods for N-of-1 tumor genome analysis, and the path for such tools to the clinic.
Gene networks are rapidly growing in size and number, raising the question of which networks are most appropriate for a particular application. Here, we evaluate 21 human genome-wide interaction networks for their ability to recover gene sets associated with 446 different diseases and 9 cancer hallmarks. While all networks have some ability in these recovery tasks, we observe a wide range of performance with STRING, GeneMANIA and GIANT networks having the best performance overall. A general tendency is that performance scales with network size, suggesting that new interaction discovery currently outweighs the detrimental effects of false positives. Correcting for size, we find that the DIP network provides the highest efficiency (value per interaction). Based on these results we create a parsimonious composite network with both high efficiency and absolute performance, which outperforms any single resource. This work provides a benchmark for selection of molecular networks in human disease research. Citation Format: Justin K. Huang, Daniel E. Carlin, Michael K. Yu, Wei Zhang, Jason F. Kreisberg, Pablo Tamayo, Trey Ideker. Systematic evaluation of gene networks for discovery of disease genes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 1310.
Gene networks are rapidly growing in size and number, raising the question of which networks are most appropriate for particular applications. Here, we evaluate 21 human genome-wide interaction networks for their ability to recover 446 disease gene sets identified through literature curation, gene expression profiling, or genome-wide association studies. While all networks have some ability to recover disease genes, we observe a wide range of performance with STRING, ConsensusPathDB, and GIANT networks having the best performance overall. A general tendency is that performance scales with network size, suggesting that new interaction discovery currently outweighs the detrimental effects of false positives. Correcting for size, we find that the DIP network provides the highest efficiency (value per interaction). Based on these results, we create a parsimonious composite network with both high efficiency and performance. This work provides a benchmark for selection of molecular networks in human disease research.
We present a unified GenomeSpace recipe that combines the results of a high throughput CRISPR genetic screen and a biological network to return a subnetwork that suggests a mechanistic explanation of the screen’s results. The explanatory subnetwork is found by network propagation, a popular systems biology approach. We demonstrate our pipeline on an alpha toxin screen, revealing a subnetwork that is both highly interconnected and highly enriched for hits in the screen.
SummaryWe present pyNBS: a modularized Python 2.7 implementation of the network-based stratification (NBS) algorithm for stratifying tumor somatic mutation profiles into molecularly and clinically relevant subtypes. In addition to release of the software, we benchmark its key parameters and provide a compact cancer reference network that increases the significance of tumor stratification using the NBS algorithm. The structure of the code exposes key steps of the algorithm to foster further collaborative development.Availability and implementationThe package, along with examples and data, can be downloaded and installed from the URL https://github.com/idekerlab/pyNBS.
We report the results of a DREAM challenge designed to predict relative genetic essentialities based on a novel dataset testing 98,000 shRNAs against 149 molecularly characterized cancer cell lines. We analyzed the results of over 3,000 submissions over a period of 4 months. We found that algorithms combining essentiality data across multiple genes demonstrated increased accuracy; gene expression was the most informative molecular data type; the identity of the gene being predicted was far more important than the modeling strategy; well-predicted genes and selected molecular features showed enrichment in functional categories; and frequently selected expression features correlated with survival in primary tumors. This study establishes benchmarks for gene essentiality prediction, presents a community resource for future comparison with this benchmark, and provides insights into factors influencing the ability to predict gene essentiality from functional genetic screens. This study also demonstrates the value of releasing pre-publication data publicly to engage the community in an open research collaboration.
BACKGROUND:RAS protein interactions have predominantly been studied in the context of the RAF and PI3kinase oncogenic pathways. Structural modeling and X-ray crystallography have demonstrated that RAS isoforms bind to canonical downstream effector proteins in these pathways using the highly conserved switch I and II regions. Other non-canonical RAS protein interactions have been experimentally identified, however it is not clear whether these proteins also interact with RAS via the switch regions.RESULTS:To address this question we constructed a RAS isoform-specific protein-protein interaction network and predicted 3D complexes involving RAS isoforms and interaction partners to identify the most probable interaction interfaces. The resulting models correctly captured the binding interfaces for well-studied effectors, and additionally implicated residues in the allosteric and hyper-variable regions of RAS proteins as the predominant binding site for non-canonical effectors. Several partners binding to this new interface (SRC, LGALS1, RABGEF1, CALM and RARRES3) have been implicated as important regulators of oncogenic RAS signaling. We further used these models to investigate competitive binding and multi-protein complexes compatible with RAS surface occupancy and the putative effects of somatic mutations on RAS protein interactions.CONCLUSIONS:We discuss our findings in the context of RAS localization to the plasma membrane versus within the cytoplasm and provide a list of RAS protein interactions with possible cancer-related consequences, which could help guide future therapeutic strategies to target RAS proteins.
Network propagation is an important and widely used algorithm in systems biology, with applications in protein function prediction, disease gene prioritization, and patient stratification. However, up to this point it has required significant expertise to run. Here we extend the popular network analysis program Cytoscape to perform network propagation as an integrated function. Such integration greatly increases the access to network propagation by putting it in the hands of biologists and linking it to the many other types of network analysis and visualization available through Cytoscape. We demonstrate the power and utility of the algorithm by identifying mutations conferring resistance to Vemurafenib.
One commonly performed bioinformatics task is to infer functional regulation of transcription factors by observing differential expression under a knockout, and integrating DNA binding information of that transcription factor. However, until now, this task has required dedicated bioinformatics support to perform the necessary data integration. GenomeSpace provides a protocol, or “recipe”, and a user interface with inter-operating software tools to identify protein occupancies along the genome from a ChIP-seq experiment and associated differentially regulated genes from a RNA-Seq experiment. By integrating RNA-Seq and ChIP-seq analyses, a user is easily able to associate differing expression phenotypes with changing epigenetic landscapes.
We introduce a novel method called Prophetic Granger Causality (PGC) for inferring gene regulatory networks (GRNs) from protein-level time series data. The method uses an L1-penalized regression adaptation of Granger Causality to model protein levels as a function of time, stimuli, and other perturbations. When combined with a data-independent network prior, the framework outperformed all other methods submitted to the HPN-DREAM 8 breast cancer network inference challenge. Our investigations reveal that PGC provides complementary information to other approaches, raising the performance of ensemble learners, while on its own achieves moderate performance. Thus, PGC serves as a valuable new tool in the bioinformatics toolkit for analyzing temporal datasets. We investigate the general and cell-specific interactions predicted by our method and find several novel interactions, demonstrating the utility of the approach in charting new tumor wiring.
Pablo Tamayo合作论文数Theoretical Division and Advanced Computing Laboratory, Los Alamos National Laboratory, Los Alamos, NM2