To understand phenotypic evolution, it is essential to investigate the underlying gene regulatory networks (GRNs). However, most comparative GRN analyzes remain descriptive due to the low signal-to-noise ratio inherent in single-cell transcriptomics data. To address this, we introduce CroCoNet (Cross-species Comparison of Networks), an R-package for quantitative GRN comparison across species. CroCoNet builds comparable network modules centered on putative regulators and compares module topologies within and between species, distinguishing true evolutionary divergence from technical and biological confounders. We demonstrate its utility by comparing early neural differentiation across primates and validating results with a CRISPRi analysis of the diverged POU5F1 module.
The identification of cell types remains a major challenge. Even after a decade of single-cell RNA sequencing (scRNA-seq), reasonable cell type annotations almost always include manual non-automated steps. The identification of orthologous cell types across species complicates matters even more, but at the same time strengthens the confidence in the assignment. Here, we generate and analyze a dataset consisting of embryoid bodies (EBs) derived from induced pluripotent stem cells (iPSCs) of four primate species: humans, orangutans, cynomolgus, and rhesus macaques. This kind of data includes a continuum of developmental cell types, multiple batch effects (i.e. species and individuals) and uneven cell type compositions and hence poses many challenges. We developed a semi-automated computational pipeline combining classification and marker-based cluster annotation to identify orthologous cell types across primates. This approach enabled the investigation of cross-species conservation of gene expression. Consistent with previous studies, our data confirm that broadly expressed genes are more conserved than cell type-specific genes, raising the question of how conserved, inherently cell type-specific, marker genes are. Our analyses reveal that human marker genes are less effective in macaques and vice versa, highlighting the limited transferability of markers across species. Overall, our study advances the identification of orthologous cell types across species, provides a well-curated cell type reference for future in vitro studies and informs the transferability of marker genes across species.
BACKGROUND Gene regulatory changes play a central role in shaping cellular phenotypes across species. To understand how these phenotypes evolve, it is essential to investigate the underlying gene regulatory networks (GRNs). However, most comparative analyses of GRNs remain qualitative and are therefore sensitive to false positives and false negatives. RESULTS To address this limitation, we introduce CroCoNet ( Cro ss-species Co mparison of Net works), an R package for the quantitative comparison of GRNs across species. CroCoNet constructs comparable network modules centered on known transcriptional regulators and quantifies the variability in module topology within and between species. By contrasting these levels of variability, CroCoNet can distinguish true evolutionary divergence from technical and biological confounders. Applying CroCoNet to scRNA-seq data from the early neural differentiation of human, gorilla, and cynomolgus macaque, we identified 20 conserved and 24 diverged modules. Despite the conserved expression pattern of the pluripotency factor POU5F1 (OCT4), its associated module was among the most diverged. This result was independently confirmed through cross-species CRISPRi perturbations coupled with single-cell RNA-seq as a readout. Moreover, we found that great ape- and human-specific LTR7 elements are enriched near POU5F1 module genes, potentially contributing to the cross-species differences in network topology. CONCLUSIONS These findings demonstrate that CroCoNet can resolve regulatory rewiring and provides a robust framework for studying GRN evolution across closely related species. CroCoNet is available as an open-source R package at . ### Competing Interest Statement The authors have declared no competing interest. * ATAC-seq : assay for transposase-accessible chromatin using sequencing CCDS : canonical coding sequence CDS : coding sequence ChIP-seq : chromatin immunoprecipitation sequencing cor.adj : correlation of adjacencies cor.kIM : correlation of intramodular connectivities CRE : cis-regulatory element CRISPRi : clustered regularly interspaced short palindromic repeats interference DE : differential expression/differentially expressed DR : differential regulation/differentially regulated FDR : false discovery rate GEX : gene expression GRN : gene regulatory network iPSC : induced pluripotent stem cell kIM : intramodular connectivity MOI : multiplicity of infection MRCA : most recent common ancestor NPC : neural precursor cell PCBC : Progenitor Cell Biology Consortium PGLS : Phylogenetic Generalized Least Squares scRNA-seq : single-cell RNA sequencing SNP : single nucleotide polymorphisms TE : transposable element TR : transcriptional regulator TSS : transcription start site UMI : unique molecular identifier UTR : untranslated region Deutsche Forschungsgemeinschaft, https://ror.org/018mejw64, 458888224, 458247426, 407541155, INST 86/2050-1 FUGG
Comparisons of molecular phenotypes across primates provide unique information to understand human biology and evolution, and single-cell RNA-seq CRISPR interference (CRISPRi) screens are a powerful approach to analyze them. Here, we generate and validate three human, three gorilla, and two cynomolgus iPS cell lines that carry a dox-inducible KRAB-dCas9 construct at the AAVS1 locus. We show that despite variable expression levels of KRAB-dCas9 among lines, comparable downregulation of target genes and comparable phenotypic effects are observed in a single-cell RNA-seq CRISPRi screen. Hence, we provide valuable resources for performing and further extending CRISPRi in human and non-human primates.
Pleiotropy, measured as expression breadth across tissues, is one of the best predictors for protein sequence and expression conservation. In this study, we investigated its effect on the evolution ofcis-regulatory elements (CREs). To this end, we carefully reanalyzed the Epigenomics Roadmap data for nine fetal tissues, assigning a measure of pleiotropic degree to nearly half a million CREs. To assess the functional conservation of CREs, we generated ATAC-seq and RNA-seq data from humans and macaques. We found that more pleiotropic CREs exhibit greater conservation in accessibility, and the mRNA expression levels of the associated genes are more conserved. This trend of higher conservation for higher degrees of pleiotropy persists when analyzing the transcription factor binding repertoire. In contrast, simple DNA sequence conservation of orthologous sites between species tends to be even lower for pleiotropic CREs than for species-specific CREs. Combining various lines of evidence, we propose that the lack of sequence conservation in functionally conserved pleiotropic CREs is owing to within-element compensatory evolution. In summary, our findings suggest that pleiotropy is also a good predictor for the functional conservation of CREs, even though this is not reflected in the sequence conservation of pleiotropic CREs.
Background In droplet-based single-cell and single-nucleus RNA-seq experiments, not all reads associated with one cell barcode originate from the encapsulated cell. Such background noise is attributed to spillage from cell-free ambient RNA or barcode swapping events. Results Here, we characterize this background noise exemplified by three scRNA-seq and two snRNA-seq replicates of mouse kidneys. For each experiment, cells from two mouse subspecies are pooled, allowing to identify cross-genotype contaminating molecules and thus profile background noise. Background noise is highly variable across replicates and cells, making up on average 3–35% of the total counts (UMIs) per cell and we find that noise levels are directly proportional to the specificity and detectability of marker genes. In search of the source of background noise, we find multiple lines of evidence that the majority of background molecules originates from ambient RNA. Finally, we use our genotype-based estimates to evaluate the performance of three methods (CellBender, DecontX, SoupX) that are designed to quantify and remove background noise. We find that CellBender provides the most precise estimates of background noise levels and also yields the highest improvement for marker gene detection. By contrast, clustering and classification of cells are fairly robust towards background noise and only small improvements can be achieved by background removal that may come at the cost of distortions in fine structure. Conclusions Our findings help to better understand the extent, sources and impact of background noise in single-cell experiments and provide guidance on how to deal with it.
Figure S1 Recurrently mutated genes and comparison to TCGA cohort. Figure S2 Stabililty of recurrently mutated genes over disease course. Figure S3 Variant allele frequency plots from diagnosis to relapse for each individual patient. Figure S4 Mutation patterns of individual genes. Lines represent mutations in individual patients. Figure S5 Detection of the pre-existence of gained variants at initial diagnosis. Figure S6 Mutations in AML-associated functional pathways. Figure S7 Deletion of exons 3-10 of the KDM6A gene in the MM-6 cell line. The sister cell line MM-1 is not affected. Deletions were detected by quantitative MLPA analysis and the peak ratio for each KDM6A exon specific probe is shown. Graph represents the results from two independent experiments. Figure S8 H3K27 methylation in the AML cell lines MM-1 and MM-6. The global mono-, di- and tri-methylation of H3K27 in MM-1 and MM-6 was analyzed by Western blot and normalized to H3. The median of three independent experiments is shown. P-Values were calculated using a two-tailed, unpaired Student's t-test. Figure S9 The MM-6 cell line is resistant to cytarabine (Ara-C) but not to Daunorubicin or AC220. Error bars indicate mean {plus minus} s.d. of at least three independent experiments. **P<0.01; unpaired, two-tailed Student's t-test; n.s., not significant . Figure S10 A Correlation of KDM6A expression (low{less than or equal to}25th percentile vs. high>75th percentile) with clinical outcome according to gender in the AMLCG-99 trial (NCT00266136). Figure S11 Proportion of transversions from diagnosis to relapse. Transversion frequencies are shown for lost (blue), stable (orange) and gained (red) mutations. Dots represent individual patients. Figure S12 Median coverage in targeted sequencing of persistent and non-persistent DNMT3A variant positions at complete remission (CR). Figure S13 Clinical outcome according to mutation persistence at remission.
Brain size and cortical folding have increased and decreased recurrently during mammalian evolution. Identifying genetic elements whose sequence or functional properties co-evolve with these traits can provide unique information on evolutionary and developmental mechanisms. A good candidate for such a comparative approach is TRNP1, as it controls proliferation of neural progenitors in mice and ferrets. Here, we investigate the contribution of both regulatory and coding sequences of TRNP1 to brain size and cortical folding in over 30 mammals. We find that the rate of TRNP1 protein evolution (ω) significantly correlates with brain size, slightly less with cortical folding and much less with body size. This brain correlation is stronger than for >95% of random control proteins. This co-evolution is likely affecting TRNP1 activity, as we find that TRNP1 from species with larger brains and more cortical folding induce higher proliferation rates in neural stem cells. Furthermore, we compare the activity of putative cis-regulatory elements (CREs) of TRNP1 in a massively parallel reporter assay and identify one CRE that likely co-evolves with cortical folding in Old World monkeys and apes. Our analyses indicate that coding and regulatory changes that increased TRNP1 activity were positively selected either as a cause or a consequence of increases in brain size and cortical folding. They also provide an example how phylogenetic approaches can inform biological mechanisms, especially when combined with molecular phenotypes across several species.
Acute myeloid leukemia (AML) patients suffer dismal prognosis upon treatment resistance. To study functional heterogeneity of resistance, we generated serially transplantable patient-derived xenograft (PDX) models from one patient with AML and twelve clones thereof, each derived from a single stem cell, as proven by genetic barcoding. Transcriptome and exome sequencing segregated clones according to their origin from relapse one or two. Undetectable for sequencing, multiplex fluorochrome-guided competitive in vivo treatment trials identified a subset of relapse two clones as uniquely resistant to cytarabine treatment. Transcriptional and proteomic profiles obtained from resistant PDX clones and refractory AML patients defined a 16-gene score that was predictive of clinical outcome in a large independent patient cohort. Thus, we identified novel genes related to cytarabine resistance and provide proof of concept that intra-tumor heterogeneity reflects inter-tumor heterogeneity in AML.
Cost-efficient library generation by early barcoding has been central in propelling single-cell RNA sequencing. Here, we optimize and validate prime-seq, an early barcoding bulk RNA-seq method. We show that it performs equivalently to TruSeq, a standard bulk RNA-seq method, but is fourfold more cost-efficient due to almost 50-fold cheaper library costs. We also validate a direct RNA isolation step, show that intronic reads are derived from RNA, and compare cost-efficiencies of available protocols. We conclude that prime-seq is currently one of the best options to set up an early barcoding bulk RNA-seq protocol from which many labs would profit.
e14044 Background: Aggressive brain tumors like glioblastoma depend on support by their local environment. While the role of tumor-associated myeloid cells on glioblastoma progression is well-documented, we have only partial knowledge of the pathological impact of glioblastoma -parenchymal progenitor cells. Methods: We investigated the glioblastoma microenvironment with transgenic lineage-tracing models ( nestin-creER2, R26-tdTomato and sox2-creER2,R26-tdTomato), intravital imaging, single-cell transcriptomics, immunofluorescence and flow-cytometry as well as histopathology and characterized a previously unknown tumor-associated progenitor cell. In functional experiments, we studied the knockout of the transcription factor SOX2 in these tumor-associated cells. Results: Lineage-traced cells from mouse glioblastoma were obtained by flow-cytometry and single cell transcriptomes compared to established gene expression data from brain tumor parenchymal cells. The traced tumor-associated cells had a transcriptomic profile largely resembling myeloid cells and expressed microglia-/macrophage-markers on the protein-level. However, transgenic models and bone-marrow chimera revealed that the traced cells were clearly distinct from microglia or macrophages. The traced tumor associated cells with a myeloid expression profile derived from a SOX2-dependent progenitor cell. Consequently, conditional Sox2-knockout ablated the entire myeloid-like cell population. Remarkably, this tumor-associated cell population had a large impact on disease-progression causing significant reduction of glioblastoma –vascularization to 53%, changing vascular function and leading to a decrease in tumor volume to 42% as compared to controls. The myeloid-like progenitor cells were identified in human brain tumors by immunofluorescence and in scRNA-seq data. Conclusions: We identified a previously unacknowledged population of tumor-associated progenitor cells with a myeloid-like expression profile that transiently appeared during glioblastoma growth. These progenitors have strong impact on glioblastoma progression and point towards a new and promising therapeutic target in order to support anti-angiogenic regimen in glioblastoma.
Genomes can be seen as notebooks of evolution on successful genetic experiments. Comparing genomes across species allows identifying conserved DNA regions responsible for conserved biological functions, but additional information could be leveraged when also considering phenotypic variance across species. Here, we exemplify such a "cross-species association study" for the gene TRNP1 that is known to be important for brain growth and cortical folding in mice and ferrets.We find that the rate of TRNP1 protein evolution (ω) co-evolves with the rate of cortical folding across mammals and that TRNP1 proteins from species with more cortical folding induce higher proliferation rates in neural stem cells. Furthermore, we compare the activity of putative cis-regulatory elements (CREs) of TRNP1 in a massively parallel reporter assay (MPRA) and identify one CRE that also co-evolves with cortical folding in Old World Monkeys and Apes. Our analyses indicate that coding and regulatory changes in TRNP1 have modulated its activity to adjust cortical folding during mammalian evolution. They also provide a blueprint how cross-species association studies could help to better understand the evolution and the molecular mechanisms of phenotypes.
Aggressive brain tumors like glioblastoma depend on support by their local environment and subsets of tumor parenchymal cells may promote specific phases of disease progression. We investigated the glioblastoma microenvironment with transgenic lineage-tracing models, intravital imaging, single-cell transcriptomics, immunofluorescence analysis as well as histopathology and characterized a previously unacknowledged population of tumor-associated cells with a myeloid-like expression profile (TAMEP) that transiently appeared during glioblastoma growth. TAMEP of mice and humans were identified with specific markers. Notably, TAMEP did not derive from microglia or peripheral monocytes but were generated by a fraction of CNS-resident, SOX2-positive progenitors. Abrogation of this progenitor cell population, by conditional Sox2-knockout, drastically reduced glioblastoma vascularization and size. Hence, TAMEP emerge as a tumor parenchymal component with a strong impact on glioblastoma progression.
Genomes can be seen as notebooks of evolution that contain unique information on successful genetic experiments (Wright 2001). This allows to identify conserved genomic sequences (Zoonomia Consortium 2020) and is very useful e.g. for finding disease-associated variants (Kircher et al. 2014). Additional information from genome comparisons across species can be leveraged when considering phenotypic variance across species. Here, we exemplify this principle in a cross-species association study by testing whether structural or regulatory changes in TRNP1 correlate with changes in brain size and cortical folding in mammals. We find that the rate of TRNP1 protein evolution (ω) correlates best with the rate of cortical folding and that TRNP1 proteins from species with more cortical folding induce higher proliferation rates in neural stem cells from murine cerebral cortex. Furthermore, we compare the activity of putative cis-regulatory elements of TRNP1 in a massively parallel reporter assay and identify one element that correlates with cortical folding in Old World Monkeys and Apes. Our analyses indicate that coding and regulatory changes in TRNP1 modulated its activity to adjust cortical folding during mammalian evolution and exemplify how evolutionary information can be leveraged by cross-species association studies.
Abstract Aggressive brain tumors like glioblastoma depend on support by their local environment and subsets of tumor parenchymal cells may promote specific phases of disease progression. We investigated the glioblastoma microenvironment with transgenic lineage-tracing models, intravital imaging, single-cell transcriptomics, immunofluorescence analysis as well as histopathology and characterized a previously unacknowledged population of tumor-associated cells with a myeloid-like expression profile (TAMEP) that transiently appeared during glioblastoma growth. TAMEP of mice and humans were identified with specific markers. Notably, TAMEP did not derive from microglia or peripheral monocytes but were generated by a fraction of CNS-resident, SOX2-positive progenitors. Abrogation of this progenitor cell population, by conditional Sox2-knockout, drastically reduced glioblastoma vascularization and size. Hence, TAMEP emerge as a tumor parenchymal component with a strong impact on glioblastoma progression.
Human pluripotent stem cells (PSCs) express human endogenous retrovirus type-H (HERV-H), which exists as more than a thousand copies on the human genome and frequently produces chimeric transcripts as long-non-coding RNAs (lncRNAs) fused with downstream neighbor genes. Previous studies showed that HERV-H expression is required for the maintenance of PSC identity, and aberrant HERV-H expression attenuates neural differentiation potentials, however, little is known about the actual of function of HERV-H. In this study, we focused on ESRG, which is known as a PSC-related HERV-H-driven lncRNA. The global transcriptome data of various tissues and cell lines and quantitative expression analysis of PSCs showed that ESRG expression is much higher than other HERV-Hs and tightly silenced after differentiation. However, the loss of function by the complete excision of the entire ESRG gene body using a CRISPR/Cas9 platform revealed that ESRG is dispensable for the maintenance of the primed and naïve pluripotent states. The loss of ESRG hardly affected the global gene expression of PSCs or the differentiation potential toward trilineage. Differentiated cells derived from ESRG-deficient PSCs retained the potential to be reprogrammed into induced PSCs (iPSCs) by the forced expression of OCT3/4, SOX2, and KLF4. In conclusion, ESRG is dispensable for the maintenance and recapturing of human pluripotency.
The recent rapid spread of single cell RNA sequencing (scRNA-seq) methods has created a large variety of experimental and computational pipelines for which best practices have not yet been established. Here, we use simulations based on five scRNA-seq library protocols in combination with nine realistic differential expression (DE) setups to systematically evaluate three mapping, four imputation, seven normalisation and four differential expression testing approaches resulting in ~3000 pipelines, allowing us to also assess interactions among pipeline steps. We find that choices of normalisation and library preparation protocols have the biggest impact on scRNA-seq analyses. Specifically, we find that library preparation determines the ability to detect symmetric expression differences, while normalisation dominates pipeline performance in asymmetric DE-setups. Finally, we illustrate the importance of informed choices by showing that a good scRNA-seq pipeline can have the same impact on detecting a biological signal as quadrupling the sample size.