edgeR is an R/Bioconductor software package for differential analyses of sequencing data in the form of read counts for genes or genomic features. Over the past 15 years, edgeR has been a popular choice for statistical analysis of data from sequencing technologies such as RNA-seq or ChIP-seq. edgeR pioneered the use of the negative binomial distribution to model read count data with replicates and the use of generalized linear models to analyze complex experimental designs. edgeR implements empirical Bayes moderation methods to allow reliable inference when the number of replicates is small. This article announces edgeR version 4, which includes new developments across a range of application areas. Infrastructure improvements include support for fractional counts, implementation of model fitting in C and a new statistical treatment of the quasi-likelihood pipeline that improves accuracy for small counts. The revised package has new functionality for differential methylation analysis, differential transcript expression, differential transcript and exon usage, testing relative to a fold-change threshold and pathway analysis. This article reviews the statistical framework and computational implementation of edgeR, briefly summarizing all the existing features and functionalities but with special attention to new features and those that have not been described previously.
edgeR is an R/Bioconductor software package for differential analyses of sequencing data in the form of read counts for genes or genomic features. Over the past 15 years, edgeR has been a popular choice for statistical analysis of data from sequencing technologies such as RNA-seq or ChIP-seq. edgeR pioneered the use of the negative binomial distribution to model read count data with replicates and the use of generalized linear models to analyse complex experimental designs. edgeR implements empirical Bayes moderation methods to allow reliable inference when the number of replicates is small. This article announces edgeR version 4, which includes new developments across a range of application areas. Infrastructure improvements include support for fractional counts, implementation of model fitting in C++, and a new statistical treatment of the quasi-likelihood pipeline that improves accuracy for small counts. The revised package has new functionality for differential methylation analysis, differential transcript expression, differential transcript and exon usage, testing relative to a fold-change threshold and pathway analysis. This article reviews the statistical framework and computational implementation of edgeR, briefly summarizing all the existing features and functionalities but with special attention to new features and those that have not been described previously.### Competing Interest StatementThe authors have declared no competing interest.
Macrophages exhibit remarkable functional plasticity, a requirement for their central role in tissue homeostasis. During chronic inflammation, macrophages acquire sustained inflammatory 'states' that contribute to disease, but there is limited understanding of the regulatory mechanisms that drive their generation. Here we describe a systematic functional genomics approach that combines genome-wide phenotypic screening in primary murine macrophages with transcriptional and cytokine profiling of genetic perturbations in primary human macrophages to uncover regulatory circuits of inflammatory states. This process identifies regulators of five distinct states associated with key features of macrophage function. Among these regulators, loss of the N-6-methyladenosine (m6A) writer components abolishes m6A modification of TNF transcripts, thereby enhancing mRNA stability and TNF production associated with multiple inflammatory pathologies. Thus, phenotypic characterization of primary murine and human macrophages describes the regulatory circuits underlying distinct inflammatory states, revealing post-transcriptional control of TNF mRNA stability as an immunosuppressive mechanism in innate immunity.
Ferroptosis is an iron-dependent cell death mechanism characterized by the accumulation of toxic lipid peroxides and cell membrane rupture. GPX4 (glutathione peroxidase 4) prevents ferroptosis by reducing these lipid peroxides into lipid alcohols. Ferroptosis induction by GPX4 inhibition has emerged as a vulnerability of cancer cells, highlighting the need to identify ferroptosis regulators that may be exploited therapeutically. Through genome-wide CRISPR activation screens, we identify the SWI/SNF (switch/sucrose non-fermentable) ATPases BRM (SMARCA2) and BRG1 (SMARCA4) as ferroptosis suppressors. Mechanistically, they bind to and increase chromatin accessibility at NRF2 target loci, thus boosting NRF2 transcriptional output to counter lipid peroxidation and confer resistance to GPX4 inhibition. We further demonstrate that the BRM/BRG1 ferroptosis connection can be leveraged to enhance the paralog dependency of BRG1 mutant cancer cells on BRM. Our data reveal ferroptosis induction as a potential avenue for broadening the efficacy of BRM degraders/inhibitors and define a specific genetic context for exploiting GPX4 dependency.
Blood-borne pathogens can cause systemic inflammatory response syndrome (SIRS) followed by protracted, potentially lethal immunosuppression. The mechanisms responsible for impaired immunity post-SIRS remain unclear. We show that SIRS triggered by pathogen mimics or malaria infection leads to functional paralysis of conventional dendritic cells (cDCs). Paralysis affects several generations of cDCs and impairs immunity for 3-4 weeks. Paralyzed cDCs display distinct transcriptomic and phenotypic signatures and show impaired capacity to capture and present antigens in vivo. They also display altered cytokine production patterns upon stimulation. The paralysis program is not initiated in the bone marrow but during final cDC differentiation in peripheral tissues under the influence of local secondary signals that persist after resolution of SIRS. Vaccination with monoclonal antibodies that target cDC receptors or blockade of transforming growth factor β partially overcomes paralysis and immunosuppression. This work provides insights into the mechanisms of paralysis and describes strategies to restore immunocompetence post-SIRS.
Single-cell multiomic technologies enable holistic reconstruction of cell states and hold the potential to identify drivers of tumor evolution and drug resistance. However, the data sparsity of single-cell assays can preclude accurate estimation of transcription factor activity. In addition, gene expression alone cannot capture transcription factor activity, especially in the context of drug treatment which alters transcription factor function without suppressing gene expression. To circumvent these challenges, we have developed epiregulon, a R package that constructs gene regulatory networks and infers transcription factor (TF) activity in single cells by integrating single cell gene expression, chromatin accessibility and bulk TF occupancy data. Epiregulon applies tests of independence to identify likely transcription factor - regulatory element - target gene relationships occurring at joint probabilities exceeding the expected probabilities of independent events. With epiregulon, we are able to detect lineage factor activity at enhanced sensitivity and predict perturbations more accurately than gene expression could. We further applied this tool to understand transcription factor activity to enzalutamide treatment and identify potential drivers of enzalutamide resistance. Finally, we generated a ground truth dataset using reprogram-seq, a technology that captures multi-omic profiles of single cells upon expression of defined transcription factors. Epiregulon correctly predicts NKX2-1 targets and demonstrates that NKX2-1 expression reprograms the epigenome towards a neuroendocrine state in LNCaP cells. Epiregulon is a useful tool that constructs gene regulatory networks to model the underlying gene regulation hierarchies that drive gene expression and cell states. Citation Format: Tomasz Wlodarczyk, Jenille Tan, Kerstin Seidel, Diana Wu, Shang-Yang Chen, Aleksander Chlebowski, Timothy Keyes, Yu Guo, Aaron Lun, Christopher Siebel, Shiqi Xie, Xiaosai Yao. Epiregulon infers single-cell transcription factor activity to dissect mechanisms of lineage plasticity and drug response [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 3141.
gesel is a JavaScript package for performing gene set enrichment analyses within the browser. All calculations are performed on the client device, without any no need for a dedicated backend server. This eliminates concerns around cost, scalability, latency, and data ownership that are associated with a backend-based architecture. We demonstrate the use of gesel with a basic web application that performs enrichment analyses on user-supplied genes with sets derived from the Gene Ontology and MSigDB. Developers can also use gesel to incorporate gene set enrichment capabilities into their own applications.
We present kana, a web application for interactive single-cell ‘omics data analysis in the browser. Like, literally, in the browser: kana leverages web technologies such as WebAssembly to efficiently perform the relevant computations on the user’s machine, avoiding the need to provision and maintain a backend service. The application provides a streamlined one-click workflow for the main steps in a typical single-cell analysis, starting from a count matrix and finishing with marker detection. Results are presented in an intuitive web interface for further exploration and iterative analysis. Testing on public datasets shows that kana can analyze over 100,000 cells within 5 minutes on a typical laptop.
The success of CRISPR-mediated gene perturbation studies is highly dependent on the quality of gRNAs, and several tools have been developed to enable optimal gRNA design. However, these tools are not all adaptable to the latest CRISPR modalities or nucleases, nor do they offer comprehensive annotation methods for advanced CRISPR applications. Here, we present a new ecosystem of R packages, called crisprVerse , that enables efficient gRNA design and annotation for a multitude of CRISPR technologies. This includes CRISPR knockout (CRISPRko), CRISPR activation (CRISPRa), CRISPR interference (CRISPRi), CRISPR base editing (CRISPRbe) and CRISPR knockdown (CRISPRkd). The core package, crisprDesign , offers a user-friendly and unified interface to add off-target annotations, rich gene and SNP annotations, and on- and off-target activity scores. These functionalities are enabled for any RNA- or DNA-targeting nucleases, including Cas9, Cas12, and Cas13. The crisprVerse ecosystem is open-source and deployed through the Bioconductor project ( https://github.com/crisprVerse ).
We present kana, a web application for interactive single-cell RNA-seq (scRNA-seq) data analysis in the browser. Like, literally, in the browser: kana leverages web technologies such as WebAssembly to efficiently perform the relevant computations on the user’s machine, avoiding the need to provision and maintain a backend service. The application provides a streamlined one-click workflow for all steps in a typical scRNA-seq analysis, starting from a count matrix and finishing with marker detection. Results are presented in an intuitive web interface for further exploration and iterative analysis. Testing on public datasets shows that kana can analyze over 100,000 cells within 5 minutes on a typical laptop.
Senescence is a stress-responsive tumor suppressor mechanism associated with expression of the senescence-associated secretory phenotype (SASP). Through the SASP, senescent cells trigger their own immune-mediated elimination, which if evaded leads to tumorigenesis. Senescent parenchymal cells are separated from circulating immunocytes by the endothelium, which is targeted by microenvironmental signaling. Here we show that SASP induces endothelial cell NF-κB activity and that SASP-induced endothelial expression of the canonical NF-κB componentRelaunderpins senescence surveillance. Using human liver sinusoidal endothelial cells (LSECs), we show that SASP-induced endothelial NF-κB activity regulates a conserved transcriptional program supporting immunocyte recruitment. Furthermore, oncogenic hepatocyte senescence drives murine LSEC NF-κB activity in vivo. Critically, we show two distinct endothelial pathways in senescence surveillance. First, endothelial-specific loss ofRelaprevents development of Stat1-expressing CD4+T lymphocytes. Second, the SASP up-regulates ICOSLG on LSECs, with the ICOS–ICOSLG axis contributing to senescence cell clearance. Our results show that the endothelium is a nonautonomous SASP target and an organizing center for immune-mediated senescence surveillance.
Doublets are prevalent in single-cell sequencing data and can lead to artifactual findings. A number of strategies have therefore been proposed to detect them. Building on the strengths of existing approaches, we developed scDblFinder, a fast, flexible and accurate Bioconductor-based doublet detection method. Here we present the method, justify its design choices, demonstrate its performance on both single-cell RNA and accessibility (ATAC) sequencing data, and provide some observations on doublet formation, detection, and enrichment analysis. Even in complex datasets, scDblFinder can accurately identify most heterotypic doublets, and was already found by an independent benchmark to outcompete alternatives.
Summary SpatialExperiment is a new data infrastructure for storing and accessing spatially resolved transcriptomics data, implemented within the R/Bioconductor framework, which provides advantages of modularity, interoperability, standardized operations, and comprehensive documentation. Here, we demonstrate the structure and user interface with examples from the 10x Genomics Visium and seqFISH platforms, and provide access to example datasets and visualization tools in the STexampleData, TENxVisiumData, and ggspavis packages. Availability and Implementation The SpatialExperiment, STexampleData, TENxVisiumData, and ggspavis packages are available from Bioconductor. The package versions described in this manuscript are available in Bioconductor version 3.15 onwards. Contact risso.davide@gmail.com, shicks19@jhu.edu Supplementary Information Supplementary Tables and Figures are available online.
basilisk is an R/Bioconductor package for managing Python environments within the Bioconductor package ecosystem.Developers of other Bioconductor packages can use basilisk to automatically provision and load custom Python environments, providing a streamlined experience for their end-users by avoiding the need for any manual system configuration.basilisk also enables robust execution of Python code via reticulate in complex analysis workflows involving multiple Python environments.This package aims to provide a standardized mechanism for integration of Python functionality into the Bioconductor code base.
Transposable elements (TEs) regulate diverse biological processes, from early development to cancer. Expression of young TEs is difficult to measure with next-generation, single-cell sequencing technologies because their highly repetitive nature means that short complementary DNA reads cannot be unambiguously mapped to a specific locus. Single CELl LOng-read RNA-sequencing (CELLO-seq) combines long-read single cell RNA-sequencing with computational analyses to measure TE expression at unique loci. We used CELLO-seq to assess the widespread expression of TEs in two-cell mouse blastomeres as well as in human induced pluripotent stem cells. Across both species, old and young TEs showed evidence of locus-specific expression with simulations demonstrating that only a small number of very young elements in the mouse could not be mapped back to the reference with high confidence. Exploring the relationship between the expression of individual elements and putative regulators revealed large heterogeneity, with TEs within a class showing different patterns of correlation and suggesting distinct regulatory mechanisms.
Genome stability relies on proper coordination of mitosis and cytokinesis, where dynamic microtubules capture and faithfully segregate chromosomes into daughter cells. With a high-content RNAi imaging screen targeting more than 2,000 human lncRNAs, we identify numerous lncRNAs involved in key steps of cell division such as chromosome segregation, mitotic duration and cytokinesis. Here, we provide evidence that the chromatin-associated lncRNA, linc00899 , leads to robust mitotic delay upon its depletion in multiple cell types. We perform transcriptome analysis of linc00899 -depleted cells and identify the neuronal microtubule-binding protein, TPPP/p25 , as a target of linc00899 . We further show that linc00899 binds TPPP/p25 and suppresses its transcription. In cells depleted of linc00899 , upregulation of TPPP/p25 alters microtubule dynamics and delays mitosis. Overall, our comprehensive screen uncovers several lncRNAs involved in genome stability and reveals a lncRNA that controls microtubule behaviour with functional implications beyond cell division.
Abstract Assessing the response of cancer cell lines to drugs and treatments that affect their growth is the cornerstone of drug discovery and development. On one hand, large screens are performed across many lines and drugs in a semi-automated manner. On the other hand, small-scale studies, for example focused on factors that contribute to sensitivity and resistance, are generally performed in labs with limited automation. Data from these complementary approaches are rarely handled in the same manner: commercial software are available for the former case, whereas experimentalists in the latter case often handle files manually and process data in spreadsheets. Here we propose a suite of computational tools that enable the processing, archiving, and visualization of drug response data from any experiment, thus ensuring reproducibility and implementation of the FAIR principles. The core of the gDR suite is a database designed to host drug response data from large screens, single-agent treatments, combination experiments, as well as data from complex experimental designs including ligand co-treatments or shRNA. Experimentalists can upload metadata files along with unprocessed output files from plate readers in an interactive web application built with R Shiny. Data normalization and quality control are performed by the software before the data is stored in the database. Data from commercial software or other databases can also be easily pushed into the database through a REST API. To fetch the processed data, we developed an R package as well as another R shiny-based web application. Both tools allow users to efficiently search the database, easily plot the data from the selected experiments, and perform basic analyses. Overall, the gDR suite is a modular software that provides an end-to-end solution for managing drug response data. At the conference, we will describe our implementation and demonstrate how to use the gDR suite. We hope that this tool will facilitate the handling of drug response data and thus contribute to the quality of published data. In addition, we hope that the community will contribute to the gDR suite by adding new functionalities for either importing, plotting, or analyzing the data. Citation Format: Marc Hafner, Arkadiusz Gladki, Jane Li, Eva Lin, Aaron Lun, Scott Martin, Natalia Potocka, Dariusz Scigocki, Steffan Vartanian, Jan Vogel, Allison Vuong. gDR suite: an integrative solution for handling cell line drug response data [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 828.
The role of Transposable Elements (TEs) in regulating diverse biological processes, from early development to cancer, is becoming increasing appreciated. However, unlike other biological processes, next generation single-cell sequencing technologies are ill-suited for assaying TE expression: in particular, their highly repetitive nature means that short cDNA reads cannot be unambiguously mapped to a specific locus. Consequently, it is extremely challenging to understand the mechanisms by which TE expression is regulated and how they might themselves regulate other protein coding genes. To resolve this, we introduce CELLO-seq, a novel method and computational framework for performing long-read RNA sequencing at single cell resolution. CELLO-seq allows for full-length RNA sequencing and enables measurement of allelic, isoform and TE expression at unique loci. We use CELLO-seq to assess the widespread expression of TEs in 2-cell mouse blastomeres as well as human induced pluripotent stem cells (hiPSCs). Across both species, old and young TEs showed evidence of locus-specific expression, with simulations demonstrating that only a small number of very young elements in the mouse could not be mapped back to with high confidence. Exploring the relationship between the expression of individual elements and putative regulators revealed surprising heterogeneity, with TEs within a class showing different patterns of correlation, suggesting distinct regulatory mechanisms. ### Competing Interest Statement The authors have declared no competing interest. * (CELLO-seq) : CELl LOng read RNA sequencing (PacBio) : Pacific Biosciences (MaLRs) : Mammalian apparent LTR-retrotransposons (SNP) : single nucleotide polymorphism (sc) : single cell (TE) : Transposable element (ERV) : endogenous retrovirus (LINE) : Long interspersed element (SINE) : Short interspersed element (RNAseq) : RNA sequencing (hiPSCs) : Human induced pluripotent stem cells (RT) : Reverse Transcription (LTRs) : long terminal repeat elements (NGS) : next generation sequencing (ONT) : Oxford Nanopore technologies (UMIs) : unique molecular identifiers (TSO) : template switch oligo (ZNFs) : Zinc finger nucleases (nt) : Nucleotides (bp) : Base pairs (TSS) : Transcription start site (TES) : Transcription end site (ERCC) : External RNA Controls Consortium
Conventional human embryonic stem cells are considered to be primed pluripotent but can be induced to enter a naive state. However, the transcriptional features associated with naive and primed pluripotency are still not fully understood. Here we used single-cell RNA sequencing to characterize the differences between these conditions. We observed that both naive and primed populations were mostly homogeneous with no clear lineage-related structure and identified an intermediate sub-population of naive cells with primed-like expression. We found that the naive-primed pluripotency axis is preserved across species, although the timing of the transition to a primed state is species specific. We also identified markers for distinguishing human naive and primed pluripotency as well as strong coregulatory relationships between lineage markers and epigenetic regulators that were exclusive to naive cells. Our data provide valuable insights into the transcriptional landscape of human pluripotency at a cellular and genome-wide resolution.