The increasing number of single-cell gene expression atlases available represent a potential revolution in understanding physio-pathological processes. To fully leverage this single-cell revolution, we need to enhance data integration and cell annotation strategies, with a particular emphasis on addressing the challenges posed by imbalanced cell type proportions and substantial batch effects. scMusketeers, a deep learning model, optimizes the latent data representation and solves all at once these challenges. scMusketeers features three neural modules: (1) an autoencoder for noise and dimensionality reductions; (2) a focal loss classifier to enhance rare cell type predictions; and (3) an adversarial domain adaptation (DANN) module for batch effect correction. Benchmarking against state-of-the-art tools, including the UCE foundation model, showed that scMusketeers performs on par or better, particularly in identifying rare cell types. It also allows to transfer cell labels from single-cell RNA sequencing to spatial transcriptomics. With its modular and adaptable design, scMusketeers offers a versatile framework that can be generalized to other large-scale biological projects requiring deep learning approaches, establishing itself as a valuable tool for single-cell data integration and analysis. ### Competing Interest Statement The authors have declared no competing interest.
ABSTRACTOrgan- and body-scale cell atlases have the potential to transform our understanding of human biology. To capture the variability present in the population, these atlases must include diverse demographics such as age and ethnicity from both healthy and diseased individuals. The growth in both size and number of single-cell datasets, combined with recent advances in computational techniques, for the first time makes it possible to generate such comprehensive large-scale atlases through integration of multiple datasets. Here, we present the integrated Human Lung Cell Atlas (HLCA) combining 46 datasets of the human respiratory system into a single atlas spanning over 2.2 million cells from 444 individuals across health and disease. The HLCA contains a consensus re-annotation of published and newly generated datasets, resolving under- or misannotation of 59% of cells in the original datasets. The HLCA enables recovery of rare cell types, provides consensus marker genes for each cell type, and uncovers gene modules associated with demographic covariates and anatomical location within the respiratory system. To facilitate the use of the HLCA as a reference for single-cell lung research and allow rapid analysis of new data, we provide an interactive web portal to project datasets onto the HLCA. Finally, we demonstrate the value of the HLCA reference for interpreting disease-associated changes. Thus, the HLCA outlines a roadmap for the development and use of organ-scale cell atlases within the Human Cell Atlas.
The expanding genus Yersinia is composed of multiple nonpathogenic species and a few pathogenic species, including the deadly etiologic agent of plague, Yersinia pestis . In 2 decades, the number of genomic, transcriptomic, and proteomic studies on Yersinia grew massively, delivering a wealth of data.
Listeria monocytogenes is a foodborne intracellular bacterial pathogen leading to human listeriosis. Despite a high mortality rate and increasing antibiotic resistance no clinically approved vaccine against Listeria is available. Attenuated Listeria strains offer protection and are tested as antitumor vaccine vectors, but would benefit from a better knowledge on immunodominant vector antigens. To identify novel antigens, we screen for Listeria peptides presented on the surface of infected human cell lines by mass spectrometry-based immunopeptidomics. In between more than 15,000 human self-peptides, we detect 68 Listeria immunopeptides from 42 different bacterial proteins, including several known antigens. Peptides presented on different cell lines are often derived from the same bacterial surface proteins, classifying these antigens as potential vaccine candidates. Encoding these highly presented antigens in lipid nanoparticle mRNA vaccine formulations results in specific CD8 + T-cell responses and induces protection in vaccination challenge experiments in mice. Our results can serve as a starting point for the development of a clinical mRNA vaccine against Listeria and aid to improve attenuated Listeria vaccines and vectors, demonstrating the power of immunopeptidomics for next-generation bacterial vaccine development.
ABSTRACTListeria monocytogenesis a foodborne intracellular bacterial pathogen leading to human listeriosis. Despite a high mortality rate and increasing antibiotic resistance no clinically approved vaccine againstListeriais available. AttenuatedListeriastrains offer protection and are tested as antitumor vaccine vectors, but would benefit from a better knowledge on immunodominant vector antigens. To identify novel antigens, we screened forListeriaepitopes presented on the surface of infected human cell lines by mass spectrometry-based immunopeptidomics. In between more than 15,000 human self-peptides, we detected 68Listeriaepitopes from 42 different bacterial proteins, including several known antigens. Peptide epitopes presented on different cell lines were often derived from the same bacterial surface proteins, classifying these antigens as potential vaccine candidates. Encoding these highly presented antigens in lipid nanoparticle mRNA vaccine formulations resulted in specific CD8+ T-cell responses and high levels of protection in vaccination challenge experiments in mice. Our results pave the way for the development of a clinical mRNA vaccine againstListeriaand aid to improve attenuatedListeriavaccines and vectors, demonstrating the power of immunopeptidomics for next-generation bacterial vaccine development.
Angiotensin-converting enzyme 2 (ACE2) and accessory proteases (TMPRSS2 and CTSL) are needed for severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) cellular entry, and their expression may shed light on viral tropism and impact across the body. We assessed the cell-type-specific expression of ACE2, TMPRSS2 and CTSL across 107 single-cell RNA-sequencing studies from different tissues. ACE2, TMPRSS2 and CTSL are coexpressed in specific subsets of respiratory epithelial cells in the nasal passages, airways and alveoli, and in cells from other organs associated with coronavirus disease 2019 (COVID-19) transmission or pathology. We performed a meta-analysis of 31 lung single-cell RNA-sequencing studies with 1,320,896 cells from 377 nasal, airway and lung parenchyma samples from 228 individuals. This revealed cell-type-specific associations of age, sex and smoking with expression levels of ACE2, TMPRSS2 and CTSL. Expression of entry factors increased with age and in males, including in airway secretory cells and alveolar type 2 cells. Expression programs shared by ACE2+TMPRSS2+ cells in nasal, lung and gut tissues included genes that may mediate viral entry, key immune functions and epithelial–macrophage cross-talk, such as genes involved in the interleukin-6, interleukin-1, tumor necrosis factor and complement pathways. Cell-type-specific expression patterns may contribute to the pathogenesis of COVID-19, and our work highlights putative molecular pathways for therapeutic intervention. An integrated analysis of over 100 single-cell and single-nucleus transcriptomics studies illustrates severe acute respiratory syndrome coronavirus 2 viral entry gene coexpression patterns across different human tissues, and shows association of age, smoking status and sex with viral entry gene expression in respiratory cell populations.
The intestinal microbiota modulates host physiology and gene expression via mechanisms that are not fully understood. Here we examine whether host epitranscriptomic marks are affected by the gut microbiota. We use methylated RNA-immunoprecipitation and sequencing (MeRIP-seq) to identify N6-methyladenosine (m(6)A) modifications in mRNA of mice carrying conventional, modified, or no microbiota. We find that variations in the gut microbiota correlate with m(6)A modifications in the cecum, and to a lesser extent in the liver, affecting pathways related to metabolism, inflammation and antimicrobial responses. We analyze expression levels of several known writer and eraser enzymes, and find that the methyltransferase Mettl16 is downregulated in absence of a microbiota, and one of its target mRNAs, encoding S-adenosylmethionine synthase Mat2a, is less methylated. We furthermore show that Akkermansia muciniphila and Lactobacillus plantarum affect specific m(6)A modifications in mono-associated mice. Our results highlight epitranscriptomic modifications as an additional level of interaction between commensal bacteria and their host.
We investigated SARS-CoV-2 tropism by surveying expression of viral entry-associated genes in single-cell RNA-seq data from human tissues. We co-detected these transcripts in specific respiratory, corneal, and intestinal epithelial cells, potentially explaining the high efficiency of SARS-CoV-2 transmission. These genes are co-expressed in nasal epithelial cells with those involved in innate immunity, highlighting their potential roles in initial viral infection, spread and clearance. The data are available online at covid19cellatlas.org.
ABSTRACT The COVID-19 pandemic, caused by the novel coronavirus SARS-CoV-2, creates an urgent need for identifying molecular mechanisms that mediate viral entry, propagation, and tissue pathology. Cell membrane bound angiotensin-converting enzyme 2 (ACE2) and associated proteases, transmembrane protease serine 2 (TMPRSS2) and Cathepsin L (CTSL), were previously identified as mediators of SARS-CoV2 cellular entry. Here, we assess the cell type-specific RNA expression of ACE2 , TMPRSS2 , and CTSL through an integrated analysis of 107 single-cell and single-nucleus RNA-Seq studies, including 22 lung and airways datasets (16 unpublished), and 85 datasets from other diverse organs. Joint expression of ACE2 and the accessory proteases identifies specific subsets of respiratory epithelial cells as putative targets of viral infection in the nasal passages, airways, and alveoli. Cells that co-express ACE2 and proteases are also identified in cells from other organs, some of which have been associated with COVID-19 transmission or pathology, including gut enterocytes, corneal epithelial cells, cardiomyocytes, heart pericytes, olfactory sustentacular cells, and renal epithelial cells. Performing the first meta-analyses of scRNA-seq studies, we analyzed 1,176,683 cells from 282 nasal, airway, and lung parenchyma samples from 164 donors spanning fetal, childhood, adult, and elderly age groups, associate increased levels of ACE2 , TMPRSS2 , and CTSL in specific cell types with increasing age, male gender, and smoking, all of which are epidemiologically linked to COVID-19 susceptibility and outcomes. Notably, there was a particularly low expression of ACE2 in the few young pediatric samples in the analysis. Further analysis reveals a gene expression program shared by ACE2 + TMPRSS2 + cells in nasal, lung and gut tissues, including genes that may mediate viral entry, subtend key immune functions, and mediate epithelial-macrophage cross-talk. Amongst these are IL6, its receptor and co-receptor, IL1R , TNF response pathways, and complement genes. Cell type specificity in the lung and airways and smoking effects were conserved in mice. Our analyses suggest that differences in the cell type-specific expression of mediators of SARS-CoV-2 viral entry may be responsible for aspects of COVID-19 epidemiology and clinical course, and point to putative molecular pathways involved in disease susceptibility and pathogenesis.
SUMMARY To face up to the exponential growth of heterogeneous datasets of various organisms, we developed a user-friendly platform for building multi-omics websites, which is named Bacnet. This platform helps bioinformaticians to construct four key web interfaces: (i) an interactive genome viewer; (ii) an expression and protein atlas; (iii) an interface for analysis of co-expression network; (iv) an interface for exploring homolog presence. We believe our platform will help the bioinformaticians to construct personalized user interfaces dedicated to biologists studying non-reference organisms. AVAILABILITY AND IMPLEMENTATION https://github.com/becavin-lab/bacnet; Java;Eclipse RAP;Eclipse RCP.
Rationale The respiratory tract constitutes an elaborated line of defense based on a unique cellular ecosystem. Single-cell profiling methods enable the investigation of cell population distributions and transcriptional changes along the airways. Methods We have explored cellular heterogeneity of the human airway epithelium in 10 healthy living volunteers by single-cell RNA profiling. 77,969 cells were collected by bronchoscopy at 35 distinct locations, from the nose to the 12th division of the airway tree. Results The resulting atlas is composed of a high percentage of epithelial cells (89.1%), but also immune (6.2%) and stromal (4.7%) cells with peculiar cellular proportions in different sites of the airways. It reveals differential gene expression between identical cell types (suprabasal, secretory, and multiciliated cells) from the nose (MUC4, PI3, SIX3) and tracheobronchial (SCGB1A1, TFF3) airways. By contrast, cell-type specific gene expression was stable across all tracheobronchial samples. Our atlas improves the description of ionocytes, pulmonary neuro-endocrine (PNEC) and brush cells, which are likely derived from a common population of precursor cells. We also report a population of KRT13 positive cells with a high percentage of dividing cells which are reminiscent of “hillock” cells previously described in mouse. Conclusions Robust characterization of this unprecedented large single-cell cohort establishes an important resource for future investigations. The precise description of the continuum existing from nasal epithelium to successive divisions of lung airways and the stable gene expression profile of these regions better defines conditions under which relevant tracheobronchial proxies of human respiratory diseases can be developed.
We investigated SARS-CoV-2 potential tropism by surveying expression of viral entry-associated genes in single-cell RNA-sequencing data from multiple tissues from healthy human donors. We co-detected these transcripts in specific respiratory, corneal and intestinal epithelial cells, potentially explaining the high efficiency of SARS-CoV-2 transmission. These genes are co-expressed in nasal epithelial cells with genes involved in innate immunity, highlighting the cells' potential role in initial viral infection, spread and clearance. The study offers a useful resource for further lines of inquiry with valuable clinical samples from COVID-19 patients and we provide our data in a comprehensive, open and user-friendly fashion at www.covid19cellatlas.org.
The SARS-CoV-2 coronavirus, the etiologic agent responsible for COVID-19 coronavirus disease, is a global threat. To better understand viral tropism, we assessed the RNA expression of the coronavirus receptor, ACE2, as well as the viral S protein priming protease TMPRSS2 thought to govern viral entry in single-cell RNA-sequencing (scRNA-seq) datasets from healthy individuals generated by the Human Cell Atlas consortium. We found that ACE2, as well as the protease TMPRSS2, are differentially expressed in respiratory and gut epithelial cells. In-depth analysis of epithelial cells in the respiratory tree reveals that nasal epithelial cells, specifically goblet/secretory cells and ciliated cells, display the highest ACE2 expression of all the epithelial cells analyzed. The skewed expression of viral receptors/entry-associated proteins towards the upper airway may be correlated with enhanced transmissivity. Finally, we showed that many of the top genes associated with ACE2 airway epithelial expression are innate immune-associated, antiviral genes, highly enriched in the nasal epithelial cells. This association with immune pathways might have clinical implications for the course of infection and viral pathology, and highlights the specific significance of nasal epithelia in viral infection. Our findings underscore the importance of the availability of the Human Cell Atlas as a reference dataset. In this instance, analysis of the compendium of data points to a particularly relevant role for nasal goblet and ciliated cells as early viral targets and potential reservoirs of SARS-CoV-2 infection. This, in turn, serves as a biological framework for dissecting viral transmission and developing clinical strategies for prevention and therapy.
Summary Recent studies have reported on the presence of bacterial RNA within or outside extracellular membrane vesicles, possibly as ribonucleoprotein complexes. Proteins that bind and stabilize bacterial RNAs in the extracellular environment have not been reported. Here, we show that the bacterial pathogen Listeria monocytogenes secretes a small RNA binding protein that we named Zea. We show that Zea binds and stabilizes a subset of L. monocytogenes RNAs causing their accumulation in the extracellular medium. Furthermore, Zea binds RIG-I, the vertebrate non-self-RNA innate immunity sensor and potentiates RIG-I-signaling leading to interferon β production. By performing in vivo infection, we finally show that Zea modulates L. monocytogenes virulence. Together, this study reveals that bacterial extracellular RNAs and RNA binding proteins can affect the host-pathogen crosstalk.
RNA binding proteins (RBPs) perform key cellular activities by controlling the function of bound RNAs. The widely held assumption that RBPs are strictly intracellular has been challenged by the discovery of secreted RBPs. While extracellular RBPs have been described in mammals, secreted bacterial RBPs have not been reported. Here, we show that the bacterial pathogen L. monocytogenes secretes a small RBP that we named Zea. We show that Zea binds a subset of L. monocytogenes RNAs causing their accumulation in the extracellular medium. Furthermore, during L. monocytogenes infection, Zea binds RIG-I, the non-self-RNA innate immunity sensor, potentiating interferon β production. Mouse infection studies revealed that Zea modulates L. monocytogenes virulence. Together, this study uncovered the presence of an extracellular ribonucleoprotein complex from bacteria and its involvement in host-pathogen crosstalk
RNA-binding proteins (RBPs) perform key cellular activities by controlling the function of bound RNAs. The widely held assumption that RBPs are strictly intracellular has been challenged by the discovery of secreted RBPs. However, extracellular RBPs have been described in eukaryotes, while secreted bacterial RBPs have not been reported. Here, we show that the bacterial pathogen Listeria monocytogenes secretes a small RBP that we named Zea. We show that Zea binds a subset of L. monocytogenes RNAs, causing their accumulation in the extracellular medium. Furthermore, during L. monocytogenes infection, Zea binds RIG-I, the non-self-RNA innate immunity sensor, potentiating interferon-β production. Mouse infection studies reveal that Zea affects L. monocytogenes virulence. Together, our results unveil that bacterial RNAs can be present extracellularly in association with RBPs, acting as "social RNAs" to trigger a host response during infection.
This Article contains a URL for a publically available whole-genome browser ( http://nterm.listeriomics.pasteur.fr ). However, due to technical constraint, this website has been replaced with an alternative ( https://listeriomics.pasteur.fr ).
The main outcome of efficient CRISPR-Cas9 cleavage in the chromosome of bacteria is cell death. This can be conveniently used to eliminate specific genotypes from a mixed population of bacteria, which can be achieved both in vitro , e.g. to select mutants, or in vivo as an antimicrobial strategy. The efficiency with which Cas9 kills bacteria has been observed to be quite variable depending on the specific target sequence, but little is known about the sequence determinants and mechanisms involved. Here we performed a genome-wide screen of Cas9 cleavage in the chromosome of E. coli to determine the efficiency with which each guide RNA kills the cell. Surprisingly we observed a large-scale pattern where guides targeting some regions of the chromosome are more rapidly depleted than others. Unexpectedly, this pattern arises from the influence of degrading specific chromosomal regions on the copy number of the plasmid carrying the guide RNA library. After taking this effect into account, it is possible to train a neural network to predict Cas9 efficiency based on the target sequence. We show that our model learns different features than previous models trained on Eukaryotic CRISPR-Cas9 knockout libraries. Our results highlight the need for specific models to design efficient CRISPR-Cas9 tools in bacteria.
High-throughput genetic screens are powerful methods to identify genes linked to a given phenotype. The catalytic null mutant of the Cas9 RNA-guided nuclease (dCas9) can be conveniently used to silence genes of interest in a method also known as CRISPRi. Here, we report a genome-wide CRISPR-dCas9 screen using a starting pool of ~ 92,000 sgRNAs which target random positions in the chromosome of E. coli. To benchmark our method, we first investigate its utility to predict gene essentiality in the genome of E. coli during growth in rich medium. We could identify 79% of the genes previously reported as essential and demonstrate the non-essentiality of some genes annotated as essential. In addition, we took advantage of the intermediate repression levels obtained when targeting the template strand of genes to show that cells are very sensitive to the expression level of a limited set of essential genes. Our data can be visualized on CRISPRbrowser, a custom web interface available at crispr.pasteur.fr. We then apply the screen to discover E. coli genes required by phages λ, T4 and 186 to kill their host, highlighting the involvement of diverse host pathways in the infection process of the three tested phages. We also identify colanic acid capsule synthesis as a shared resistance mechanism to all three phages. Finally, using a plasmid packaging system and a transduction assay, we identify genes required for the formation of functional λ capsids, thus covering the entire phage cycle. This study demonstrates the usefulness and convenience of pooled genome-wide CRISPR-dCas9 screens in bacteria and paves the way for their broader use as a powerful tool in bacterial genomics.