Despite the promise of single-cell transcriptomics for understanding cell states in heterogeneous populations, widely used platforms have limited ability to link transcriptional states to somatic mutations within the same cells. Here, we introduce Genotyping in Fixed Transcriptomes (GIFT) for the simultaneous detection of large numbers of targeted genetic variants with whole transcriptome profiles in single cells. The core innovation of GIFT is a rationally designed gapfilling reaction between adjacent single-stranded DNA (ssDNA) probes that barcodes native transcript sequence to enable highly-specific targeted mutation detection. GIFT achieves greater than 99% genotyping accuracy and flexible capture of hundreds of mutations per cell, including in formalin-fixed, paraffin-embedded (FFPE) tissue, enabling clonal lineage tracing in heterogeneous settings. We demonstrate the unique scalability of GIFT by profiling more than 700,000 cells from 35 donors with myeloproliferative neoplasms (MPN), revealing mutation-dependent hematopoietic responses to systemic inflammation associated with the characteristic JAK2V617 mutation, including an allelic dose gradient of interferon-associated transcriptional programs and priming of hematopoietic stem cells that develop into divergent disease states. The technical advantages of GIFT enable direct resolution of genotype-to-phenotype relationships via clonal tracing with comprehensive cell-state measurements at single-cell resolution.
Technologies to study localized host-pathogen interactions are urgently needed. Here, we present a spatial transcriptomics approach to simultaneously capture host and pathogen transcriptome-wide spatial gene expression information from human formalin-fixed paraffin-embedded (FFPE) tissue sections at a near single-cell resolution. We demonstrate this methodology in lung samples from COVID-19 patients and validate our spatial detection of SARS-CoV-2 against RNAScope and in situ sequencing. Host-pathogen colocalization analysis identified putative modulators of SARS-CoV-2 infection in human lung cells. Our approach provides new insights into host response to pathogen infection through the simultaneous, unbiased detection of two transcriptomes in FFPE samples.
To advance our understanding of cellular host-pathogen interactions, technologies that facilitate the co-capture of both host and pathogen spatial transcriptome information are needed. Here, we present an approach to simultaneously capture host and pathogen spatial gene expression information from the same formalin-fixed paraffin embedded (FFPE) tissue section using the spatial transcriptomics technology. We applied the method to COVID-19 patient lung samples and enabled the dual detection of human and SARS-CoV-2 transcriptomes at 55 μm resolution. We validated our spatial detection of SARS-CoV-2 and identified an average specificity of 94.92% in comparison to RNAScope and 82.20% in comparison to in situ sequencing (ISS). COVID-19 tissues showed an upregulation of host immune response, such as increased expression of inflammatory cytokines, lymphocyte and fibroblast markers. Our colocalization analysis revealed that SARS-CoV-2+ spots presented shifts in host RNA metabolism, autophagy, NFκB, and interferon response pathways. Future applications of our approach will enable new insights into host response to pathogen infection through the simultaneous, unbiased detection of two transcriptomes.### Competing Interest StatementH.S., Y.M., S.G. are scientific advisors to 10x Genomics, Inc. that holds IP rights to the ST technology and previously acquired ReadCoor and Cartana and their accompanying intellectual property rights. S.G. holds 10x Genomics stocks. H.S., Y.M., E.B., A.H. and S.G. are co-inventors on patent filings relating to this work. E.B., A.H., A.J. and A.N. are employees of 10x Genomics and hold stock options. All other authors declare no competing interests.
The CD8(+) T cell response to an antigen is composed of many T cell clones with unique T cell receptors, together forming a heterogeneous repertoire of effector and memory cells. How individual T cell clones contribute to this heterogeneity throughout immune responses remains largely unknown. In this study, we longitudinally track human CD8(+) T cell clones expanding in response to yellow fever virus (YFV) vaccination at the single-cell level. We observed a drop in clonal diversity in blood from the acute to memory phase, suggesting that clonal selection shapes the circulating memory repertoire. Clones in the memory phase display biased differentiation trajectories along a gradient from stem cell to terminally differentiated effector memory fates. In secondary responses, YFV- and influenza-specific CD8(+) T cell clones are poised to recapitulate skewed differentiation trajectories. Collectively, we show that the sum of distinct clonal phenotypes results in the multifaceted human T cell response to acute viral infections.
The future of human genomics is one that seeks to resolve the entirety of genetic variation through sequencing. The prospect of utilizing genomics for medical purposes require cost-efficient and accurate base calling, long-range haplotyping capability, and reliable calling of structural variants. Short-read sequencing has lead the development towards such a future but has struggled to meet the latter two of these needs. To address this limitation, we developed a technology that preserves the molecular origin of short sequencing reads, with an insignificant increase to sequencing costs. We demonstrate a novel library preparation method for high throughput barcoding of short reads where millions of random barcodes can be used to reconstruct megabase-scale phase blocks.
Accurate variant calling and genotyping represent major limiting factors for downstream applications of single-cell genomics. Here, we report Conbase for the identification of somatic mutations in single-cell DNA sequencing data. Conbase leverages phased read data from multiple samples in a dataset to achieve increased confidence in somatic variant calls and genotype predictions. Comparing the performance of Conbase to three other methods, we find that Conbase performs best in terms of false discovery rate and specificity and provides superior robustness on simulated data, in vitro expanded fibroblasts and clonal lymphocyte populations isolated directly from a healthy human donor.
CD8+ T cells play essential roles in immunity to viral and bacterial infections, and to guard against malignant cells. The CD8+ T cell response to an antigen is composed of many T cell clones with unique T cell receptors, together forming a heterogenous repertoire of phenotypically and functionally distinct effector and memory cells[1][1], [2][2]. How individual T cell clones contribute to this heterogeneity during an immune response is key to understand immunity but remains largely unknown. Here, we longitudinally tracked CD8+ T cell clones expanding in response to yellow fever virus vaccination at the single cell level in humans. We show that only a fraction of the clones detected in the acute response persists as circulating memory T cells, indicative of clonal selection. Clones persisting in the memory phase displayed biased differentiation trajectories along a gradient of stem cell memory (SCM) towards terminally differentiated effector memory (EMRA) fates. Reactivation of single memory CD8+ T cells revealed that they were poised to recapitulate skewed differentiation trajectories in secondary responses, and this was generalizable across individuals for both yellow fever and influenza virus. Together, we show that the sum of distinct clonal differentiation repertoires results in the multifaceted T cell response to acute viral infections in humans. [1]: #ref-1 [2]: #ref-2
Spatial Transcriptomics has been shown to be a persuasive RNA sequencingtechnology for analyzing cellular heterogeneity within tissue sections. Thetechnology efficiently captures and barcodes 3’ ta ...
ABSTRACT The future of human genomics is one that seeks to resolve the entirety of genetic variation through sequencing. The prospect of utilizing genomics for medical purposes require cost-efficient and accurate base calling, long-range haplotyping capability, and reliable calling of structural variants. Short read sequencing has lead the development towards such a future but has struggled to meet the latter two of these needs 1 . To address this limitation, we developed a technology that preserves the molecular origin of short sequencing reads, with an insignificant increase to sequencing costs. We demonstrate a novel library preparation method which enables whole genome haplotyping, long-range phasing of single DNA molecules, and de novo genome assembly through barcode-linked reads (BLR). Millions of random barcodes are used to reconstruct megabase-scale phase blocks and call structural variants. We also highlight the versatility of our technology by generating libraries from different organisms using only picograms to nanograms of input material.
Here we report the development of Conbase, a software application for the identification of somatic mutations in single cell DNA sequencing data with high rates of allelic dropout and at low read depth. Conbase leverages data from multiple samples in a dataset and utilizes read phasing to call somatic single nucleotide variants and to accurately predict genotypes in whole genome amplified single cells in somatic variant loci. We demonstrate the accuracy of Conbase on simulated datasets, in vitro expanded fibroblasts and clonally in vivo expanded lymphocyte populations isolated directly from a healthy human donor.
Data produced with short-read sequencing technologies result in ambiguous haplotyping and a limited capacity to investigate the full repertoire of biologically relevant forms of genetic variation. The notion of haplotype-resolved sequencing data has recently gained traction to reduce this unwanted ambiguity and enable exploration of other forms of genetic variation; beyond studies of just nucleotide polymorphisms, such as compound heterozygosity and structural variations. Here we describe Droplet Barcode Sequencing, a novel approach for creating linked-read sequencing libraries by uniquely barcoding the information within single DNA molecules in emulsion droplets, without the aid of specialty reagents or microfluidic devices. Barcode generation and template amplification is performed simultaneously in a single enzymatic reaction, greatly simplifying the workflow and minimizing assay costs compared to alternative approaches. The method has been applied to phase multiple loci targeting all exons of the highly variable Human Leukocyte Antigen A (HLA-A) gene, with DNA from eight individuals present in the same assay. Barcode-based clustering of sequencing reads confirmed analysis of over 2000 independently assayed template molecules, with an average of 753 reads in support of called polymorphisms. Our results show unequivocal characterization of all alleles present, validated by correspondence against confirmed HLA database entries and haplotyping results from previous studies.
BackgroundWhole genome amplification (WGA) is currently a prerequisite for single cell whole genome or exome sequencing. Depending on the method used the rate of artifact formation, allelic dropout and sequence coverage over the genome may differ significantly.ResultsThe largest difference between the evaluated protocols was observed when analyzing the target coverage and read depth distribution. These differences also had impact on the downstream variant calling. Conclusively, the products from the AMPLI1 and MALBAC kits were shown to be most similar to the bulk samples and are therefore recommended for WGA of single cells.DiscussionIn this study four commercial kits for WGA (AMPLI1, MALBAC, Repli-G and PicoPlex) were used to amplify human single cells. The WGA products were exome sequenced together with non-amplified bulk samples from the same source. The resulting data was evaluated in terms of genomic coverage, allelic dropout and SNP calling.
The most widely used method for the preservation of clinical tissue specimens is formalin fixation and paraffin embedding (FFPE). Simultaneous analysis of RNA and DNA from samples preserved using t ...
The work presented in this thesis describes methodologies developed for integration and accurate interpretation of barcoded DNA, to empower large-scale-omics analysis. The objectives mainly aim at enabling multiplexed proteomic measurements in high-throughput format through DNA barcoding and massive parallel sequencing. The thesis is based on four scientific papers that focus on three main criteria; (i) to prepare reagents for large-scale affinity-proteomics, (ii) to present technical advances in barcoding systems for parallel protein detection, and (iii) address challenges in complex sequencing data analysis.In the first part, bio-conjugation of antibodies is assessed at significantly downscaled reagent quantities. This allows for selection of affinity binders without restrictions to accessibility in large amounts and purity from amine-containing buffers or stabilizer materials (Paper I). This is followed by DNA barcoding of antibodies using minimal reagent quantities. The procedure additionally enables efficient purification of barcoded antibodies from free remaining DNA residues to improve sensitivity and accuracy of the subsequent measurements (Paper II). By utilizing a solid-phase approach on magnetic beads, a high-throughput set-up is ready to be facilitated by automation. Subsequently, the applicability of prepared bio-conjugates for parallel protein detection is demonstrated in different types of standard immunoassays (Papers I and II).As the second part, the method immuno-sequencing (I-Seq) is presented for DNAmediated protein detection using barcoded antibodies. I-Seq achieved the detection of clinically relevant proteins in human blood plasma by parallel DNA readout (Paper II). The methodology is further developed to track antibody-antigen interaction events on suspension bead arrays, while being encapsulated in barcoded emulsion droplets (Paper III). The method, denoted compartmentalized immuno-sequencing (cI-Seq), is potent to perform specific detections with paired antibodies and can provide information on details of joint recognition events.Recent progress in technical developments of DNA sequencing has increased the interest in large-scale studies to analyze higher number of samples in parallel. The third part of this thesis focuses on addressing challenges of large-scale sequencing analysis. Decoding of a huge DNA-barcoded data is presented, aiming at phase-defined sequence investigation of canine MHC loci in over 3000 samples (Paper IV). The analysis revealed new single nucleotide variations and a notable number of novel haplotypes for the 2nd exon of DLA DRB1.Taken together, this thesis demonstrates emerging applications of barcoded sequencing in protein and DNA detection. Improvements through the barcoding systems for assay parallelization, de-convolution of antigen-antibody interactions, sequence variant analysis, as well as large-scale data interpretation would aid biomedical studies to achieve a deeper understanding of biological processes. The future perspectives of the developed methodologies may therefore stem for advancing large-scale omics investigations, particularly in the promising field of DNA-mediated proteomics, for highly multiplex studies of numerous samples at a notably improved molecular resolution.
Because human white adipocytes display a high turnover throughout adulthood, a continuous supply of precursor cells is required to maintain adipogenesis. Bone marrow (BM)-derived progenitor cells may contribute to mammalian adipogenesis; however, results in animal models are conflicting. Here we demonstrate in 65 subjects who underwent allogeneic BM or peripheral blood stem cell (PBSC) transplantation that, over the entire lifespan, BM/PBSC-derived progenitor cells contribute ∼10% to the subcutaneous adipocyte population. While this is independent of gender, age, and different transplantation-related parameters, body fat mass exerts a strong influence, with up to 2.5-fold increased donor cell contribution in obese individuals. Exome and whole-genome sequencing of single adipocytes suggests that BM/PBSC-derived progenitors contribute to adipose tissue via both differentiation and cell fusion. Thus, at least in the setting of transplantation, BM serves as a reservoir for adipocyte progenitors, particularly in obese subjects.
High-throughput sequencing platforms mainly produce short-read data, resulting in a loss of phasing information for many of the genetic variants analysed. For certain applications, it is vital to know which variant alleles are connected to each individual DNA molecule. Here we demonstrate a method for massively parallel barcoding and phasing of single DNA molecules. First, a primer library with millions of uniquely barcoded beads is generated. When compartmentalized with single DNA molecules, the beads can be used to amplify and tag any target sequences of interest, enabling coupling of the biological information from multiple loci. We apply the assay to bacterial 16S sequencing and up to 94% of the hypothesized phasing events are shown to originate from single molecules. The method enables use of widely available short-read-sequencing platforms to study long single molecules within a complex sample, without losing phase information.
Polyguanine nucleotide repeats exhibit a greater degree of variation than the average for the genome as a whole. This is partly due to polymerase slippage that causes insertions or deletions in the ...