Abstract Recent advancements in precision medicine, while highly promising, presents a major technical challenge to researchers due to disease heterogeneity. The emergence of single-cell technologies has greatly refined the resolution in which sample diversity can be investigated, enhancing the efficiency of selecting appropriate molecular targets. Additionally, applying multi-omics analysis on single cells would further improve the understanding of cell-to-cell heterogeneity by providing unique insights on cellular and genetic composition. Using a two-step droplet microfluidic technology, the Mission Bio Tapestri Platform enables multiplex-PCR based high-throughput targeted DNA sequencing in single cells to obtain single-nucleotide variation (SNV) and copy number variation (CNV) information. By leveraging this technology, a new workflow is developed to detect protein expression in addition to DNA genotype in the same single cells. In this approach, cells are labeled with a pool of oligonucleotide-conjugated antibodies prior to loading the cells into the Tapestri Instrument. Sequencing libraries are then prepared from both antibody oligonucleotides and the amplified DNA sequences, followed by identification of single-cell DNA genotypes and protein signatures from the sequencing readout. The number of protein targets can be in the range from 6 to over 40, which is beyond the limit for a single flow cytometry run. This method has been successfully performed on cell lines, fresh and frozen PBMCs, as well as clinical samples. In an acute myeloid leukemia (AML) sample, combined single-cell SNV, CNV, and protein expression data illustrated the heterogeneity within the sample. The data clearly identified CD3+ T cells and CD19+ B cells without pathogenic SNVs and CNVs. CD34hiCD11blo and CD34loCD11bhi subpopulations were also identified within the cells carrying the same pathogenic SNVs and CNVs. We believe that this novel multi-omics technology will facilitate new discoveries in the complex relationship between genotype and phenotype, enable a better understanding of disease biology, and subsequently improve the design of diagnostics and therapies. Citation Format: Aik Ooi, Dalia Dhingra, Adam Sciambi, Kate Thompson, Jacqueline Marin, Saurabh Parikh, Mani Manivannan, David Ruff. Single-cell multi-omics analysis of SNV, CNV, and protein expression [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 2259.
Abstract Single cell multiomics assays targeting RNA and Protein from the same cell provide a high-resolution view of the heterogeneity of the sample. However both analytes target the phenotype of the cells and unambiguous inference that a cellular phenotype is caused by a genotype can only be achieved by their measurement from the same single cell. To address this gap, we have developed the Tapestri multi-omics workflow to analyze the DNA and Protein information from the same cell. After pre-processing the reads, cell calling is performed using DNA reads. Variant calling and filtering is carried out using DNA reads to identify the high quality genetic variants within each cell and the variant-cell matrix is then analyzed to identify clones. The Protein reads for the valid cells are log normalized. A systemic artefact resulting in random uniform amplification of antibodies is corrected for by learning the scaling factor from the read count distribution. This reduces the number of false positives. Furthermore read depth dependence of expression is corrected by identifying the correlation of the mean expression and background counts with the total reads in the cell. The mean signal and background for each cell is learned using a 2 component gaussian mixture model. Then, z-score normalization is applied to prevent high expressors from skewing further analysis. The scaled values are dimensionally reduced using PCA. A graph structure is created using KNN with weights calculated using Jaccard similarity following which community detection is employed to identify the cell types. We test this method on a model system with two donor PBMC, 2 cell lines mixture titrated at a 47:47:5:1 ratio with a 312 amplicon DNA panel and 45 plex antibody panel. We filter cells using isotype controls from the antibody panel and mixed cell signatures using the DNA panel. We can identify 4 clones using the genetic variants. We can identify multiple cell types including the major populations such as T-cells, B-cells, Monocytes and NK-cells using the phenotypic expression in the PBMC donors. Moreover, we are able to correlate the proportions of each PBMC cell type for the identified clones with that of the individual donors. Due to the corrections applied during the normalization of protein counts we are also able to accurately identify various expected expression profiles in PBMCs such as CD45RA-CD45RO and CD8-CD4 mutual exclusion, and the trimodal expression of CD4. Citation Format: Saurabh Parikh, Manimozhi Manivannan, Jacqueline Marin, Kate Thompson, Daniel Mendoza, Benjamin Schroeder, Saurabh Gulati, Shu Wang, Aik Ooi. Method to analyze mutational and phenotypic profiles from single cell for clonal evolution [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 246.
Abstract Background: It is now possible to interrogate thousands of cells in a single experiment for studying genetic variability with the advancements in single-cell sequencing technologies. Single-cell DNA platforms like Tapestri is still susceptible to errors from polymerase incorporations, structure induced template switching, PCR mediated recombination in Tapestri workflow or DNA-damage. Errors from sequencing could propagate from cluster amplification, cycle sequencing or image analysis. All together these errors can be divided into substitutions, insertions and deletion errors and can range from 0.5% to 2% depending on the sequencer. This makes rare variant and minimal residual disease detection challenging. To address these challenges, we developed deep learning models for correcting the errors, reduce false-positive rates and predict true variants. Method: First we build a consensus sequence from several reads to predict the correct sequence. The initial layers learn the motifs and local sequence contexts in classifying the patterns. The output of this network is a probability distribution over possible bases and the prediction is the base with highest probability. The bases in the reads are subsequently corrected to the predicted base from the first step model. After error correcting the reads, we used the variants called by Genome Analysis Toolkit to feed into a multi-class classifier network. Our features consists of percent of cells mutated, and the different genotype features including depth, AF and quality of each variant in these cells. The truth labels are generated using tapestri instrument from multiple experiments with known bulk truth. We trained the network on over 200k cells from 13 samples and tested on a larger set of samples. Class imbalance was handled using upsampling the truth data. Our training samples include diverse samples from cell mixtures at various dilution uptill 0.1% and clinical samples processed through tapestri instrument and sequenced on a diverse set of sequencers including miseq and novaseq. Conclusion: To validate this method, we used two different targeted panels on a Latin square model system with known truth mutations. With our 2-step workflow using error correction and variant prediction model, we significantly improved our median PPV 2-3 fold at 0.5% LOD while maintaining the sensitivity. We are further optimizing the model by adding more training samples and feature optimization. Citation Format: Manimozhi Manivannan, Sombeet Sahu, Kim Dong, Shu Wang, Saurabh Gulati, Saurabh Parikh, Nigel Beard, Anup Parikh. Improvements in variant calling sensitivity and specificity in single-cell DNA sequencing using deep learning [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 861.
Abstract Background: While scRNA-seq helps in obtaining high-resolution views of single-cell heterogeneity through characterization of the functional state of cells, our understanding of the cellular properties and population architectures of heterogeneous tissues will be greatly advanced by the multi-omics investigation of single cells. Recently, several computational methods have been developed to integrate data from multiple different single cell experiments measuring individual analytes. But unambiguous inference that a cellular phenotype is caused by a genotype can only be achieved by their measurement from the same single cell. To address this gap, we have developed the Tapestri multi-omics workflow to analyze the RNA and DNA information from the same cell. Methods: After pre-processing of the reads, cell calling is done by identifying the good barcodes using both DNA and RNA reads. DNA reads are processed to identify genetic variants in each cell. The variant cell matrix is then filtered for data completeness to ensure only high-quality data is used in downstream processing. Ploidy is estimated using by normalizing DNA reads, genetic variants and ploidy information together is used to identify subclones. The RNA reads are log normalized and scaled within each cell. Next we set the mean expression of each transcript to 0 and scale the variance to 1. This avoids downstream analysis being skewed by high expressors. We then trained a random forest classifier to identify significantly differentially expressed transcripts across subclones which were identified using genetic variants and ploidy information. Results: Using the top differentially expressed transcripts we performed dimensionality reduction followed by clustering of cell types. The resulting visualization showed how well the genotypic and transcriptomic datasets integrated with one another. We tested this method on a model system with Raji and KG1 cell lines titrated at 50:50 ratio and were clearly able to associate the transcriptional variation with the genotypic variation of the 2 different cell lines. The method was also validated on a PBMC sample to ensure robustness of methods. We were able to identify the different cell types present the sample and were able to overlay that information with genetic variants to identify sub-clones in the identified cell types. Citation Format: Saurabh Gulati, Shu Wang, Saurabh Parikh, Ben Liu, Kaustubh Gokhale, Manimozhi Manivannan, Sombeet Sahu, Dong Kim, Anup Parikh. Using machine learning to detect heterogeneity in single cell multi-omics datasets [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 865.
Abstract Background: To realize the promise of precision medicine for cancer, assessing genetic variation present in rare cells and understanding the role that these cells play in the evolution of tumor progression is essential. High throughput single cell DNA targeted sequencing enables detection of rare mutations in cells and identification of subclones defined by co-occurrence of mutations. The big challenge with multiplex sequencing at single cell level is the non-uniform amplification of the targeted regions during PCR. This results in inadequate coverage of mutations of interest in the panel and hence makes genotyping challenging. To address this challenge, we developed a machine learning engine to optimize amplicon design for uniform amplification by making reliable performance prediction. Methods: Multiple panels with various sizes were designed with amplicons spanning a wide range of design properties such as amplicon GC, length, secondary structure prediction, primer specificity. These panels were synthesized and processed through Tapestri single cell DNA platform. The tested amplicons are classified into low-performer, OK-performer and high flyer based on their normalized reads-per-cell value. Design properties and property distribution of the amplicons and the panel are the features. We used random forest classifier to calculate feature importance and analyzed the range of the top features for each class and their significance of variance between classes. These ranges were then used as parameters in the assay design pipeline. Next, we train machine learning models with performance data to develop a performance prediction engine. Results: To test the performance of the design pipeline with new parameters, we designed a small (31), medium (128) and large (287) amplicon panel. Multiple runs were conducted for each panel with different cell types. We were able to achieve high panel performance of 97%, 92% and 88% across the three panels. The new parameters resulted in ~10-20% improvement in panel uniformity. We are working on further optimizing the performance prediction engine by using different ML classification models with K-fold cross validation, training using larger group of amplicons and optimizing features using combinations of properties. Citation Format: Shu Wang, Saurabh Gulati, Dong Kim, Sombeet Sahu, Saurabh Parikh, Nianzhen Li, Manimozhi Manivannan, Nigel Beard. Amplicon design algorithm for single cell targeted DNA sequencing using machine learning [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 2109.
Myeloid malignancies, including acute myeloid leukaemia (AML), arise from the expansion of haematopoietic stem and progenitor cells that acquire somatic mutations. Bulk molecular profiling has suggested that mutations are acquired in a stepwise fashion: mutant genes with high variant allele frequencies appear early in leukaemogenesis, and mutations with lower variant allele frequencies are thought to be acquired later1–3. Although bulk sequencing can provide information about leukaemia biology and prognosis, it cannot distinguish which mutations occur in the same clone(s), accurately measure clonal complexity, or definitively elucidate the order of mutations. To delineate the clonal framework of myeloid malignancies, we performed single-cell mutational profiling on 146 samples from 123 patients. Here we show that AML is dominated by a small number of clones, which frequently harbour co-occurring mutations in epigenetic regulators. Conversely, mutations in signalling genes often occur more than once in distinct subclones, consistent with increasing clonal diversity. We mapped clonal trajectories for each sample and uncovered combinations of mutations that synergized to promote clonal expansion and dominance. Finally, we combined protein expression with mutational analysis to map somatic genotype and clonal architecture with immunophenotype. Our findings provide insights into the pathogenesis of myeloid transformation and how clonal complexity evolves with disease progression. The evolution of myeloid malignancies is investigated using combined single-cell sequencing and immunophenotypic analysis.
With the advancements in single-cell sequencing technologies, it is now possible to interrogate thousands of cells in a single experiment for studying genetic variability. Single-Cell DNA platforms like Tapestri is susceptible to errors primarily from PCR and sequencing with rates ranging from 0.5% - 2%. This makes variant calling and minimal residual disease detection challenging. To address these challenges, we developed a novel consensus sequence-based method for correcting the errors, reduce false-positive rates and predict true variants. First, we build a consensus sequence from several reads to predict the correct sequence. The initial layers learn the motifs and local sequence contexts in classifying the patterns. The output of this network is a probability distribution over possible bases and the prediction is the base with highest probability. The bases in the reads are subsequently corrected to the predicted base from the first step model. After error correcting the reads, we used the variants called by Genome Analysis Toolkit to feed into a multi-class classifier network. Our features consist of percent of cells mutated, and the different genotype features including depth, AF and quality of each variant in these cells. The truth labels are generated using tapestri instrument from multiple experiments with known truth. We trained the network on over 200k cells from 13 samples and tested on a larger set of samples. Class imbalance was handled using upsampling the truth data. Our training samples include diverse samples from cell mixtures at various dilution uptill 0.1% and clinical samples processed through tapestri instrument and sequenced on a diverse set of sequencers including miseq and novaseq. With our 2-step error correction and variant prediction model, we significantly improved our median PPV 2-3 fold at 0.5% LOD. This will enable researchers in finding the rare subclone for characterizing MRD.
We validated this method on clinical samples and admixture samples with cell lines mixed at known ratios. CNV alone confidently detects subclones while when combined with mutational analysis, rare subclones of ∼1% prevalence was detected. Integration of CNVs and SNVs facilitates more accurate reconstruction of tumor evolution to better understand cancer progression mechanisms as well for quality control of gene edited cells, to further advance cancer research and therapy.
Abstract The TCGA, ICGC and other research consortiums has shown that an integral, multi-omic analysis of tumors is necessary in order to understand a more complete picture of the signaling pathways involved in development, evolution, prognosis and treatment resistance of tumors. However, in such studies, because different data modalities (“omics”) are analyzed separately, the information on how the co-existence of in the same cell is lacking. We report a new approach to measure multiple modalities simultaneously from up to 10,000 individual cells using high-throughput droplet microfluidics on the Tapestri platform, paired with next generation sequencing. Our triomic methodology evaluates targeted protein levels, mRNA transcript levels and somatic gDNA sequence variations (SNV & CNV) from the same cell. We employ oligonucleotide-conjugated antibody panels to probe cellular surface markers and targeted RNA and DNA amplification to resolve gene expression levels and genomic variants. Cell suspensions are first stained with cocktails of oligonucleotide-antibodies. These cells are loaded onto the Mission Bio Tapestri® microfluidic cartridge for generation of the first droplet. This droplet biochemistry allows for concurrent cell lysis and release of gDNA from chromatin scaffolding. A second droplet formation event brings together a barcoded bead and multiplex PCR amplification reagents. After amplification, the combined libraries are sequenced and demultiplexed bioinformatically yielding a multi-omics readout from the single-cells. Six protein surface markers, 35 transcripts and 88 genomic regions in 68 genes were interrogated with a 50:50 cell mixture of MCF7 and GM12878 cells. These targets play key roles in cellular migration, stemness, differentiation and proliferation. We intend to further explore the relationship between genotype-to-phenotype in breast carcinoma cell lines cultured in patient-derived scaffolds mimicking tumor microenvironments under different drug regimen conditions. We report the dynamic multimodal resolution and demonstrate utility of triomics in further illuminating cellular states and biological responses. Citation Format: Pedro Mendez, Dalia Dhingra, Aik Ooi, Shu Wang, Saurabh Gulati, Nigel Beard, Manimozhi Manivannan, Jens Björkman, Mikael Kubista, Göran Landberg, David Ruff, Anders Ståhlberg. Simultaneous DNA, RNA and protein analysis from single cells using a high-throughput microfluidic workflow for resolution of genotype-to-phenotype modalities [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 2506.
Summary Myeloid malignancies, including acute myeloid leukemia (AML), arise from the proliferation and expansion of hematopoietic stem and progenitor cells which acquire somatic mutations. Bulk molecular profiling studies on patient samples have suggested that somatic mutations are obtained in a step-wise fashion, where mutant genes with high variant allele frequencies (VAFs) are proposed to occur early in disease development and mutations with lower VAFs are thought to be acquired later in disease progression 1–3 . Although bulk sequencing informs leukemia biology and prognostication, it cannot distinguish which mutations occur in the same clone(s), accurately measure clonal complexity and clone size, or offer definitive evidence of mutational order. To elucidate the clonal framework of myeloid malignancies, we performed single cell mutational profiling on 146 samples from 123 patients. We found AML is most commonly comprised of a small number of dominant clones, which in many cases harbor co-occurring mutations in epigenetic regulators. Conversely, mutations in signaling genes often occur more than once in distinct subclones consistent with increasing clonal diversity. We also used these data to map the clonal trajectory of each patient and found that specific mutation combinations ( FLT3-ITD + NPM1 c ) synergize to promote clonal expansion and dominance. We combined cell surface protein expression with single cell mutational analysis to map somatic genotype and clonal architecture with immunophenotype. Our studies of clonal architecture at a single cell level provide novel insights into the pathogenesis of myeloid transformation and how clonal complexity contributes to disease progression.
Background: FMS-like tyrosine kinase 3 receptor-internal tandem duplication (FLT3-ITD) commonly occurs in one-quarter of patients with acute myeloid leukemia. Acute leukemia has a poor prognosis, mainly due to relapse. Single-Cell DNA sequencing technologies such as Tapestri platform enables us to understand the clonal heterogeneity of AML patient samples. Large indel calling is prone to errors from library preparation, sequencing biases, and algorithm artifacts. These errors contribute to false positives often in the form of multiple representations of the same variant. Here we present an algorithm to identify these large indels and reduce false positives to accurately measure the clonal heterogeneity and enable precision diagnostics. Methods The Tapestri analytical workflow involves obtaining raw reads from the sequencer, removing adapters, aligning and mapping the reads, calling individual cells and identifying genetic variants within each cell. We use a soft-clip based approach to detect the internal tandem duplications found in the FLT3 gene. The targeted panel has two amplicons targeting exons 14 and 15 in the FLT3 gene. The soft-clipped reads from these 2 amplicons are scanned for possible insertion events. The observed insertion event is qualified as an ITD variant if the total number of reads at the loci is greater than 10 and at least 20% of the reads support the insertion. The ITD variant is called homozygous if the allele frequency is greater than 0.9 and heterozygous otherwise. We then apply a generalized median string in Levenshtein space to collapse the different indel variants. The generalized median string is defined as a string that has the smallest sum of distances to the elements of a given set of strings. To do this, we first identify the candidate ITD size bins from the frequency peaks of all the called ITD variants and group the individual variants that are within 20bp boundaries of the frequency peaks into their respective bins. We project the ITD sequence strings within a bin on to Levenshtein vector space domain and calculate the median distance between all strings. We then use the string with the median distance to collapse the ITDs to the consensus sequence and report it in the vcf file. Results We processed AML samples with known FLT3 ITDs through Tapestri platform. We analyzed the raw data via Tapestri analytical workflow including our large indel and ITD detection algorithm. Using this method, we were able to accurately identify the ITDs and reproduce the true positive clones for the sample. We are currently optimizing this approach using different samples with a wide range of known ITDs.
Background Single-cell sequencing has the potential to provide unique insights on the cellular and genetic composition, drivers, and signatures of cancer at unparalleled sensitivity. Single-cell DNA sequencing has recently gained traction due to technologies like Tapestri Single-Cell DNA platform. Tapestri is a high-throughput single-cell DNA analysis platform that leverages droplet microfluidics and a multiplex-PCR based targeted DNA sequencing approach. Often times in droplet-based protocols, cell doublet is an issue where two cells get captured in the same droplet and result in spurious signals. This doublet rate is usually proportional to the total number of cells to the bead ratio. Hence many technologies try to keep this ratio low thereby inhibiting their ability to interrogate a higher number of cells. Also, signals from such droplets contribute to false positives in genotyping because reads from the doublets are misinterpreted as reads from a single cell. To address both these issues, we present for the first time a doublet identification tool based on deep neural networks. Methods We processed 60 samples that are cell line mixtures of Raji and K562 titrated at 50% each via Tapestri platform. We used 4 known loci that are genotypically distinct between the two cell lines to assign each cell to K562, RAJI or a doublet. We also used fluorescent images of cells to confirm the doublet rate. This data is used as ground truth for our classifier. We train our neural networks on a binary classification label using the ground truth on the amplicon-cell read distribution matrix. We set up a densely connected neural network classifier using tensorflow. The number of hidden layers in the classifier is equal to the number of amplicons in the targeted panel. Since our amplicon performance is based on amplicon competition in a single droplet we use a dense network instead of a convolution layer. We split the data from the 60 samples into training and test datasets. We train and test separately on multiple replicates of experiments. The hyperparameters were further optimized for low and high performing amplicons since the distribution of reads were noisier for low performing amplicons. To further improve accuracy we introduced artificial doublet by randomly sampling barcodes without replacement from the pool and adding to the pool. Results Preliminary results from this method look promising. We were able to detect ~97% of doublets confirmed by ground truth and remove them from downstream processing thereby reducing false positive clones. Introduction of artificial doublet also improved the accuracy of estimation. Further training of the neural network with more cells and hyper-parameter optimization is needed to improve accuracy.
Abstract Background: High throughput single-cell DNA sequencing allows for detection of rare mutations in cells and identification of subclones defined by zygosity and co-occurrence of mutations. This enables researchers to characterize tumor heterogeneity and progression which cannot be achieved by standard bulk sequencing. The big challenge with targeted sequencing is the non-uniform amplification of the targeted regions during PCR. This results in inadequate coverage of mutations of interest in the panel and hence makes genotyping challenging at the single-cell level. To address this challenge, we developed a machine learning engine to optimize amplicon design for uniform amplification by making reliable performance prediction. Methods: 10 targeted sequencing panels were designed with amplicons spanning a wide range of design properties such as amplicon GC%, length, secondary structure prediction, and primer specificity. These panels were synthesized and processed through Mission Bio TapestriⓇ Platform. We pre-processed the raw reads, mapped the reads, called cells and generated the amplicon-cell read matrix using the Tapestri Insights analytical pipeline. The tested amplicons are classified into low-performer, OK-performer and high-flyer based on their normalized reads-per-cell value. The design properties of the amplicon are the features. Highly correlated features were identified and pruned. We used random forest classifier to calculate feature importance. Top features were identified using two different feature selection methods. We then analyzed the range of the top features for each class and their significance of variance between classes. These ranges were then used as parameters in the assay design pipeline. Results: To test the performance of the design pipeline with new parameters, we designed a small (31), medium (128) and large (287) amplicon panel. Multiple runs were conducted for each panel with different cell types. We achieved high panel uniformity of 97%, 92% and 88% across the three panels. The new parameters resulted in ~10-20% improvement in panel uniformity. We are working on further optimizing the performance prediction engine by using different ML classification models with K-fold cross validation, training using larger group of amplicons, and optimizing features using combination of properties.
Background With the advancements in single-cell sequencing technologies, it is now possible to interrogate thousands of cells in a single experiment for studying genetic variability. Although Single-Cell DNA sequencing platforms like Tapestri has been readily available in the community, it is still susceptible to errors from polymerase incorporations, structure induced template switching, PCR mediated recombination in Tapestri workflow or DNA-damage. Errors from sequencing could propagate from cluster amplification, cycle sequencing or image analysis. All together these errors can be divided into substitutions, insertions and deletion errors and can range from 0.5% to 2% depending on the sequencer. This makes variant calling and minimal residual disease detection challenging. To address these challenges, we developed a novel consensus sequence-based method for correcting the errors and reduce false-positive rates. Methods Here we present a model to correct base position errors in Tapestri single-cell DNA analytical workflow. The error correction method involves 2 steps. First, we train the sample set to generate base transition probabilities. Then we use the base transition probability to error correct reads per amplicon. We show that base substitution rates could have a significant coefficient of variation across different Tapestri runs/ sequencing runs and could have standard deviations as high as 73% of the mean error rate. Hence the first step of training on the fly per run becomes important. To train the transition probabilities we parse the reads from cell bam files having cell barcode as the read group identifier. A pileup is generated per amplicon and we calculate the probability of each base occurring along the amplicon. We also see that training across all the reads tends to flatten out and hence to improve compute efficiency we stop training after a few million bases. After we generate the multinomial transition probabilities we apply error correction on each read per cell. We calculate the frequency at which each non-reference base is observed per loci. We use the transition probability matrix and correct to reference based on a threshold of significance. Also, to ensure that we filter out the noisy reads before passing the data to our variant caller we suppress the quality scores of reads having very low coverage. Since error correcting on deeper sequencing runs is very computationally intensive, we used several methods to parallelize this process to make use of modern compute-architectures. Results To validate this method, we used two different targeted panels on a Latin square model system with known truth mutations. We performed titrations experiments of 4 cell line mixtures with 98.4%, 1%, 0.5% and 0.1% dilutions. We processed the samples through Tapestri platform and sequenced over multiple Illumina sequencers (Hiseq 2500, Miseq). We ran the Tapestri analytical workflow with and without error correction. With the error correction pipeline, we were able to reduce our false positive rates by ~60-80% while maintaining our sensitivity.
Recent advancements in single cell analysis technologies are now able to provide insights into genomic DNA content, RNA expression and protein surface markers. In bulk assays, the effect of genetic variation on gene expression would be masked by the heterogeneity inherent in tumor cells. Additionally, for cancer immunotherapy studies that rely on gene editing, single cell resolution is necessary to minimize possible off target effects. However, the ability to simultaneously interrogate multiple intracellular analytes, such as genomic DNA and RNA, have proved difficult to implement in a high throughput single cell workflow. We report the development of a complete solution that enables both targeted genomic DNA and RNA sequencing from individual cells. This workflow relies on the Tapestri microfluidic droplet platform, where up to 20,000 cells can be sequenced in each run. Leveraging proprietary cell barcoding, novel primer design strategies and enzymatic manipulation of cellular contents, DNA and RNA multiplex targeted sequencing panels provide for independent barcoded sequence information from both overlapping mRNA and corresponding genomic DNA regions. In addition, amplification primers can be designed to target separate gDNA and non-overlapping RNA transcript. Sequencing is followed by an integrated analysis solution that assigns the reads from both the gDNA and RNA to each cell. Feasibility of the targeted nucleic acid workflow has been shown with inputs from mixed cancer cell lines. A targeted sequencing panel with overlapping mRNA and gDNA regions was designed covering oncogenes, tumor suppressor genes, and known fusions. Expected SNVs and indels were detected and gene expression measured for thousands of cells per run with high cell recovery. This complete solution for single cell multiomics on the Tapestri platform has the power to quantitatively and unambiguously link genotypic and phenotypic data, giving insight into cancer progression. Citation Format: Dalia Dhingra, Kaustubh Gokhale, Nianzhen Li, Pedro Mendez, Shu Wang, Manimozhi Manivannan, Adam Sciambi, Keith Jones, Charlie Silver, Dennis Eastburn, David Ruff. A complete solution for high throughput single cell targeted multiomic DNA and RNA sequencing for cancer research [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 3540.
Introduction: The challenge in precision medicine has been improving the understanding of cancer heterogeneity and clonal evolution, which has major implications in targeted therapy selection and disease monitoring. However, current bulk sequencing methods are unable to unambiguously identify rare pathogenic or drug-resistant cell populations and determine whether mutations co-occur within the same cell. Single-cell sequencing has the potential to provide unique insights on the cellular and genetic composition, drivers, and signatures of cancer at unparalleled sensitivity. Methods: Previously we have developed a high-throughput single-cell DNA analysis platform (TapestriTM) that leverages droplet microfluidics and a multiplex-PCR based targeted DNA sequencing approach, and demonstrated the generation of high-resolution maps of clonal architecture from acute myeloid leukemia (AML) tumors. Here we present an update to the Tapestri Platform which employs new biochemistry and features improved firmware, software, workflow, and data analysis solution resulting in higher throughput, better sensitivity, specificity and unprecedented flexibility. Results: From cell prep to sequencing-ready libraries, the workflow can be completed within 2 days, and new modifications have doubled the throughput to up to 20,000 genotyped cells per run from 10,000 shown previously. We have validated the performance of an AML (19 genes, 50 amplicons) and a CLL (chronic lymphocytic leukemia) (34 genes, 286 amplicons) panel. We also developed a robust web-based design portal for custom targets. The updated biochemistry enables easy addition of new gene and loci targets into existing panels for improved coverage and updated studies. Using longitudinal AML and CLL samples, we were able to detect rare subclones of <0.1% prevalence, identify mutation co-occurrence, and characterize clonal evolution due to disease progression and drug treatment. Conclusion: We demonstrate that single-cell DNA sequencing can reveal the heterogeneity of blood cancers and map the clonal architecture and clonal evolution with higher sensitivity than bulk NGS methods. This is critical in patient stratification and drug selection over the entire course of treatment. Besides the catalog AML and CLL panels, the flexibility of system allows for analyzing SNV and indel mutations of any custom cancer DNA targets. Additionally, the system provides capabilities for quality control of gene edited cells, further advancing research into cancer therapies. Citation Format: Nianzhen Li, Daniel Mendoza, Adam Sciambi, Mani Manivannan, Jacob Ho, Kaustubh Gokhale, Jacqueline Marin, Kathryn Thompson, Jamie Yates, Vasu Sharma, Steven Chow, Sombeet Sahu, Shu Wang, Dennis Eastburn, Keith Jones. High-throughput single-cell targeted DNA sequencing using an updated TapestriTM Platform reveals rare clones and clonal evolution for multiple blood cancers [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 4696.