Performance of Single Cell RNA sequencing (scRNA-seq) experiments depends on multiple factors including the number of cells loaded, cell recovery rates, and sequencing coverage. We optimized such factors in 10x Genomics gene expression profiling and compared targeted panels to whole transcriptome sequencing using Human donor PBMC samples. We expect that our findings will apply to scRNA-seq studies of other sample types as well. Optimizing cell recovery in scRNA-seq experiments is essential in order to ensure that changes in low-abundance cell populations (e.g. Treg) can be accurately quantified. Statistical analysis indicated that at least 100 cells are needed to assess cell-type specific state changes. We expect to recover 40-65% of cells loaded into the 10x chip. We found that loading 20k-30k cells/lane combined with cell hashtags for improved doublet detection improved cell recovery over the recommended loading of 10K-16K cells. We also surveyed internal single cell experiments and observed variable cell recovery across different sample types. The amount of ambient RNA in the library was correlated with lower performance in single cell studies, suggesting that cell death or damage was causing lower recovery. We found that using PBS for cell resuspension (which is less likely to cause cell rupture) performs as well as the nuclease-free water suggested by 10x. We also assessed the impact of sequencing depth and protocol on scRNA-seq data quality. Sequencing depth impacts both experiment cost and data quality due to dropouts. We performed computational simulations of lower depth sequencing by subsampling various number of reads from PBMC experiments to obtain coverages of 10K, 20K, 40K and 80K reads per cell. We observed that 40K reads per cell, yielding about 70% sequencing saturation, provided a good balance between cost and sequencing depth. Finally, we compared dropout rates and cell type annotation for two targeted panels, the Human Gene Signature and Human Immunology panels to whole transcriptome scRNAseq. On average we detected only 200 genes out of over 1000 represented in each panel. Dropout rates for targeted panels were only reduced at low coverage; at read depths higher than 40K reads per cell the whole transcriptome and targeted panels had similar dropout rates. Both methods detected major immune cell types, but targeted sequencing could not accurately identify some subtypes due to low number of detected genes. We conclude that for immune cell profiling whole-transcriptome analysis at coverage of 40K reads per cell or higher with inputs of 20k-30k cells and use of sample hashtag antibodies provides the best balance of experiment cost, cell recovery and transcriptome coverage. Although these guidelines were established for Human PBMCs we expect similar outcomes with other complex cell mixtures such as dissociated tissues. Citation Format: Amir Bayegan, Julien Tessier, Emma Wang, Adalis Maisonet, Shu Yan, Shannon McGrath, Donald G. Jackson, Jack Pollard. Practical guidelines for the design of single cell sequencing studies [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 762.
BackgroundReinvigoration of anti-tumor immunity via immune checkpoint blockade (ICB) has transformed outcomes in a-NSCLC. However, a majority of patients are innately resistant to ICB, and a better understanding of the resistance mechanisms may guide the development of new treatment strategies and therapies for patients.MethodsBiopsies performed immediately before treatment with single agent ICB in patients with a-NSCLC (MATCH-R trial [NCT02517892]) were analyzed. The stromal microenvironment and immune context were characterized via an integrated analysis of whole transcriptome (RNA-seq), whole exome sequencing (WES), and immunohistochemistry (IHC) of CD3, CD8, FOXP3 and PDL1. Specifically, the immune context and the relative abundance of 10 immune and stromal cell types were assessed with integrated IHC and Cell Populations-counter (MCP-counter) [1] analysis of the RNA-seq. Somatic mutations and Tumor Mutation Burden (TMB) were evaluated. The transcriptional state of the tumor and its microenvironment were assessed by GSVA analysis [2] of the MSigDB collection [3]. Patient‘s outcome was associated to molecular data. Primary resistance to ICB was defined as PD (progressive disease) in the first radiological examination, or a median PFS inferior to 3 months.ResultsFifty-two patients with NSCLC were enrolled (43 adeno, 6 squamous, and 3 other carcinoma): Median age was 61 (34–93), 18 were female, 46 were smokers, 22 were responders, and 30 were non-responders. Median tumor cellularity was 60% (30%–90%).Patients may be divided into two groups (HIGH and LOW) at baseline based on their degree of immune infiltration as assessed by RNAseq or IHC. A hallmark of the HIGH infiltration group is an increase in Interferon Gamma (IFN-γ) pathway signature [4]. In contrast, patients in the LOW infiltration group (relative to the HIGH infiltration group) exhibit a decrease in IFN-γ pathway signaling and concomitantly an increase in hypoxia and gluconeogenic pathway signatures. Response rates to ICB were not associated to immune infiltration groups at baseline, but an analysis within each infiltration group revealed that high TMB is only associated to response in the HIGH infiltration group. Furthermore, only in the LOW infiltration group was increased the transforming growth factor (TGF-β) pathway signature associated to ICB response.ConclusionsThis study suggests that the tumor and its microenvironment influence baseline immune infiltration. Tumors with LOW baseline infiltration show altered metabolism such as gluconeogenic activation and hypoxia activation. In contrast, factors such as TMB are not associated with baseline infiltration
Abstract The successful application of Next Generation Sequencing (NGS) to drug discovery requires systems to manage and document each step of the sequencing process from sample receipt through data generation and data processing. We combined BenchlingTM, a solution for tracking NGS lab processes, with FONDA (Framework Of Next generation sequencing Data Analysis) an internally developed data processing platform, to support multiple types of NGS data generation and processing. Benchling combines a digital notebook and a laboratory information management system (LIMS). The system documents and automates steps in the NGS process including: sample registration, nucleic acid extraction, library construction, flow cell construction, sequencer sample sheet generation and BCL2FASTQ conversion. This enables wet lab scientists to easily retrieve an appropriate protocol for each sample and sequencing library type. We connected our sequencers to Benchling in order to monitor each sequencing run and to keep track of the quality of NGS data. In addition, it generates “analysis ready sample sheet” (contains project and study information, location of FASTQ, sample species and library type) and uploads it into designated S3 buckets for data processing. Benchling dashboards provide overviews of NGS sample preparation, data generation and quality control. In summary, Benchling interconnects the original sample, the labels, the barcodes, the cDNA/DNA, the library, and all the QC results. We process NGS data using pipelines implemented in FONDA on a dockerized Amazon Web Services cloud platform. Analyses can be configured automatically from information exported by Benchling or launched manually. After data processing is completed, output files such as gene expression counts or variant calls are deposited into project-specific folders, ready for secondary analysis. In the current FONDA version (as of Nov 2020), we have developed pipelines for single cell multi-omics (CITE-seq and single-cell immune profiling) and bulk RNA-seq. The modular design of FONDA facilitates the development, the updating, and the extension of pipelines to new sequencing technologies. In summary, Benchling and FONDA enable high quality sample and NGS data flows from the lab for target identification, understanding mechanism of action, patient stratification and biomarker discovery. Availability and implementation: FONDA is implemented in Java and released under the Apache License 2.0. FONDA can be downloaded from GitHub at https://github.com/epam/fonda. Citation Format: Chandra Sekhar Pedamallu, Joon Sang Lee, Shu Yan, Adalis Maisonet, Aleksandr Sidoruk, Tengui Chen, Yulia Kamyshova, Mariia Zueva, Mark Magid, Quan Wan, Jeffrey Thompson, Valerie Zebrouck, Immanuel Gadaczek, Mikhail Alperovich, Brian McNatt, Alexei Protopopov, Donald Jackson, Jack Pollard. A comprehensive sample tracking and data processing workflow for next generation sequencing [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 2280.