Building predictive models of the cell requires systematically mapping how perturbations reshape each cell's state, function, and behavior. Here, we present Tahoe-100M, a giga-scale single-cell atlas of 100 million transcriptomic profiles measuring how each of 1,100 small-molecule perturbations impact cells across 50 cancer cell lines. Our high-throughput Mosaic platform, composed of a highly diverse and optimally balanced 'cell village', reduces batch effects and enables parallel profiling of thousands of conditions at single-cell resolution at an unprecedented scale. As the largest single-cell dataset to date, Tahoe-100M enables artificial-intelligence (AI)-driven models to learn context-dependent functions, capturing fundamental principles of gene regulation and network dynamics. Although we leverage cancer models and pharmacological compounds to create this resource, Tahoe-100M is fundamentally designed as a broadly applicable perturbation atlas and supports deeper insights into cell biology across multiple tissues and contexts. By publicly releasing this atlas, we aim to accelerate the creation and development of robust AI frameworks for systems biology, ultimately improving our ability to predict and manipulate cellular behaviors across a wide range of applications. ### Competing Interest Statement Jesse Zhang, Airol A Ubas, Richard de Borja, Valentine Svensson, Nicole Thomas, Neha Thakar, Ian Lai, Aidan Winters, Umair Khan, Matthew G. Jones, Daniele Merico, Nima Alidoust, Hani Goodarzi, Johnny Yu, have received compensation and/or equity from Vevo Therapeutics. Vuong Tran, Joseph Pangallo, Efthymia Papalexi, Ajay Sapre, Hoai Nguyen, Oliver Sanderson, Maria Nigos, Olivia Kaplan, Sarah Schroeder, Bryan Hariadi, Simone Marrujo, Crina Curca Alec Salvino, Guillermo Gallareta Olivares, Ryan Koehler, Gary Geiss, Alexander Rosenberg, Charles Roco, have received compensation and/or equity from Parse Biosciences.
Single-cell RNA-sequencing (scRNA-seq) has rapidly spread across multiple research fields, leading to new discoveries. Many applications of scRNA-seq are focused on cell type identification, gene regulatory networks, or biomarker discovery which require the interrogation of specific sets of well-characterized genes. In these cases, sequencing the entire transcriptome may be adding unnecessary project costs. To increase throughput and minimize sequencing costs, the development of a targeted gene enrichment method is required. Here, we extend our whole transcriptome (WT) split-pool combinatorial barcoding technology to enrich a subset of genes in 16 libraries representing hundreds of thousands of human bone marrow mononuclear cells (BMMCs) from three acute myeloid leukemia (AML) and one acute lymphocytic leukemia (ALL) donors. We used our 1,000 immune gene panel to enrich genes representing canonical immune cell markers and pathways. Our method increased the percent of reads on target from as low as 7% in the whole transcriptome libraries to 75% in the targeted libraries. Furthermore, despite a nearly ten-fold reduction in sequencing reads between unenriched and enriched libraries, the resulting clustering yielded high concordance of cell type identities and preserved leukemia-specific signatures such as FLT3, MKI67, and CD19. Overall, we demonstrate our modular enrichment strategy preserves biological structure and allows for deep characterization of gene signatures in health and disease. We envision our approach will enable researchers to simultaneously reduce sequencing costs while drastically scaling up the number of cells and samples across experiments.
Single cell RNA sequencing (scRNA-seq) has become a core tool for researchers to understand biology. As scRNA-seq has become more ubiquitous, many applications demand higher scalability and sensitivity. Split-pool combinatorial barcoding makes it possible to scale projects to hundreds of samples and millions of cells, overcoming limitations of previous droplet based technologies. However, there is still a need for increased sensitivity for both droplet and combinatorial barcoding based scRNA-seq technologies. To meet this need, here we introduce an updated combinatorial barcoding method for scRNA-seq with dramatically improved sensitivity. To assess performance, we profile a variety of sample types, including cell lines, human peripheral blood mononuclear cells (PBMCs), mouse brain nuclei, and mouse liver nuclei. When compared to the previously best performing approach, we find up to a 2.6-fold increase in unique transcripts detected per cell and up to a 1.8-fold increase in genes detected per cell. These improvements to transcript and gene detection increase the resolution of the resulting data, making it easier to distinguish cell types and states in heterogeneous samples. Split-pool combinatorial barcoding already enables scaling to millions of cells, the ability to perform scRNA-seq on previously fixed and frozen samples, and access to scRNA-seq without the need to purchase specialized lab equipment. Our hope is that by combining these previous advantages with the dramatic improvements to sensitivity presented here, we will elevate the standards and capabilities of scRNA-seq for the broader community.
Next-generation deep sequencing of gene panels is being adopted as a diagnostic test to identify actionable mutations in cancer patient samples. However, clinical samples, such as formalin-fixed, paraffin-embedded specimens, frequently provide low quantities of degraded, poor quality DNA. To overcome these issues, many sequencing assays rely on extensive PCR amplification leading to an accumulation of bias and artifacts. Thus, there is a need for a targeted sequencing assay that performs well with DNA of low quality and quantity without relying on extensive PCR amplification. We evaluate the performance of a targeted sequencing assay based on Oligonucleotide Selective Sequencing, which permits the enrichment of genes and regions of interest and the identification of sequence variants from low amounts of damaged DNA. This assay utilizes a repair process adapted to clinical FFPE samples, followed by adaptor ligation to single stranded DNA and a primer-based capture technique. Our approach generates sequence libraries of high fidelity with reduced reliance on extensive PCR amplification—this facilitates the accurate assessment of copy number alterations in addition to delivering accurate single nucleotide variant and insertion/deletion detection. We apply this method to capture and sequence the exons of a panel of 130 cancer-related genes, from which we obtain high read coverage uniformity across the targeted regions at starting input DNA amounts as low as 10 ng per sample. We demonstrate the performance using a series of reference DNA samples, and by identifying sequence variants in DNA from matched clinical samples originating from different tissue types.
Abstract Cancer tumor profiling by targeted resequencing of actionable cancer genes is rapidly becoming the standard approach for selecting targeted therapies and clinical trials in refractory cancer patients. In this clinical scenario, a tumor sample is obtained from an FFPE block and sequenced by targeted next-generation sequencing (NGS) to uncover actionable somatic mutations in relevant cancer genes. Some of the challenges that arise in analyzing tumor-derived NGS data include distinguishing between somatic and germline variants in the absence of normal tissue data, recognizing pathogenic germline variants, and identifying sequencing errors (which occur at about 0.5% rate). Additional challenges arise when considering other clinical applications of NGS such as sequencing cell-free tumor DNA (cf-DNA) from plasma samples to monitor disease response or disease recurrence. Here we present a principled approach to identify both single-nucleotide and small insertion/deletion somatic mutations and germline variants from NGS data of tumor tissue that leverages the allelic fraction patterns in tumors and prior information from external databases through the use of a Bayesian Network algorithm. Our approach allows us to score each putative mutation or variant with respect to its probability of belonging to each variant class, versus classification as a sequencing error. The method enables the joint calling of related samples form the same patient, such as cases where a cf-DNA sample and primary tumor sample are both profiled improving sensitivity and specificity. We validated our method by analyzing data obtained with the TOMA OS-Seq targeted sequencing RUO assay for 98 cancer genes from a mixture of well-known genomes, patient case triads (where normal, tumor and cf-DNA are available), and a retrospective analysis of tumor patient data that underwent clinical tumor profiling for therapy selection. Citation Format: Francisco M. De La Vega, Ryan T. Koehler, Yannick Pouliot, Yosr Bouhlal, Austin So, Federico Goodsaid, Sean Irvine, Len Trigg, Lincoln Nadauld. Joint somatic mutation and germline variant identification and scoring from tumor molecular profiling and ct-DNA monitoring of cancer patients by high-throughput sequencing. [abstract]. In: Proceedings of the 107th Annual Meeting of the American Association for Cancer Research; 2016 Apr 16-20; New Orleans, LA. Philadelphia (PA): AACR; Cancer Res 2016;76(14 Suppl):Abstract nr 2712.
Determining the chromosomal phase of pairs of sequence variants - the arrangement of specific alleles as haplotypes - is a routine challenge in molecular genetics. Here we describe Drop-Phase, a molecular method for quickly ascertaining the phase of pairs of DNA sequence variants (separated by 1-200 kb) without cloning or manual single-molecule dilution. In each Drop-Phase reaction, genomic DNA segments are isolated in tens of thousands of nanoliter-sized droplets together with allele-specific fluorescence probes, in a single reaction well. Physically linked alleles partition into the same droplets, revealing their chromosomal phase in the co-distribution of fluorophores across droplets. We demonstrated the accuracy of this method by phasing members of trios (revealing 100% concordance with inheritance information), and demonstrate a common clinical application by phasing CFTR alleles at genomic distances of 11-116 kb in the genomes of cystic fibrosis patients. Drop-Phase is rapid (requiring less than 4 hours), scalable (to hundreds of samples), and effective at long genomic distances (200 kb).
The human epidermal growth factor receptor 2 (HER2, also known as erbB2) gene is involved in signal transduction for cell growth and differentiation. It is a cell surface receptor tyrosine kinase and a proto-oncogene. Overexpression of HER2 is of clinical relevance in breast cancer due to its prognostic value correlating elevated expression with worsening clinical outcome. At the same time, HER2 assessment is also of importance because successful anti-tumor treatment with Herceptin® is strongly correlated with HER2 overexpression in the tumor (approximately 30% of all breast tumors overexpress HER2). In a comprehensive national study, Wolff et al. [1] state that “Approximately 20% of current HER2 testing may be inaccurate” which underscores the importance of developing more accurate methods to determine HER2 status. Droplet Digital™ PCR (ddPCR™) has the potential to improve upon HER2 measurements due to its ability to quantitate DNA and RNA targets with high precision and accuracy. Here we present a study which investigates whether ddPCR can be used to assess HER2 transcript levels in formalin-fixed paraffin embedded (FFPE) human breast tumors and whether these ddPCR measurements agree with prior assessments of these same samples by pathologists using immunohistochemistry (IHC) and in some cases fluorescence in situ hybridization (FISH). We also determined the copy number of HER2 in these samples as compared to the CEP17 reference gene. Results: Clinical FFPE samples were successfully studied using ddPCR and compared to results from standard FISH and IHC methodology. The results demonstrate that ddPCR can rank order the samples in complete agreement with the current standard methods and that ddPCR has the added benefit of providing quantitative results, rather than relying on the expert skill of a seasoned pathologist for determination.
BACKGROUND:Human epidermal growth factor receptor 2 (HER2) testing is routinely performed by immunohistochemistry (IHC) and/or fluorescence in situ hybridization (FISH) analyses for all new cases of invasive breast carcinoma. IHC is easier to perform, but analysis can be subjective and variable. FISH offers better diagnostic accuracy and added confidence, particularly when it is used to supplement weak IHC signals, but it is more labor intensive and costly than IHC. We examined the performance of droplet digital PCR (ddPCR) as a more precise and less subjective alternative for quantifying HER2 DNA amplification. METHODS:Thirty-nine cases of invasive breast carcinoma containing ≥30% tumor were classified as positive or negative for HER2 by IHC, FISH, or both. DNA templates for these cases were prepared from formalin-fixed paraffin-embedded (FFPE) tissues to determine the HER2 copy number by ddPCR. ddPCR involved emulsifying hydrolysis probe-based PCR reaction mixtures containing the ERBB2 [v-erb-b2 erythroblastic leukemia viral oncogene homolog 2, neuro/glioblastoma derived oncogene homolog (avian); also known as HER2] gene and chromosome 17 centromere assays into nanoliter-sized droplets for thermal cycling and analysis. RESULTS:ddPCR distinguished, through differences in the level of HER2 amplification, the 10 HER2-positive samples from the 29 HER2-negative samples with 100% concordance to HER2 status obtained by FISH and IHC analysis. ddPCR results agreed with the FISH results for the 6 cases that were equivocal by IHC analyses, confirming 2 of these samples as positive for HER2 and the other 4 as negative. CONCLUSIONS:ddPCR can be used as a molecular-analysis tool to precisely measure copy number alterations in FFPE samples of heterogeneous breast tumor tissue.
Two years ago, we described the first droplet digital PCR (ddPCR) system aimed at empowering all researchers with a tool that removes the substantial uncertainties associated with using the analogue standard, quantitative real-time PCR (qPCR). This system enabled TaqMan hydrolysis probe-based assays for the absolute quantification of nucleic acids. Due to significant advancements in droplet chemistry and buoyed by the multiple benefits associated with dye-based target detection, we have created a "second generation" ddPCR system compatible with both TaqMan-probe and DNA-binding dye detection chemistries. Herein, we describe the operating characteristics of DNA-binding dye based ddPCR and offer a side-by-side comparison to TaqMan probe detection. By partitioning each sample prior to thermal cycling, we demonstrate that it is now possible to use a DNA-binding dye for the quantification of multiple target species from a single reaction. The increased resolution associated with partitioning also made it possible to visualize and account for signals arising from nonspecific amplification products. We expect that the ability to combine the precision of ddPCR with both DNA-binding dye and TaqMan probe detection chemistries will further enable the research community to answer complex and diverse genetic questions.
Abstract Molecular tests for genetic mutations play an important role in the diagnosis of cancer. Somatic mutations that drive the pathological features of most tumors have increasing promise as biomarkers for cancer prognosis and therapeutic efficacy. The detection of somatic mutations poses an analytical challenge due to the heterogeneous nature of most samples, where a gene carrying a mutation may differ from the highly abundant wild type sequence by only a single nucleotide. Although a variety of methods exist for mutation analysis, many have poor selectivity and fail to detect mutant sequence below 1 in 100 wildtype sequences. Methods that provide better discrimination and quantitation of somatic mutations are desirable. Here we present a simple strategy using droplet digital™ PCR (ddPCR™) for the detection of somatic mutations with high selectivity and sensitivity. Based on the simple principle of sample partitioning into water-in-oil microdroplets, this ddPCR method increases the abundance of a mutant DNA sequence up to 20,000 times compared to an equivalent bulk PCR reaction. Using conventional TaqMan chemistries and workflow, selectivities of up to 1/100,000 can readily be achieved in any laboratory. Here we present results on the use of ddPCR for the detection and quantitation of several clinically important mutations, including KRAS, c-KIT D816V and JAK2 from clinical samples such as bone marrow aspirates and FFPE. Results from ddPCR are compared to those of conventional approaches including allele specific real-time PCR and sequencing. This ddPCR method may play an important role in the earlier detection of cancer, monitoring the progress of disease and response to therapeutics. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 103rd Annual Meeting of the American Association for Cancer Research; 2012 Mar 31-Apr 4; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2012;72(8 Suppl):Abstract nr 4859. doi:1538-7445.AM2012-4859