While several tools have been developed to map axes of variation among individual cells, no analogous approaches exist for identifying axes of variation among multicellular biospecimens profiled at single-cell resolution. For this purpose, we developed ‘phenotypic earth mover’s distance’ (PhEMD). PhEMD is a general method for embedding a ‘manifold of manifolds’, in which each datapoint in the higher-level manifold (of biospecimens) represents a collection of points that span a lower-level manifold (of cells). We apply PhEMD to a newly generated drug-screen dataset and demonstrate that PhEMD uncovers axes of cell subpopulational variation among a large set of perturbation conditions. Moreover, we show that PhEMD can be used to infer the phenotypes of biospecimens not directly profiled. Applied to clinical datasets, PhEMD generates a map of the patient-state space that highlights sources of patient-to-patient variation. PhEMD is scalable, compatible with leading batch-effect correction techniques and generalizable to multiple experimental designs. Phenotypic earth mover’s distance (PhEMD) facilitates the comparison of single-cell experimental conditions, each of which is a high-dimensional dataset, and identifies axes of variation among multicellular biospecimens.
The high-dimensional data created by high-throughput technologies require visualization tools that reveal data structure and patterns in an intuitive form. We present PHATE, a visualization method that captures both local and global nonlinear structure using an information-geometric distance between data points. We compare PHATE to other tools on a variety of artificial and biological datasets, and find that it consistently preserves a range of patterns in data, including continual progressions, branches and clusters, better than other tools. We define a manifold preservation metric, which we call denoised embedding manifold preservation (DEMaP), and show that PHATE produces lower-dimensional embeddings that are quantitatively better denoised as compared to existing visualization methods. An analysis of a newly generated single-cell RNA sequencing dataset on human germ-layer differentiation demonstrates how PHATE reveals unique biological insight into the main developmental branches, including identification of three previously undescribed subpopulations. We also show that PHATE is applicable to a wide variety of data types, including mass cytometry, single-cell RNA sequencing, Hi-C and gut microbiome data.
It is currently challenging to analyze single-cell data consisting of many cells and samples, and to address variations arising from batch effects and different sample preparations. For this purpose, we present SAUCIE, a deep neural network that combines parallelization and scalability offered by neural networks, with the deep representation of data that can be learned by them to perform many single-cell data analysis tasks. Our regularizations (penalties) render features learned in hidden layers of the neural network interpretable. On large, multi-patient datasets, SAUCIE’s various hidden layers contain denoised and batch-corrected data, a low-dimensional visualization and unsupervised clustering, as well as other information that can be used to explore the data. We analyze a 180-sample dataset consisting of 11 million T cells from dengue patients in India, measured with mass cytometry. SAUCIE can batch correct and identify cluster-based signatures of acute dengue infection and create a patient manifold, stratifying immune response to dengue.
Previously, the effect of a drug on a cell population was measured based on simple metrics such as cell viability. However, as single-cell technologies are becoming more advanced, drug screen experiments can now be conducted with more complex readouts such as gene expression profiles of individual cells. The increasing complexity of measurements from these multi-sample experiments calls for more sophisticated analytical approaches than are currently available. We developed a novel method called PhEMD (Phenotypic Earth Mover’s Distance) and show that it can be used to embed the space of drug perturbations on the basis of the drugs’ effects on cell populations. When testing PhEMD on a newly-generated, 300-sample CyTOF kinase inhibition screen experiment, we find that the state space of the perturbation conditions is surprisingly low-dimensional and that the network of drugs demonstrates manifold structure. We show that because of the fairly simple manifold geometry of the 300 samples, we can accurately capture the full range of drug effects using a dictionary of only 30 experimental conditions. We also show that new drugs can be added to our PhEMD embedding using similarities inferred from other characterizations of drugs using a technique called Nystrom extension. Our findings suggest that large-scale drug screens can be conducted by measuring only a small fraction of the drugs using the most expensive high-throughput single-cell technologies—the effects of other drugs may be inferred by mapping and extending the perturbation space. We additionally show that PhEMD can be useful for analyzing other types of single-cell samples, such as patient tumor biopsies, by mapping the patient state space in a similar way as the drug state space. We demonstrate that PhEMD is scalable, compatible with leading batch effect correction techniques, and generalizable to multiple experimental designs. Altogether, our analyses suggest that PhEMD may facilitate drug discovery efforts and help uncover the network geometry of a collection of single-cell samples.
With the advent of high-throughput technologies measuring high-dimensional biological data, there is a pressing need for visualization tools that reveal the structure and emergent patterns of data in an intuitive form. We present PHATE, a visualization method that captures both local and global nonlinear structure in data by an information-geometric distance between datapoints. We perform extensive comparison between PHATE and other tools on a variety of artificial and biological datasets, and find that it consistently preserves a range of patterns in data including continual progressions, branches, and clusters. We define a manifold preservation metric DEMaP to show that PHATE produces quantitatively better denoised embeddings than existing visualization methods. We show that PHATE is able to gain unique insight from a newly generated scRNA-seq dataset of human germ layer differentiation. Here, PHATE reveals a dynamic picture of the main developmental branches in unparalleled detail, including the identification of three novel subpopulations. Finally, we show that PHATE is applicable to a wide variety of datatypes including mass cytometry, single-cell RNA-sequencing, Hi-C, and gut microbiome data, where it can generate interpretable insights into the underlying systems.
Single-cell data are now being collected in large quantities across multiple samples and gene profiling runs. This introduces the need for computational methods that can compare and stratify samples that are represented themselves as complex, high-dimensional objects. We introduce PhEMD as an analytical approach that can be used for this purpose. PhEMD uses Earth Mover’s Distance (EMD), a distance between probability distributions that is sensitive to differences at multiple levels of granularity, in order to compute an accurate measure of dissimilarity between single-cell samples. PhEMD then generates a low-dimensional embedding of the samples based on this dissimilarity. We demonstrate the utility of the PhEMD sample embedding by using it to subtype melanoma and clear-cell renal cell carcinomas based on their immune cell profiles. These analyses reveal sources of inter-sample heterogeneity that have potentially clinically actionable implications, given the recent adoption of immunotherapy as an effective treatment for these cancers. We also apply PhEMD to a newly-generated 300-sample CyTOF drug screen experiment, where the effects of 233 kinase inhibitors are measured at the single-cell resolution in 33 protein dimensions. In doing so, we find that PhEMD reveals novel insights into the effects of small-molecule inhibitors on breast cancer cell subpopulations undergoing epithelial-to-mesenchymal transition. Finally, by leveraging the Nystrom extension method for diffusion maps, we demonstrate that the results of PhEMD can be integrated with other data sources and data types to predict the single-cell phenotypes of samples not directly profiled. Our analyses demonstrate that PhEMD is highly scalable and compatible with leading batch effect correction techniques, allowing for the simultaneous comparison of many single-cell samples.
Abstract Background: A leading model of cancer metastasis is epithelial-to-mesenchymal transition (EMT). We sought to determine whether single-cell inhibition data targeting potential mediators of EMT could uncover mechanistic insights into the EMT process. Methods: EMT was artificially induced on Py2T murine breast cancer cells by TGFb treatment. Additionally, a unique drug inhibitor was added to each well of a multiplexed CyTOF experiment. 37 transcription factors and cell surface markers were measured in each cell to assess epithelial and mesenchymal states, SMAD, AKT, and MAPK signaling activity, cell cycle regulation, and apoptosis pathway activation. The final single-cell dataset consisted of 300 inhibition and control conditions (cell populations), which we aimed to characterize in relation to one another with respect to effect on EMT. Analyzing the similarity between drug inhibitions amounts to a novel type of clustering problem that involves computing the similarity between diverse cell populations generated by each inhibitor. Traditional methods for comparing cell populations are not robust to the intra-population heterogeneity we observed amongst cells undergoing EMT. Thus, we developed Phenotypic Earth Mover’s Distance (PhEMD). This method for comparing cell populations leverages the insight that only a limited number of “cell subtypes” (e.g. mesenchymal, epithelial, transitional) are observed in unperturbed and perturbed EMT. By classifying each cell as one of these distinct subtypes using community-detection based clustering, PhEMD represents an inhibition or control condition as its relative abundance of each cell subtype. It then uses Earth Mover’s Distance (EMD) to compare two relative abundance distributions (i.e. heterogeneous cell populations). PhEMD thus derives a single value representing the dissimilarity between two inhibition conditions. Using PhEMD as measure of dissimilarity between each pair of inhibition and control conditions, we constructed an inhibitor-inhibitor graph and used graph clustering to identify groups of inhibitors that had similar effects to one another. Results: PhEMD analysis revealed that MEK, EGFR and Src inhibitors significantly halted EMT by generating far fewer mesenchymal cells and maintaining a large epithelial cell subpopulation. PhEMD also revealed that different mTOR, PI3K, and Akt inhibitors tended to have similar effects to one another and collectively generated a distinct, transitional cell subpopulation with altered pS6 expression. The pS6 dysregulation may be explained by the fact that ribosomal protein S6 is downstream of PI3K/AKT/mTOR, and the intermediate levels of both E-cadherin and vimentin suggest that this subset of cells may have been halted mid-transition. Finally, several Aurora Kinase and CDK inhibitors resulted in a relatively high percentage of apoptotic cells, suggesting these kinases may be important cell cycle regulators in the context of EMT. Citation Format: William S. Chen, Nevena Zivanovic, Dana Pe'er, Bernd Bodenmiller, Smita Krishnaswamy. Phenotypic analysis of single-cell breast cancer inhibition data reveals insights into EMT [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 977. doi:10.1158/1538-7445.AM2017-977