Cell Painting images offer valuable insights into a cell's state and enable many biological applications, but publicly available arrayed datasets only include hundreds of genes perturbed. The JUMP Cell Painting Consortium perturbed roughly 75% of the protein-coding genome in human U-2 OS cells, generating a rich resource of single-cell images and extracted features. These profiles capture the phenotypic impacts of perturbing 15,243 human genes, including overexpressing 12,609 genes (using open reading frames) and knocking out 7,975 genes (using CRISPR-Cas9). Here we mitigated technical artifacts by rigorously evaluating data processing options and validated the dataset's robustness and biological relevance. Analysis of phenotypic profiles revealed previously undiscovered gene clusters and functional relationships, including those associated with mitochondrial function, cancer and neural processes. The JUMP Cell Painting genetic dataset is a valuable resource for exploring gene relationships and uncovering previously unknown functions.
The identification of genetic and chemical perturbations with similar impacts on cell morphology can elucidate compounds' mechanisms of action or novel regulators of genetic pathways. Research on methods for identifying such similarities has lagged due to a lack of carefully designed and well-annotated image sets of cells treated with chemical and genetic perturbations. Here we create such a Resource dataset, CPJUMP1, in which each perturbed gene's product is a known target of at least two chemical compounds in the dataset. We systematically explore the directionality of correlations among perturbations that target the same protein encoded by a given gene, and we find that identifying matches between chemical and genetic perturbations is a challenging task. Our dataset and baseline analyses provide a benchmark for evaluating methods that measure perturbation similarities and impact, and more generally, learn effective representations of cellular state from microscopy images. Such advancements would accelerate the applications of image-based profiling of cellular states, such as uncovering drug mode of action or probing functional genomics. The CPJUMP1 Resource comprises Cell Painting images and profiles of 75 million cells treated with hundreds of chemical and genetic perturbations. The dataset enables exploration of their relationships and lays the foundation for the development of advanced methods to match perturbations.
In image-based profiling, software extracts thousands of morphological features of cells from multi-channel fluorescence microscopy images, yielding single-cell profiles that can be used for basic research and drug discovery. Powerful applications have been proven, including clustering chemical and genetic perturbations on the basis of their similar morphological impact, identifying disease phenotypes by observing differences in profiles between healthy and diseased cells and predicting assay outcomes by using machine learning, among many others. Here, we provide an updated protocol for the most popular assay for image-based profiling, Cell Painting. Introduced in 2013, it uses six stains imaged in five channels and labels eight diverse components of the cell: DNA, cytoplasmic RNA, nucleoli, actin, Golgi apparatus, plasma membrane, endoplasmic reticulum and mitochondria. The original protocol was updated in 2016 on the basis of several years’ experience running it at two sites, after optimizing it by visual stain quality. Here, we describe the work of the Joint Undertaking for Morphological Profiling Cell Painting Consortium, to improve upon the assay via quantitative optimization by measuring the assay’s ability to detect morphological phenotypes and group similar perturbations together. The assay gives very robust outputs despite various changes to the protocol, and two vendors’ dyes work equivalently well. We present Cell Painting version 3, in which some steps are simplified and several stain concentrations can be reduced, saving costs. Cell culture and image acquisition take 1–2 weeks for typically sized batches of ≤20 plates; feature extraction and data analysis take an additional 1–2 weeks. This protocol is an update to Nat. Protoc. 11, 1757–1774 (2016): https://doi.org/10.1038/nprot.2016.105 We provide an updated protocol for image-based profiling with Cell Painting. A detailed procedure, with standardized conditions for the assay, is presented, along with a comprehensive description of parameters to be considered when optimizing the assay.
Image-based profiling has emerged as a powerful technology for various steps in basic biological and pharmaceutical discovery, but the community has lacked a large, public reference set of data from chemical and genetic perturbations. Here we present data generated by the Joint Undertaking for Morphological Profiling (JUMP)-Cell Painting Consortium, a collaboration between 10 pharmaceutical companies, six supporting technology companies, and two non-profit partners. When completed, the dataset will contain images and profiles from the Cell Painting assay for over 116,750 unique compounds, over-expression of 12,602 genes, and knockout of 7,975 genes using CRISPR-Cas9, all in human osteosarcoma cells (U2OS). The dataset is estimated to be 115 TB in size and capturing 1.6 billion cells and their single-cell profiles. File quality control and upload is underway and will be completed over the coming months at the Cell Painting Gallery: https://registry.opendata.aws/cellpainting-gallery . A portal to visualize a subset of the data is available at https://phenaid.ardigen.com/jumpcpexplorer/ .
Pooled variant expression libraries can test the phenotypes of thousands of variants of a gene in a single multiplexed experiment. In a library encoding all single-amino-acid substitutions of a protein, each variant differs from its reference only at a single codon-position located anywhere along the coding sequence. Consequently, accurately identifying these variants by sequencing is a major technical challenge. A popular but expensive brute-force approach is to divide the pool of variants into multiple smaller sub-libraries that each contains variants of a small region and that must each be constructed and screened individually, but that can then be PCR-amplified and fully sequenced with a single read to allow direct readout of variant abundance. Here we present an approach to screen very large variant libraries with mutations spanning a wide region in a single pool, including library design criteria and mutant-detection algorithms that permit reliable calling and counting of variants from large-scale sequencing data.### Competing Interest StatementW.C.H. is a consultant for ThermoFisher, Solasta, MPM Capital, iTeos, Jubilant Therapeutics, Tyra Therapeutics, RAPPTA Therapeutics, Frontier Medicines, KSQ Therapeutics and Paraxel. A.J.A has consulted for Oncorus, Inc., Arrakis Therapeutics, and Merck & Co., Inc, and has research funding from Mirati Therapeutics, Syros, Deerfield, Inc., and Novo Ventures that is unrelated to this work. A.O.G. is a share and option holder of 10X Genomics. D.E.R. receives research funding from members of the Functional Genomics Consortium (Abbvie, BMS, Jannsen, Merck, Vir), and is a director of Addgene, Inc.
Abstract FANCJ (BRIP1/BACH1) is a hereditary breast and ovarian cancer (HBOC) gene encoding a DNA helicase. Similar to HBOC genes, BRCA1 and BRCA2, FANCJ is critical for processing DNA inter-strand crosslinks (ICL) induced by chemotherapeutics, such as cisplatin. Consequently, cells deficient in FANCJ or its catalytic activity are sensitive to ICL-inducing agents. Unfortunately, the majority of FANCJ clinical mutations remain uncharacterized, limiting therapeutic opportunities to effectively use cisplatin to treat tumors with mutated FANCJ. Here, we sought to perform a comprehensive screen to identify FANCJ loss-of-function (LOF) mutations. We developed a FANCJ lentivirus mutation library representing approximately 450 patient–derived FANCJ nonsense and missense mutations to introduce FANCJ mutants into FANCJ knockout (K/O) HeLa cells. We performed a high-throughput screen to identify FANCJ LOF mutants that, as compared with wild-type FANCJ, fail to robustly restore resistance to ICL-inducing agents, cisplatin or mitomycin C (MMC). On the basis of the failure to confer resistance to either cisplatin or MMC, we identified 26 missense and 25 nonsense LOF mutations. Nonsense mutations elucidated a relationship between location of truncation and ICL sensitivity, as the majority of nonsense mutations before amino acid 860 confer ICL sensitivity. Further validation of a subset of LOF mutations confirmed the ability of the screen to identify FANCJ mutations unable to confer ICL resistance. Finally, mapping the location of LOF mutations to a new homology model provides additional functional information. Implications: We identify 51 FANCJ LOF mutations, providing important classification of FANCJ mutations that will afford additional therapeutic strategies for affected patients.
Abstract The brain is the foremost non-gonadal tissue for expression of non-coding RNAs of unclear function. Yet, whether such transcripts are truly non-coding or rather the source of non-canonical protein translation is unknown. Here, we used functional genomic screens to establish the cellular bioactivity of non-canonical proteins located in putative non-coding RNAs or untranslated regions of protein-coding genes. We experimentally interrogated 553 open reading frames (ORFs) identified by ribosome profiling for three major phenotypes: 257 (46%) demonstrated protein translation when ectopically expressed in HEK293T cells, 401 (73%) induced gene expression changes following ectopic expression across 4 cancer cell types, and 57 (10%) induced a viability defect when the endogenous ORF was knocked out using CRISPR/Cas9 in 8 human cancer cell lines. CRISPR tiling and start codon mutagenesis indicated that the biological impact of these non-canonical ORFs required their translation as opposed to RNA-mediated effects. We functionally characterized one of these ORFs, G029442—renamed GREP1 (Glycine-Rich Extracellular Protein-1)—as a cancer-implicated gene with high expression in multiple cancer types, such as gliomas. GREP1 knockout in >200 cancer cell lines reduced cell viability in multiple cancer types, including glioblastoma, in a cell-autonomous manner and produced cell cycle arrest via single-cell RNA sequencing. Analysis of the secretome of GREP1-expressing cells showed increased abundance of the oncogenic cytokine GDF15, and GDF15 supplementation mitigated the growth inhibitory effect of GREP1 knock-out. Taken together, these experiments suggest that the non-canonical ORFeome is surprisingly rich in biologically active proteins and potential cancer therapeutic targets deserving of further study.
A key question in genome research is whether biologically active proteins are restricted to the ∼20,000 canonical, well-annotated genes, or rather extend to the many non-canonical open reading frames (ORFs) predicted by genomic analyses. To address this, we experimentally interrogated 553 ORFs nominated in ribosome profiling datasets. Of these 553 ORFs, 57 (10%) induced a viability defect when the endogenous ORF was knocked out using CRISPR/Cas9 in 8 human cancer cell lines, 257 (46%) showed evidence of protein translation when ectopically expressed in HEK293T cells, and 401 (73%) induced gene expression changes measured by transcriptional profiling following ectopic expression across 4 cell types. CRISPR tiling and start codon mutagenesis indicated that the biological effects of these non-canonical ORFs required their translation as opposed to RNA-mediated effects. We selected one of these ORFs,G029442--renamedGREP1(Glycine-Rich Extracellular Protein-1)--for further characterization. We found thatGREP1encodes a secreted protein highly expressed in breast cancer, and its knock-out in 263 cancer cell lines showed preferential essentiality in breast cancer derived lines. Analysis of the secretome of GREP1-expressing cells showed increased abundance of the oncogenic cytokine GDF15, and GDF15 supplementation mitigated the growth inhibitory effect ofGREP1knock-out. Taken together, these experiments suggest that the non-canonical ORFeome is surprisingly rich in biologically active proteins and potential cancer therapeutic targets deserving of further study.
Unlike most tumor suppressor genes, the most common genetic alterations in tumor protein p53 (TP53) are missense mutations(1,2) . Mutant p53 protein is often abundantly expressed in cancers and specific allelic variants exhibit dominant-negative or gain-of-function activities in experimental models(3-8). To gain a systematic view of p53 function, we interrogated loss-of-function screens conducted in hundreds of human cancer cell lines and performed TP53 saturation mutagenesis screens in an isogenic pair of TP53 wild-type and null cell lines. We found that loss or dominant-negative inhibition of wild-type p53 function reliably enhanced cellular fitness. By integrating these data with the Catalog of Somatic Mutations in Cancer (COSMIC) mutational signatures database(9,10), we developed a statistical model that describes the TP53 mutational spectrum as a function of the baseline probability of acquiring each mutation and the fitness advantage conferred by attenuation of p53 activity. Collectively, these observations show that widely-acting and tissue-specific mutational processes combine with phenotypic selection to dictate the frequencies of recurrent TP53 mutations.