Advances in high-throughput microscopy have enabled the rapid acquisition of large numbers of high-content microscopy images. Whether by deep learning or classical algorithms, image analysis pipelines then produce single-cell features. To process these single-cells for downstream applications, we present Pycytominer, a user-friendly, open-source python package that implements the bioinformatics steps, known as image-based profiling. We demonstrate Pycytominers usefulness in a machine learning project to predict nuisance compounds that cause undesirable cell injuries.
Cell Painting images offer valuable insights into a cell's state and enable many biological applications, but publicly available arrayed datasets only include hundreds of genes perturbed. The JUMP Cell Painting Consortium perturbed roughly 75% of the protein-coding genome in human U-2 OS cells, generating a rich resource of single-cell images and extracted features. These profiles capture the phenotypic impacts of perturbing 15,243 human genes, including overexpressing 12,609 genes (using open reading frames) and knocking out 7,975 genes (using CRISPR-Cas9). Here we mitigated technical artifacts by rigorously evaluating data processing options and validated the dataset's robustness and biological relevance. Analysis of phenotypic profiles revealed previously undiscovered gene clusters and functional relationships, including those associated with mitochondrial function, cancer and neural processes. The JUMP Cell Painting genetic dataset is a valuable resource for exploring gene relationships and uncovering previously unknown functions.
We present the nELISA, a high-throughput, high-fidelity, and high-plex protein profiling platform. DNA oligonucleotides are used to pre-assemble antibody pairs on spectrally encoded microparticles and perform displacement-mediated detection. Spatial separation between non-cognate antibodies prevents the rise of reagent-driven cross-reactivity, while read-out is performed cost-efficiently and at high-throughput using flow cytometry. We assembled an inflammatory panel of 191 targets that were multiplexed without cross-reactivity or impact on performance vs 1-plex signals, with sensitivities as low as 0.1pg/mL and measurements spanning 7 orders of magnitude. We then performed a large-scale secretome perturbation screen of peripheral blood mononuclear cells (PBMCs), with cytokines as both perturbagens and read-outs, measuring 7,392 samples and generating ∼1.5M protein datapoints in under a week, a significant advance in throughput compared to other highly multiplexed immunoassays. We uncovered 447 significant cytokine responses, including multiple putatively novel ones, that were conserved across donors and stimulation conditions. We also validated the nELISA’s use in phenotypic screening, and propose its application to drug discovery.
Identifying how a given chemical of interest exerts its impact on biological systems is a critical step in developing new medicines and chemical products. The mechanism of a query compound of interest can sometimes be identified when its image-based morphological profile matches a compound in a library of well-annotated compound profiles. In this study, we demonstrate a significant improvement in classification performance by incorporating side information: gene representations. We generate these representations using the morphological profiles of cells where the level of a single gene’s expression has been artificially increased or decreased. The genes are selected as those encoding known protein targets of annotated compounds in the library. A transformer model is trained to classify gene-compound pairs, where each pair represents a potential interaction between a gene and a compound, as true or false. Subsequently, the model generates a ranked list of likely target genes for a previously unseen query compound. Although the strategy exhibits high performance only for compounds that target previously encountered genes – likely due to the limited size of our training dataset – the performance increase demonstrates a notable improvement over simply matching compound profiles directly to compound profiles or to gene profiles. Larger datasets may improve the prediction capabilities of this approach, enabling the prediction of gene targets for novel compounds, which can then be experimentally validated. ### Competing Interest Statement The Authors declare the following competing interests: S.S. and A.E.C. serve as scientific advisors for companies that use image-based profiling and Cell Painting (A.E.C: Recursion, SyzOnc, Quiver Bioscience, S.S.: Waypoint Bio, Dewpoint Therapeutics, Deepcell) and receive honoraria for occasional talks at pharmaceutical and biotechnology companies. All other authors declare no competing interests.
The identification of genetic and chemical perturbations with similar impacts on cell morphology can elucidate compounds' mechanisms of action or novel regulators of genetic pathways. Research on methods for identifying such similarities has lagged due to a lack of carefully designed and well-annotated image sets of cells treated with chemical and genetic perturbations. Here we create such a Resource dataset, CPJUMP1, in which each perturbed gene's product is a known target of at least two chemical compounds in the dataset. We systematically explore the directionality of correlations among perturbations that target the same protein encoded by a given gene, and we find that identifying matches between chemical and genetic perturbations is a challenging task. Our dataset and baseline analyses provide a benchmark for evaluating methods that measure perturbation similarities and impact, and more generally, learn effective representations of cellular state from microscopy images. Such advancements would accelerate the applications of image-based profiling of cellular states, such as uncovering drug mode of action or probing functional genomics. The CPJUMP1 Resource comprises Cell Painting images and profiles of 75 million cells treated with hundreds of chemical and genetic perturbations. The dataset enables exploration of their relationships and lays the foundation for the development of advanced methods to match perturbations.
Mechanistic studies of Geobacillus stearothermophilus tryptophanyl-tRNA synthetase (TrpRS) afford an unusually detailed description-the escapement mechanism-for the distinct steps coupling catalysis to domain motion, efficiently converting the free energy of ATP hydrolysis into biologically useful alternative forms of information and work. Further elucidation of the escapement mechanism requires understanding thermodynamic linkages between domain configuration and conformational stability. To that end, we compare experimental thermal melting of fully liganded and apo TrpRS with a computational simulation of the melting of its fully liganded form. The simulation also provides important structural cameos at successively higher temperatures, enabling more confident interpretation. Experimental and simulated melting both proceed through a succession of three transitions at successively higher temperature. The low-temperature transition occurs at approximately the growth temperature of the organism and so may be functionally relevant but remains too subtle to characterize structurally. Structural metrics from the simulation imply that the two higher-temperature transitions entail forming a molten globular state followed by unfolding of secondary structures. Ligands that stabilize the enzyme in a pre-transition (PreTS) state compress the temperature range over which these transitions occur and sharpen the transitions to the molten globule and fully denatured states, while broadening the low-temperature transition. The experimental enthalpy changes provide a key parameter necessary to convert changes in melting temperature of combinatorial mutants into mutationally induced conformational free energy changes. The TrpRS urzyme, an excerpted model representing an early ancestral form, containing virtually the entire catalytic apparatus, remains largely intact at the highest simulated temperatures.
In image-based profiling, software extracts thousands of morphological features of cells from multi-channel fluorescence microscopy images, yielding single-cell profiles that can be used for basic research and drug discovery. Powerful applications have been proven, including clustering chemical and genetic perturbations on the basis of their similar morphological impact, identifying disease phenotypes by observing differences in profiles between healthy and diseased cells and predicting assay outcomes by using machine learning, among many others. Here, we provide an updated protocol for the most popular assay for image-based profiling, Cell Painting. Introduced in 2013, it uses six stains imaged in five channels and labels eight diverse components of the cell: DNA, cytoplasmic RNA, nucleoli, actin, Golgi apparatus, plasma membrane, endoplasmic reticulum and mitochondria. The original protocol was updated in 2016 on the basis of several years’ experience running it at two sites, after optimizing it by visual stain quality. Here, we describe the work of the Joint Undertaking for Morphological Profiling Cell Painting Consortium, to improve upon the assay via quantitative optimization by measuring the assay’s ability to detect morphological phenotypes and group similar perturbations together. The assay gives very robust outputs despite various changes to the protocol, and two vendors’ dyes work equivalently well. We present Cell Painting version 3, in which some steps are simplified and several stain concentrations can be reduced, saving costs. Cell culture and image acquisition take 1–2 weeks for typically sized batches of ≤20 plates; feature extraction and data analysis take an additional 1–2 weeks. This protocol is an update to Nat. Protoc. 11, 1757–1774 (2016): https://doi.org/10.1038/nprot.2016.105 We provide an updated protocol for image-based profiling with Cell Painting. A detailed procedure, with standardized conditions for the assay, is presented, along with a comprehensive description of parameters to be considered when optimizing the assay.
Image-based profiling has emerged as a powerful technology for various steps in basic biological and pharmaceutical discovery, but the community has lacked a large, public reference set of data from chemical and genetic perturbations. Here we present data generated by the Joint Undertaking for Morphological Profiling (JUMP)-Cell Painting Consortium, a collaboration between 10 pharmaceutical companies, six supporting technology companies, and two non-profit partners. When completed, the dataset will contain images and profiles from the Cell Painting assay for over 116,750 unique compounds, over-expression of 12,602 genes, and knockout of 7,975 genes using CRISPR-Cas9, all in human osteosarcoma cells (U2OS). The dataset is estimated to be 115 TB in size and capturing 1.6 billion cells and their single-cell profiles. File quality control and upload is underway and will be completed over the coming months at the Cell Painting Gallery: https://registry.opendata.aws/cellpainting-gallery . A portal to visualize a subset of the data is available at https://phenaid.ardigen.com/jumpcpexplorer/ .
Morphological and gene expression profiling can cost-effectively capture thousands of features in thousands of samples across perturbations by disease, mutation, or drug treatments, but it is unclear to what extent the two modalities capture overlapping versus complementary information. Here, using both the L1000 and Cell Painting assays to profile gene expression and cell morphology, respectively, we perturb human A549 lung cancer cells with 1,327 small molecules from the Drug Repurposing Hub across six doses, providing a data resource including dose-response data from both assays. The two assays capture both shared and complementary information for mapping cell state. Cell Painting profiles from compound perturbations are more reproducible and show more diversity but measure fewer distinct groups of features. Applying unsupervised and supervised methods to predict compound mechanisms of action (MOAs) and gene targets, we find that the two assays not only provide a partially shared but also a complementary view of drug mechanisms. Given the numerous applications of profiling in biology, our analyses provide guidance for planning experiments that profile cells for detecting distinct cell types, disease phenotypes, and response to chemical or genetic perturbations.
Image-based profiling is a maturing strategy by which the rich information present in biological images is reduced to a multidimensional profile, a collection of extracted image-based features. These profiles can be mined for relevant patterns, revealing unexpected biological activity that is useful for many steps in the drug discovery process. Such applications include identifying disease-associated screenable phenotypes, understanding disease mechanisms and predicting a drug’s activity, toxicity or mechanism of action. Several of these applications have been recently validated and have moved into production mode within academia and the pharmaceutical industry. Some of these have yielded disappointing results in practice but are now of renewed interest due to improved machine-learning strategies that better leverage image-based information. Although challenges remain, novel computational technologies such as deep learning and single-cell methods that better capture the biological information in images hold promise for accelerating drug discovery.
The D1 switch is a packing motif, broadly distributed in the proteome, that couples tryptophanyl-tRNA synthetase (TrpRS) domain movement to catalysis and specificity, thereby creating an escapement mechanism essential to free-energy transduction. The escapement mechanism arose from analysis of an extensive set of combinatorial mutations to this motif, which allowed us to relate mutant-induced changes quantitatively to both kinetic and computational parameters during catalysis. To further characterize the origins of this escapement mechanism in differential TrpRS conformational stabilities, we use high-throughput Thermofluor measurements for the 16 variants to extend analysis of the mutated residues to their impact on unliganded TrpRS stability. Aggregation of denatured proteins complicates thermodynamic interpretations of denaturation experiments. The free energy landscape of a liganded TrpRS complex, carried out for different purposes, closely matches the volume, helix content, and transition temperatures of Thermoflour and CD melting profiles. Regression analysis using the combinatorial design matrix accounts for >90% of the variance in Tms of both Thermofluor and CD melting profiles. We argue that the agreement of experimental melting temperatures with both computational free energy landscape and with Regression modeling means that experimental melting profiles can be used to analyze the thermodynamic impact of combinatorial mutations. Tertiary packing and aromatic stacking of Phenylalanine 37 exerts a dominant stabilizing effect on both native and molten globular states. The TrpRS Urzyme structure remains essentially intact at the highest temperatures explored by the simulations.
PATH algorithms for identifying conformational transition states provide computational parameters—time to the transition state, conformational free energy differences, and transition state activation energies—for comparison to experimental data and can be carried out sufficiently rapidly to use in the “high throughput” mode. These advantages are especially useful for interpreting results from combinatorial mutagenesis experiments. This report updates the previously published algorithm with enhancements that improve correlations between PATH convergence parameters derived from virtual variant structures generated by RosettaBackrub and previously published kinetic data for a complete, four-way combinatorial mutagenesis of a conformational switch in Tryptophanyl-tRNA synthetase.
We measured and cross-validated the energetics of networks in Bacillus stearothermophilus Tryptophanyl-tRNA synthetase (TrpRS) using both multi-mutant and modular thermodynamic cycles. Multi-dimensional combinatorial mutagenesis showed that four side chains from this “molecular switch” move coordinately with the active-site Mg2+ ion as the active site preorganizes to stabilize the transition state for amino acid activation. A modular thermodynamic cycle consisting of full-length TrpRS, its Urzyme, and the Urzyme plus each of the two domains deleted in the Urzyme gives similar energetics. These dynamic linkages, although unlikely to stabilize the transition-state directly, consign the active-site preorganization to domain motion, assuring coupled vectorial behavior.
PATH rapidly computes a path and a transition state between crystal structures by minimizing the Onsager-Machlup action. It requires input parameters whose range of values can generate different transition-state structures that cannot be uniquely compared with those generated by other methods. We outline modifications to estimate these input parameters to circumvent these difficulties and validate the PATH transition states by showing consistency between transition-states derived by different algorithms for unrelated protein systems. Although functional protein conformational change trajectories are to a degree stochastic, they nonetheless pass through a well-defined transition state whose detailed structural properties can rapidly be identified using PATH.
Aminoacyl-tRNA synthetases (aaRS) catalyze both chemical steps that translate the universal genetic code. Rodin and Ohno offered an explanation for the existence of two aaRS classes, observing that codons for the most highly conserved Class I active-site residues are anticodons for corresponding Class II active-site residues. They proposed that the two classes arose simultaneously, by translation of opposite strands from the same gene. We have characterized wild-type 46-residue peptides containing ATP-binding sites of Class I and II synthetases and those coded by a gene designed by Rosetta to encode the corresponding peptides on opposite strands. Catalysis by WT and designed peptides is saturable, and the designed peptides are sensitive to active-site residue mutation. All have comparable apparent second-order rate constants 2.9-7.0E-3 M(-1) s(-1) or ∼750,000-1,300,000 times the uncatalyzed rate. The activities of the two complementary peptides demonstrate that the unique information in a gene can have two functional interpretations, one from each complementary strand. The peptides contain phylogenetic signatures of longer, more sophisticated catalysts we call Urzymes and are short enough to bridge the gap between them and simpler uncoded peptides. Thus, they directly substantiate the sense/antisense coding ancestry of Class I and II aaRS. Furthermore, designed 46-mers achieve similar catalytic proficiency to wild-type 46-mers by significant increases in both kcat and Km values, supporting suggestions that the earliest peptide catalysts activated ATP for biosynthetic purposes.
Urzymology is the enzymological characterization of invariant cores, containing ~100 ± 40 amino acids, identified in enzyme superfamilies (1,2). Urzymes prepared from Class I and Class II aminoacyl‐tRNA synthetases (aaRS) accelerate cognate amino acid activation and tRNA acylation more than 105 times faster than necessary to support ribosome‐independent assembly of polypeptides and are therefore themselves highly evolved (3). Their high catalytic activities afford an experimental basis for combinatorial analysis of the effects of modular enhancements thought to have been involved in generating modern enzymes (4).Among the central questions now accessible via the study of Urzymes are the likely origin of Class I and II aaRS on opposite strands of the same ancestral gene (5); the origin of intramolecular communication (6,7); the modular construction of the Urzymes themselves, and hence their descent from even more primitive catalysts; and the suggestion that an operational RNA code preceded development of the canonical genetic code.Grant Funding Source: NIGMS 78227 and 40906
We previously showed (Li, L., and Carter, C. W., Jr. (2013) J. Biol. Chem. 288, 34736-34745) that increased specificity for tryptophan versus tyrosine by contemporary Bacillus stearothermophilus tryptophanyl-tRNA synthetase (TrpRS) over that of TrpRS Urzyme results entirely from coupling between the anticodon-binding domain and an insertion into the Rossmann-fold known as Connecting Peptide 1. We show that this effect is closely related to a long range catalytic effect, in which side chain repacking in a region called the D1 Switch, accounts fully for the entire catalytic contribution of the catalytic Mg(2+) ion. We report intrinsic and higher order interaction effects on the specificity ratio, (kcat/Km)Trp/(kcat/Km)Tyr, of 15 combinatorial mutants from a previous study (Weinreb, V., Li, L., and Carter, C. W., Jr. (2012) Structure 20, 128-138) of the catalytic role of the D1 Switch. Unexpectedly, the same four-way interaction both activates catalytic assist by Mg(2+) ion and contributes -4.4 kcal/mol to the free energy of the specificity ratio. A minimum action path computed for the induced-fit and catalytic conformation changes shows that repacking of the four residues precedes a decrease in the volume of the tryptophan-binding pocket. We suggest that previous efforts to alter amino acid specificities of TrpRS and glutaminyl-tRNA synthetase (GlnRS) by mutagenesis without extensive, modular substitution failed because mutations were incompatible with interdomain motions required for catalysis.
Design of a regulatable multistate protein is a challenge for protein engineering. Here we design a protein with a unique topology, called uniRapR, whose conformation is controlled by the binding of a small molecule. We confirm switching and control ability of uniRapR in silico , in vitro, and in vivo. As a proof of concept, uniRapR is used as an artificial regulatory domain to control activity of kinases. By activating Src kinase using uniRapR in single cells and whole organism, we observe two unique phenotypes consistent with its role in metastasis. Activation of Src kinase leads to rapid induction of protrusion with polarized spreading in HeLa cells, and morphological changes with loss of cell–cell contacts in the epidermal tissue of zebrafish. The rational creation of uniRapR exemplifies the strength of computational protein design, and offers a powerful means for targeted activation of many pathways to study signaling in living organisms.
A widespread consensus holds that protein synthesis according to a genetic code was launched entirely by sophisticated RNA molecules that played both coding and functional roles. This belief persists, unsupported by phylogenetic evidence for ancestral ribozymes that catalyzed either amino acid activation or tRNA aminoacylation. By contrast, we have adduced strong experimental evidence that the most highly conserved portions of contemporary aminoacyl-tRNA synthetases (aaRS) accelerate both reactions well in excess of rates achieved by RNA aptomers derived from combinatorial libraries and of rates required for primordial protein synthesis. Such ancestral enzymes, or “Urzymes”, characterized for Class I (TrpRS (Pham et al., 2010, 2007) and LeuRS (Collier et al., 2013); 130 residues) and Class II (HisRS; 120–140 residues; (Li et al., 2011)) synthetases generally have promiscuous amino acid specificities, whereas ATP and cognate tRNA affinities are within an order of magnitude of those for contemporary enzymes. These characteristics match or exceed expectations for the primordial catalysts necessary to launch protein synthesis. Structural hierarchies in Class I and II aaRS also exhibit plateaus of increasing enzymatic activity, suggesting that catalysis by peptides similar to the Aleph motif identified by Trifonov (Sobolevsky et al.) may have been both necessary and sufficient to launch protein synthesis. Sense/antisense alignments of TrpRS and HisRS Urzyme coding sequences reveal unexpectedly high middle-base complementarity that increases in reconstructed ancestral nodes (Chandrasekaran et al.), consistent with the proposal of Rodin and Ohno (Rodin & Ohno, 1995). Thus, these ancestors were likely coded by opposite strands of the same gene, favoring simultaneous expression of aaRS activating both hydrophobic (core) and hydrophilic (surface) amino acids. Our results support the view that aaRS coevolved with cognate tRNAs from a much earlier stage than that envisioned under the RNA World hypothesis, and that their descendants make up appreciable portions of the proteome.
We tested the idea that ancestral class I and II aminoacyl-tRNA synthetases arose on opposite strands of the same gene. We assembled excerpted 94-residue Urgenes for class I tryptophanyl-tRNA synthetase (TrpRS) and class II Histidyl-tRNA synthetase (HisRS) from a diverse group of species, by identifying and catenating three blocks coding for secondary structures that position the most highly conserved, active-site residues. The codon middle-base pairing frequency was 0.35 ± 0.0002 in all-by-all sense/antisense alignments for 211 TrpRS and 207 HisRS sequences, compared with frequencies between 0.22 ± 0.0009 and 0.27 ± 0.0005 for eight different representations of the null hypothesis. Clustering algorithms demonstrate further that profiles of middle-base pairing in the synthetase antisense alignments are correlated along the sequences from one species-pair to another, whereas this is not the case for similar operations on sets representing the null hypothesis. Most probable reconstructed sequences for ancestral nodes of maximum likelihood trees show that middle-base pairing frequency increases to approximately 0.42 ± 0.002 as bacterial trees approach their roots; ancestral nodes from trees including archaeal sequences show a less pronounced increase. Thus, contemporary and reconstructed sequences all validate important bioinformatic predictions based on descent from opposite strands of the same ancestral gene. They further provide novel evidence for the hypothesis that bacteria lie closer than archaea to the origin of translation. Moreover, the inverse polarity of genetic coding, together with a priori α-helix propensities suggest that in-frame coding on opposite strands leads to similar secondary structures with opposite polarity, as observed in TrpRS and HisRS crystal structures.