Advances in high-throughput microscopy have enabled the rapid acquisition of large numbers of high-content microscopy images. Whether by deep learning or classical algorithms, image analysis pipelines then produce single-cell features. To process these single-cells for downstream applications, we present Pycytominer, a user-friendly, open-source python package that implements the bioinformatics steps, known as image-based profiling. We demonstrate Pycytominers usefulness in a machine learning project to predict nuisance compounds that cause undesirable cell injuries.
Large-scale profiling assays capture a cell population's state by measuring thousands of biological properties per cell or sample. However, evaluating profile strength and similarity remains challenging due to the high dimensionality and non-linear, heterogeneous nature of measurements. Here, we develop a statistical framework using mean average precision (mAP) as a single, data-driven metric to address this challenge. We validate the mAP framework against established metrics through simulations and real-world data, revealing its ability to capture subtle and meaningful biological differences in cell state. Specifically, we use mAP to assess a sample's phenotypic activity relative to controls, as well as the phenotypic consistency of groups of perturbations (or samples). We evaluate the framework across diverse datasets and on different profile types (image, protein, mRNA), perturbations (CRISPR, gene overexpression, small molecules), and resolutions (single-cell, bulk). The mAP framework, together with our open-source software package copairs, is useful for evaluating high-dimensional profiling data in biological research and drug discovery.
Cell Painting images offer valuable insights into a cell's state and enable many biological applications, but publicly available arrayed datasets only include hundreds of genes perturbed. The JUMP Cell Painting Consortium perturbed roughly 75% of the protein-coding genome in human U-2 OS cells, generating a rich resource of single-cell images and extracted features. These profiles capture the phenotypic impacts of perturbing 15,243 human genes, including overexpressing 12,609 genes (using open reading frames) and knocking out 7,975 genes (using CRISPR-Cas9). Here we mitigated technical artifacts by rigorously evaluating data processing options and validated the dataset's robustness and biological relevance. Analysis of phenotypic profiles revealed previously undiscovered gene clusters and functional relationships, including those associated with mitochondrial function, cancer and neural processes. The JUMP Cell Painting genetic dataset is a valuable resource for exploring gene relationships and uncovering previously unknown functions.
Introduction: Parkinson’s disease (PD) is a progressive neurological disorder that often leads to impairments in speech (dysarthria) and facial expressiveness (hypomimia). These signs are clinically relevant but often evaluated subjectively. Current methods treat these impairments separately. We propose a multimodal deep learning framework that jointly analyzes audio and video data to detect patterns associated with dysarthria and hypomimia.Method: Facial movement and voice signals were encoded using deep neural representations to identify PD-related patterns. A retrospective study involved 14 native Colombian Spanish speakers: 7 PD patients (mean age 65 ± 4, 4M/3F) and 7 controls (mean age 61 ± 3, 2M/5F). Diagnoses were confirmed by a neurologist and partially labeled via the Hoehn and Yahr scale; 2 patients were off medication and 5 under Levodopa. Data were acquired in controlled conditions using a Nikon D3500 camera (1080p, 60 fps) with monaural audio (48 kHz). Participants performed vowel, phoneme, and word tasks to elicit PD symptoms. This is the first Spanish-speaking synchronized audio-video corpus for PD orofacial analysis. Voice recordings were transformed into Mel spectrograms. Three 2D CNN architectures (ResNet-50, VGG-16, and a custom model) were trained to classify PD vs. controls. Facial videos were processed using a compact 3D-CNN and an inflated 3D (I3D) model to extract spatiotemporal features linked to hypomimia. A late fusion strategy combined audio and video predictions via weighted linear integration. Models were evaluated using leave-one-out cross-validation across pronunciation tasks, reporting precision, recall, F1-score, accuracy, and AUC. All networks were trained with binary cross-entropy, Adam optimizer, early stopping, and a learning rate of 1 \times 10^{-5} over 25 epochs.Results: Audio-based classification using Mel spectrograms showed that VGG-16 outperformed ResNet-50 and a baseline CNN, achieving an average AUC of 69.72%, with peak performance (73.55%) during vowel articulation. For video analysis, a custom 3D-CNN surpassed both a 2D ResNet-50 and an inflated 3D (I3D) model, effectively capturing spatiotemporal facial patterns linked to hypomimia. Multimodal fusion of VGG-16 (audio) and 3D-CNN (video) improved diagnostic accuracy across tasks, with the highest AUC (85.39%) observed for phonemes at fusion weight λ = 0.5. This approach demonstrates the complementary value of combining speech and facial cues.Discussion: The multimodal approach effectively integrates facial and vocal biomarkers. Phoneme-based facial cues via 3D-CNN yielded superior performance (AUC = 82.59%), likely due to expressive dynamics from plosive articulations. Vowel-based audio features processed via VGG-16 achieved an AUC of 73%, benefiting from stable acoustic phases. Fusion enhanced classification across all tasks, supporting the hypothesis that synchronized motor and speech signals offer complementary diagnostic value. Complementary studies have focused on audio-only representations, achieving high accuracy but overlooking facial-vocal interplay. Our synchronized dataset contributes a valuable resource for integrated modeling. Correlations between jaw motion and respiratory control support multimodal analysis. Compared to prior fusion methods, our approach captures richer temporal dynamics and articulatory complexity.Conclusions:The proposed framework improves PD classification by combining deep audio and visual embeddings, with phoneme articulation as the most informative task. Convex fusion outperformed unimodal models, with highest gains for phonemes (AUC = 85.39 at λ = 0.5), reinforcing the clinical relevance of synchronized motor-speech impairments. Future work will explore end-to-end models capable of learning fusion weights and temporal dependencies directly from raw inputs. Expanding the dataset to include spontaneous and emotionally expressive speech, as well as broader phonetic coverage, will enhance ecological validity. Longitudinal studies will track disease progression and therapy response, enabling personalized diagnostic tools and early intervention.
BACKGROUND:Cell Painting, the leading image-based profiling assay, involves staining plated cells with six dyes that mark the different compartments in a cell. Such profiles can then be used to discover connections between samples (whether different cell lines, different genetic treatments, or different compound treatments) as well as to assess particular features impacted by each treatment. Researchers may wish to vary the standard dye panel to assess particular phenotypes, or image cells live while maintaining the ability to cluster profiles overall. METHODS:In this study, we evaluate the performance of dyes that can either replace or augment the traditional Cell Painting dyes or enable tracking live cell dynamics. We perturbed U2OS cells with 90 different compounds and subsequently stained them with either standard Cell Painting dyes (Revvity), or with MitoBrilliant (Tocris) replacing MitoTracker or Phenovue phalloidin 400LS (Revvity) replacing phalloidin. We also tested the live-cell compatible ChromaLive dye (Saguaro). RESULTS:All dye sets effectively separated biological replicates of the same sample vs. negative controls (phenotypic activity), although separating from replicates of all other compounds (phenotypic distinctiveness) proved challenging for all dye sets. While individual dye substitutions within the standard Cell Painting panel had minimal impact on assay performance, the live cell dye exhibited distinct performance profiles across different compound classes compared to the standard panel, with later time points more distinct than earlier ones. DISCUSSION:Substituting MitoBrilliant or Phenovue phalloidin 400LS for standard mitochondrial or actin dyes minimally impacted Cell Painting assay performance. Phenovue phalloidin 400LS offers the advantage of isolating actin features from Golgi or plasma membrane while accommodating an additional 568 nm dye. Live cell imaging, enabled by ChromaLive dye, provides real-time assessment of compound-induced morphological changes. Combining this with the standard Cell Painting assay significantly expands the feature space for enhanced cellular profiling. Our findings provide data-driven options for researchers selecting dye sets for image-based profiling.
Predicting drug efficacy and safety in vivo requires information on biological responses (e.g., cell morphology and gene expression) to small molecule perturbations. However, current molecular representation learning methods do not provide a comprehensive view of cell states under these perturbations and struggle to remove noise, hindering model generalization. We introduce the Information Alignment (InfoAlign) approach to learn molecular representations through the information bottleneck method in cells. We integrate molecules and cellular response data as nodes into a context graph, connecting them with weighted edges based on chemical, biological, and computational criteria. For each molecule in a training batch, InfoAlign optimizes the encoder's latent representation with a minimality objective to discard redundant structural information. A sufficiency objective decodes the representation to align with different feature spaces from the molecule's neighborhood in the context graph. We demonstrate that the proposed sufficiency objective for alignment is tighter than existing encoder-based contrastive methods. Empirically, we validate representations from InfoAlign in two downstream applications: molecular property prediction against up to 27 baseline methods across four datasets, plus zero-shot molecule-morphology matching.
Image-based cell profiling is a powerful tool that compares perturbed cell populations by measuring thousands of single-cell features and summarizing them into profiles. Typically a sample is represented by averaging across cells, but this fails to capture the heterogeneity within cell populations. We introduce CytoSummaryNet: a Deep Sets-based approach that improves mechanism of action prediction by 30-68% in mean average precision compared to average profiling on a public dataset. CytoSummaryNet uses self-supervised contrastive learning in a multiple-instance learning framework, providing an easier-to-apply method for aggregating single-cell feature data than previously published strategies. Interpretability analysis suggests that the model achieves this improvement by downweighting small mitotic cells or those with debris and prioritizing large uncrowded cells. The approach requires only perturbation labels for training, which are readily available in all cell profiling datasets. CytoSummaryNet offers a straightforward post-processing step for single-cell profiles that can significantly boost retrieval performance on image-based profiling datasets.
Identifying how a given chemical of interest exerts its impact on biological systems is a critical step in developing new medicines and chemical products. The mechanism of a query compound of interest can sometimes be identified when its image-based morphological profile matches a compound in a library of well-annotated compound profiles. In this study, we demonstrate a significant improvement in classification performance by incorporating side information: gene representations. We generate these representations using the morphological profiles of cells where the level of a single gene’s expression has been artificially increased or decreased. The genes are selected as those encoding known protein targets of annotated compounds in the library. A transformer model is trained to classify gene-compound pairs, where each pair represents a potential interaction between a gene and a compound, as true or false. Subsequently, the model generates a ranked list of likely target genes for a previously unseen query compound. Although the strategy exhibits high performance only for compounds that target previously encountered genes – likely due to the limited size of our training dataset – the performance increase demonstrates a notable improvement over simply matching compound profiles directly to compound profiles or to gene profiles. Larger datasets may improve the prediction capabilities of this approach, enabling the prediction of gene targets for novel compounds, which can then be experimentally validated. ### Competing Interest Statement The Authors declare the following competing interests: S.S. and A.E.C. serve as scientific advisors for companies that use image-based profiling and Cell Painting (A.E.C: Recursion, SyzOnc, Quiver Bioscience, S.S.: Waypoint Bio, Dewpoint Therapeutics, Deepcell) and receive honoraria for occasional talks at pharmaceutical and biotechnology companies. All other authors declare no competing interests.
High-throughput image-based profiling platforms are powerful technologies capable of collecting data from billions of cells exposed to thousands of perturbations in a time- and cost-effective manner. Therefore, image-based profiling data has been increasingly used for diverse biological applications, such as predicting drug mechanism of action or gene function. However, batch effects severely limit community-wide efforts to integrate and interpret image-based profiling data collected across different laboratories and equipment. To address this problem, we benchmark ten high-performing single-cell RNA sequencing (scRNA-seq) batch correction techniques, representing diverse approaches, using a newly released Cell Painting dataset, JUMP. We focus on five scenarios with varying complexity, ranging from batches prepared in a single lab over time to batches imaged using different microscopes in multiple labs. We find that Harmony and Seurat RPCA are noteworthy, consistently ranking among the top three methods for all tested scenarios while maintaining computational efficiency. Our proposed framework, benchmark, and metrics can be used to assess new batch correction methods in the future. This work paves the way for improvements that enable the community to make the best use of public Cell Painting data for scientific discovery.
The identification of genetic and chemical perturbations with similar impacts on cell morphology can elucidate compounds' mechanisms of action or novel regulators of genetic pathways. Research on methods for identifying such similarities has lagged due to a lack of carefully designed and well-annotated image sets of cells treated with chemical and genetic perturbations. Here we create such a Resource dataset, CPJUMP1, in which each perturbed gene's product is a known target of at least two chemical compounds in the dataset. We systematically explore the directionality of correlations among perturbations that target the same protein encoded by a given gene, and we find that identifying matches between chemical and genetic perturbations is a challenging task. Our dataset and baseline analyses provide a benchmark for evaluating methods that measure perturbation similarities and impact, and more generally, learn effective representations of cellular state from microscopy images. Such advancements would accelerate the applications of image-based profiling of cellular states, such as uncovering drug mode of action or probing functional genomics. The CPJUMP1 Resource comprises Cell Painting images and profiles of 75 million cells treated with hundreds of chemical and genetic perturbations. The dataset enables exploration of their relationships and lays the foundation for the development of advanced methods to match perturbations.
Background: Early prostate cancer diagnosis from BP-MRI studies constitutes the new guidelines in PI-RADS-2 protocol. From such sequences, malignant lesions are characterized by morphological and cellular density properties. Nonetheless, such characterization is sensible to high variability among different MRI sequences and prostate zones, which often results in misdiagnosis. Current supervised deep learning representations have shown promising results to support the diagnosis. Nevertheless, such strategies require a huge amount of annotated MRI lesions. Methods: This work introduces a weakly supervised learning approach from a deep BP-MRI representation to classify malignant lesions, overcoming supervised deep learning approaches that use MP-MRI. Redundant and rich tissue patches are taken from the prostate gland, allowing to adjust a representation to discriminate between lesions and healthy tissue. This pretext task is performed under a contrastive learning scheme, learning an embedding projection that groups similar patches while maximizing the distance among different classes. Then, from such representation, it is carried out a fine-tuning process to discriminate between benign and malignant lesions related to prostate cancer lesions. Results: The proposed approach outperformed baseline studies in a public dataset, achieving a ROC-AUC of 0.85 using the 80% of the available annotated lesions. Also, using 20% of the lesions, the proposed strategy achieved a ROC-AUC of 0.80, being a promising result to transfer models to the clinical routine. Conclusions: The projected contrastive embedded space better differentiates malignant lesion regions. Also, this representation could be transferred in scenarios with scarce labeled data, approaching self-supervised learning of raw original data.
Image-based or morphological profiling is a rapidly expanding field wherein cells are "profiled" by extracting hundreds to thousands of unbiased, quantitative features from images of cells that have been perturbed by genetic or chemical perturbations. The Cell Painting assay is the most popular imaged-based profiling assay wherein six small-molecule dyes label eight cellular compartments and thousands of measurements are made, describing quantitative traits such as size, shape, intensity, and texture within the nucleus, cytoplasm, and whole cell (Cimini et al., 2023). We have created the Cell Painting Gallery, a publicly available collection of Cell Painting datasets, with granular dataset descriptions and access instructions. It is hosted by AWS on the Registry of Open Data (RODA). As of January 2024, the Cell Painting Gallery holds 656 terabytes (TB) of image and associated numerical data. It includes the largest publicly available Cell Painting dataset, in terms of perturbations tested (Joint Undertaking for Morphological Profiling or JUMP (Chandrasekaran et al., 2023)), along with many other canonical datasets using Cell Painting, close derivatives of Cell Painting (such as LipocyteProfiler (Laber et al., 2023) and Pooled Cell Painting (Ramezani et al., 2023)).
Drug-target interaction (DTI) prediction is crucial for identifying new therapeutics and detecting mechanisms of action. While structure-based methods accurately model physical interactions between a drug and its protein target, cell-based assays such as Cell Painting can better capture complex DTI interactions. This paper introduces MOTIVE, a Morphological cOmpound Target Interaction Graph dataset comprising Cell Painting features for 11,000 genes and 3,600 compounds, along with their relationships extracted from seven publicly available databases. We provide random, cold-source (new drugs), and cold-target (new genes) data splits to enable rigorous evaluation under realistic use cases. Our benchmark results show that graph neural networks that use Cell Painting features consistently outperform those that learn from graph structure alone, feature-based models, and topological heuristics. MOTIVE accelerates both graph ML research and drug discovery by promoting the development of more reliable DTI prediction models. MOTIVE resources are available at https://github.com/carpenter-singh-lab/motive.
Image-based profiling has emerged as a powerful technology for various steps in basic biological and pharmaceutical discovery, but the community has lacked a large, public reference set of data from chemical and genetic perturbations. Here we present data generated by the Joint Undertaking for Morphological Profiling (JUMP)-Cell Painting Consortium, a collaboration between 10 pharmaceutical companies, six supporting technology companies, and two non-profit partners. When completed, the dataset will contain images and profiles from the Cell Painting assay for over 116,750 unique compounds, over-expression of 12,602 genes, and knockout of 7,975 genes using CRISPR-Cas9, all in human osteosarcoma cells (U2OS). The dataset is estimated to be 115 TB in size and capturing 1.6 billion cells and their single-cell profiles. File quality control and upload is underway and will be completed over the coming months at the Cell Painting Gallery: https://registry.opendata.aws/cellpainting-gallery . A portal to visualize a subset of the data is available at https://phenaid.ardigen.com/jumpcpexplorer/ .
Objective.Multi-parametric magnetic resonance imaging (MP-MRI) has played an important role in prostate cancer diagnosis. Nevertheless, in the clinical routine, these sequences are principally analyzed from expert observations, which introduces an intrinsic variability in the diagnosis. Even worse, the isolated study of these MRI sequences trends to false positive detection due to other diseases that share similar radiological findings. Hence, the main objective of this study was to design, propose and validate a deep multimodal learning framework to support MRI-based prostate cancer diagnosis using cross-correlation modules that fuse MRI regions, coded from independent MRI parameter branches.Approach.This work introduces a multimodal scheme that integrates MP-MRI sequences and allows to characterize prostate lesions related to cancer disease. For doing so, potential 3D regions were extracted around expert annotations over different prostate zones. Then, a convolutional representation was obtained from each evaluated sequence, allowing a rich and hierarchical deep representation. Each convolutional branch representation was integrated following a special inception-like module. This module allows a redundant non-linear integration that preserves textural spatial lesion features and could obtain higher levels of representation.Main results.This strategy enhances micro-circulation, morphological, and cellular density features, which thereafter are integrated according to an inception late fusion strategy, leading to a better differentiation of prostate cancer lesions. The proposed strategy achieved a ROC-AUC of 0.82 over the PROSTATEx dataset by fusing regions ofKtransand apparent diffusion coefficient (ADC) maps coded from DWI-MRI.Significance.This study conducted an evaluation about how MP-MRI parameters can be fused, through a deep learning representation, exploiting spatial correlations among multiple lesion observations. The strategy, from a multimodal representation, learns branches representations to exploit radio-logical findings from ADC andKtrans. Besides, the proposed strategy is very compact (151 630 trainable parameters). Hence, the methodology is very fast in training (3 s for an epoch of 320 samples), being potentially applicable in clinical scenarios.
Most variants in most genes across most organisms have an unknown impact on the function of the corresponding gene. This gap in knowledge is especially acute in cancer, where clinical sequencing of tumors now routinely reveals patient-specific variants whose functional impact on the corresponding gene is unknown, impeding clinical utility. Transcriptional profiling was able to systematically distinguish these variants of unknown significance (VUS) as impactful vs. neutral in an approach called expression-based variant-impact phenotyping (eVIP). We profiled a set of lung adenocarcinoma-associated somatic variants using Cell Painting, a morphological profiling assay that captures features of cells based on microscopy using six stains of cell and organelle components. Using deep-learning-extracted features from each cell’s image, we found that cell morphological profiling (cmVIP) can predict variants’ functional impact and, particularly at the single-cell level, reveals biological insights into variants which can be explored in our public online portal. Given its low cost, convenient implementation, and single-cell resolution, cmVIP profiling therefore seems promising as an avenue for using non-gene-specific assays to systematically assess the impact of variants, including disease-associated alleles, on gene function.
Clinically significant regions (CSR), captured over multi-parametric MRI (mp-MRI) images, have emerged as a potential screening test for early prostate cancer detection and characterization. These sequences are able to quantify morphology, micro-circulation, and cellular density patterns that might be related to cancer disease. Nonetheless, this evaluation is mainly carried out by expert radiologists, introducing inter-reader variability in the diagnosis. Therefore, different deep learning models were proposed to support the diagnosis, but a proper representation of prostate lesions remains limited due to the non-alignment among sequences and the dependency of considerable amounts of labeled data for learning. The main limitation of such representation lies in the cross-entropy minimization that only exploits inter-class variation, being insufficient data augmentation and transfer learning strategies. This work introduces a Supervised Contrastive Learning (SCL) strategy that fully exploits the inter and intra-class variability of prostate lesions to robustly represent MRI regions. This strategy extracts lesion sample tuples, with positive and negative labels, regarding a query lesion. Such tuples are involved into an easy-positive, and semi-hard negative mining to project samples that better update the deep representation. The proposed learning strategy achieved an average ROC-AVC of 0.82, to characterize prostate cancer in MRI, using only the 60% of the available annotated data. Clinical relevance - A robust learning scheme that properly finds representations in limited data scenarios to classify clinically significant MRI regions on prostate cancer.
This paper considers the problem of leveraging multiple sources of information or data modalities (e.g., images and text) in neural networks. We define a novel model called gated multimodal unit (GMU), designed as an internal unit in a neural network architecture whose purpose is to find an intermediate representation based on a combination of data from different modalities.The GMU learns to decide how modalities influence the activation of the unit using multiplicative gates.The GMU can be used as a building block for different kinds of neural networks and can be seen as a form of intermediate fusion. The model was evaluated on two multimodal learning tasks in conjunction with fully connected and convolutional neural networks. We compare the GMU with other early- and late-fusion methods, outperforming classification scores in two benchmark datasets: MM-IMDb and DeepScene.
Magnetic resonance imaging (MRI) plays a valuable role in many task related with characterization of prostate cancer lesions. Recently, the DCE-MRI (Dynamic contrast Enhanced) has allowed to visualize and localize potential tumor regions. Specifically, Ktrans, from DCR-MRI, has shown to be a powerful pharmacokinetic parameter that allows to characterize tumor biology and to detect treatment responses from reconstructed coefficient maps of capillary permeability. Nevertheless, even expert-based analysis of Ktrans sequences are subject to a large false positive findings (FPF). In much of such cases, the prostate angiogenesis, or benign prostatic hyperplasia (BPH) regions are misclassified as cancer findings. This work introduces a robust deep convolutional strategy that characterizes Ktrans regions and allows an automatic prediction of cancer findings. The proposed strategy was validated over the SPIE-AAPM-NCI PROSTATEx public dataset with 320 multimodal images on peripheral, transitional and anterior fibromuscular stroma regions. The best configuration of proposal strategy achieved an area under the ROC curve (AUC) of 0.74. Additionally, the proposed strategy achieved a proper characterization by using mainly Ktrans information that together with T2-MRI-transaxial overcome baseline strategies that use additional modalities of MRI.
Representation learning methods have received a lot of attention by researchers and practitioners because of their successful application to complex problems in areas such as computer vision, speech recognition and text processing [1]. Many of these promising results are due to the development of methods to automatically learn the representation of complex objects directly from large amounts of sample data [2]. These efforts have concentrated on data involving one type of information (images, text, speech, etc.), despite data being naturally multimodal. Multimodality refers to the fact that the same real-world concept can be described by different views or data types. Addressing multimodal automatic analysis faces three main challenges: feature learning and extraction, modeling of relationships between data modalities and scalability to large multimodal collections [3, 4]. This research considers the problem of leveraging multiple sources of information or data modalities in neural networks. It defines a novel model called gated multimodal unit (GMU), designed as an internal unit in a neural network architecture whose purpose is to find an intermediate representation based on a combination of data from different modalities. The GMU learns to decide how modalities influence the activation of the unit using multiplicative gates. The GMU can be used as a building block for different kinds of neural networks and can be seen as a form of intermediate fusion. The model was evaluated on four supervised learning tasks in conjunction with fully-connected and convolutional neural networks. We compare the GMU with other early and late fusion methods, outperforming classification scores in the evaluated datasets. Strategies to understand how the model gives importance to each input were also explored. By measuring correlation between gate activations and predictions, we were able to associate modalities with classes. It was found that some classes were more correlated with some particular modality. Interesting findings in genre prediction show, for instance, that the model associates the visual information with animation movies while textual information is more associated with drama or romance movies. During the development of this project, three new benchmark datasets were built and publicly released. The BCDR-F03 dataset which contains 736 mammography images and serves as benchmark for mass lesion classification. The MM-IMDb dataset containing around 27000 movie plots, poster along with 50 metadata annotations and that motivates new research in multimodal analysis. And the Goodreads dataset, a collection of 1000 books that encourages the research on success prediction based on the book content. This research also facilitates reproducibility of the present work by releasing source code implementation of the proposed methods.