Cell Painting images offer valuable insights into a cell's state and enable many biological applications, but publicly available arrayed datasets only include hundreds of genes perturbed. The JUMP Cell Painting Consortium perturbed roughly 75% of the protein-coding genome in human U-2 OS cells, generating a rich resource of single-cell images and extracted features. These profiles capture the phenotypic impacts of perturbing 15,243 human genes, including overexpressing 12,609 genes (using open reading frames) and knocking out 7,975 genes (using CRISPR-Cas9). Here we mitigated technical artifacts by rigorously evaluating data processing options and validated the dataset's robustness and biological relevance. Analysis of phenotypic profiles revealed previously undiscovered gene clusters and functional relationships, including those associated with mitochondrial function, cancer and neural processes. The JUMP Cell Painting genetic dataset is a valuable resource for exploring gene relationships and uncovering previously unknown functions.
PurposeDeveloping innovative precision and personalized cancer therapeutics is essential to enhance cancer survivability, particularly for prevalent cancer types such as colorectal cancer. This study aims to demonstrate various approaches for discovering new targets for precision therapies using artificial intelligence (AI) on a Polish cohort of colorectal cancer patients. MethodsWe analyzed 71 patients with histopathologically confirmed advanced resectional colorectal adenocarcinoma. Whole exome sequencing was performed on tumor and peripheral blood samples, while RNA sequencing (RNAseq) was conducted on tumor samples. We employed three approaches to identify potential targets for personalized and precision therapies. First, using our in-house neoantigen calling pipeline, ARDentify, combined with an AI-based model trained on immunopeptidomics mass spectrometry data (ARDisplay), we identified neoepitopes in the cohort. Second, based on recurrent mutations found in our patient cohort, we selected corresponding cancer cell lines and utilized knock-out gene dependency scores to identify synthetic lethality genes. Third, an AI-based model trained on cancer cell line data was employed to identify cell lines with genomic profiles similar to selected patients. Copy number variants and recurrent single nucleotide variants in these cell lines, along with gene dependency data, were used to find personalized synthetic lethality pairs. ResultsWe identified approximately 8,700 unique neoepitopes, but none were shared by more than two patients, indicating limited potential for shared neoantigenic targets across our cohort. Additionally, we identified three synthetic lethality pairs: the well-known APC-CTNNB1 and BRAF-DUSP4 pairs, along with the recently described APC-TCF7L2 pair, which could be significant for patients with APC and BRAF variants. Furthermore, by leveraging the identification of similar cancer cell lines, we uncovered a potential gene pair, VPS4A and VPS4B, with therapeutic implications. ConclusionOur study highlights three distinct approaches for identifying potential therapeutic targets in cancer patients. Each approach yielded valuable insights into our cohort, underscoring the relevance and utility of these methodologies in the development of precision and personalized cancer therapies. Importantly, we developed a novel AI model that aligns tumors with representative cell lines using RNAseq and methylation data. This model enables us to identify cell lines closely resembling patient tumors, facilitating accurate selection of models needed for in vitro validation.
Background and Aims Histological disease activity in inflammatory bowel disease [IBD] is associated with clinical outcomes and is an important endpoint in drug development. We developed deep learning models for automating histological assessments in IBD. Methods Histology images of intestinal mucosa from phase 2 and phase 3 clinical trials in Crohn’s disease [CD] and ulcerative colitis [UC] were used to train artificial intelligence [AI] models to predict the Global Histology Activity Score [GHAS] for CD and Geboes histopathology score for UC. Three AI methods were compared. AI models were evaluated on held-back testing sets, and model predictions were compared against an expert central reader and five independent pathologists. Results The model based on multiple instance learning and the attention mechanism [SA-AbMILP] demonstrated the best performance among competing models. AI-modelled GHAS and Geboes subgrades matched central readings with moderate to substantial agreement, with accuracies ranging from 65% to 89%. Furthermore, the model was able to distinguish the presence and absence of pathology across four selected histological features, with accuracies for colon in both CD and UC ranging from 87% to 94% and for CD ileum ranging from 76% to 83%. For both CD and UC and across anatomical compartments [ileum and colon] in CD, comparable accuracies against central readings were found between the model-assigned scores and scores by an independent set of pathologists. Conclusions Deep learning models based upon GHAS and Geboes scoring systems were effective at distinguishing between the presence and absence of IBD microscopic disease activity.
Image-based profiling has emerged as a powerful technology for various steps in basic biological and pharmaceutical discovery, but the community has lacked a large, public reference set of data from chemical and genetic perturbations. Here we present data generated by the Joint Undertaking for Morphological Profiling (JUMP)-Cell Painting Consortium, a collaboration between 10 pharmaceutical companies, six supporting technology companies, and two non-profit partners. When completed, the dataset will contain images and profiles from the Cell Painting assay for over 116,750 unique compounds, over-expression of 12,602 genes, and knockout of 7,975 genes using CRISPR-Cas9, all in human osteosarcoma cells (U2OS). The dataset is estimated to be 115 TB in size and capturing 1.6 billion cells and their single-cell profiles. File quality control and upload is underway and will be completed over the coming months at the Cell Painting Gallery: https://registry.opendata.aws/cellpainting-gallery . A portal to visualize a subset of the data is available at https://phenaid.ardigen.com/jumpcpexplorer/ .
Even though cancer is driven by genomic alterations, the chain functions causing this disease are largely carried out by proteins. Proteins are also typically targeted in treatment. However, proteomes are harder and more expensive to measure than genomes and transcriptomes. Thus, it would be very valuable to accurately estimate protein levels using other omics data. To catalise developments of solutions to this problem, and to answer fundamental questions about transcriptional and translational control, we leveraged the power of crowdsourcing via a collaborative competition: The NCI-CPTAC DREAM Proteogenomics Challenge. The best performance for predicting protein and phosphorylation levels was achieved by an ensemble of models including as predictors transcript level of the corresponding genes, interaction between genes, conservation across tumor types and, for phosphorylation prediction, phosphosite proximity. Proteins from metabolic pathways were the best predicted, whereas complex proteins were the least well predicted. However, the performance even of the best performing model was modest, suggesting that the level for many proteins are strongly regulated through translational control and degradation. From the best-performing model, we identified common predictors, which are predictive of survival outcome. Our results shed light on the potential application of computational models to large scale proteogenomic characterization of cancer in order to better understand signaling dysregulation mechanisms in the disease.
Cancer is driven by genomic alterations, but the processes causing this disease are largely performed by proteins. However, proteins are harder and more expensive to measure than genes and transcripts. To catalyze developments of methods to infer protein levels from other omics measurements, we leveraged crowdsourcing via the NCI-CPTAC DREAM proteogenomic challenge. We asked for methods to predict protein and phosphorylation levels from genomic and transcriptomic data in cancer patients. The best performance was achieved by an ensemble of models, including as predictors transcript level of the corresponding genes, interaction between genes, conservation across tumor types, and phosphosite proximity for phosphorylation prediction. Proteins from metabolic pathways and complexes were the best and worst predicted, respectively. The performance of even the best-performing model was modest, suggesting that many proteins are strongly regulated through translational control and degradation. Our results set a reference for the limitations of computational inference in proteogenomics. A record of this paper's transparent peer review process is included in the Supplemental Information.
Designing a molecule with desired properties is one of the biggest challenges in drug development, as it requires optimization of chemical compound structures with respect to many complex properties. To improve the compound design process, we introduce Mol-CycleGAN—a CycleGAN-based model that generates optimized compounds with high structural similarity to the original ones. Namely, given a molecule our model generates a structurally similar one with an optimized value of the considered property. We evaluate the performance of the model on selected optimization objectives related to structural properties (presence of halogen groups, number of aromatic rings) and to a physicochemical property (penalized logP). In the task of optimization of penalized logP of drug-like molecules our model significantly outperforms previous results.
During the drug design process, one must develop a molecule, which structure satisfies a number of physicochemical properties. To improve this process, we introduce Mol-CycleGAN – a CycleGAN-based model that generates compounds optimized for a selected property, while aiming to retain the already optimized ones. In the task of constrained optimization of penalized logP of drug-like molecules our model significantly outperforms previous results.
To draw inference on serial extremal dependence within heavy-tailed Markov chains, Drees et al., (2015) proposed nonparametric estimators of the spectral tail process. The methodology can be extended to the more general setting of a stationary, regularly varying time series. The large-sample distribution of the estimators is derived via empirical process theory for cluster functionals. The finite-sample performance of these estimators is evaluated via Monte Carlo simulations. Moreover, two different bootstrap schemes are employed which yield confidence intervals for the pre-asymptotic spectral tail process: the stationary bootstrap and the multiplier block bootstrap. The estimators are applied to stock price data to study the persistence of positive and negative shocks.
At high levels, the asymptotic distribution of a stationary, regularly varying Markov chain is conveniently given by its tail process. The latter takes the form of a geometric random walk, the increment distribution depending on the sign of the process at the current state and on the flow of time, either forward or backward. Estimation of the tail process provides a nonparametric approach to analyze extreme values. A duality between the distributions of the forward and backward increments provides additional information that can be exploited in the construction of more efficient estimators. The large-sample distribution of such estimators is derived via empirical process theory for cluster functionals. Their finite-sample performance is evaluated via Monte Carlo simulations involving copula-based Markov models and solutions to stochastic recurrence equations. The estimators are applied to stock price data to study the absence or presence of symmetries in the succession of large gains and losses.
There is an increasing interest to understand the dependence structure of a random vector not only in the center of its distribution but also in the tails. Extreme-value theory tackles the problem of modelling the joint tail of a multivariate distribution by modelling the marginal distributions and the dependence structure separately. For estimating dependence at high levels, the stable tail dependence function and the spectral measure are particularly convenient. These objects also lie at the basis of nonparametric techniques for modelling the dependence among extremes in the max-domain of attraction setting. In case of asymptotic independence, this setting is inadequate, and more refined tail dependence coefficients exist, serving, among others, to discriminate between asymptotic dependence and independence. Throughout, the methods are illustrated on financial data.
The spectral measure plays a key role in the statistical modeling of multivariate extremes. Estimation of the spectral measure is a complex issue, given the need to obey a certain moment condition. We propose a Euclidean likelihood-based estimator for the spectral measure which is simple and explicitly defined, with its expression being free of Lagrange multipliers. Our estimator is shown to have the same limit distribution as the maximum empirical likelihood estimator of Einmahl and Segers (2009). Numerical experiments suggest an overall good performance and identical behavior to the maximum empirical likelihood estimator. We illustrate the method in an extreme temperature data analysis.