Integrating single-cell omics data at an atlas scale enhances our understanding of cell types and disease mechanisms. However, the integration of data processed by different normalization methods can lead to biases, such as unexpected batch effects and gene expression distortion, leading to misinterpretations in downstream analysis. To address these challenges, we present scDenorm, an algorithm that reverts delta-method normalized single-cell omics data to raw counts, preserving the integrity of the original measurements and ensuring consistent data processing during integration. We evaluated scDenorm's performance on large-scale datasets and benchmarked its impact on data integration and downstream analysis across 3 datasets.
Expression Atlas (https://www.ebi.ac.uk/gxa/home) is EMBL-EBI's comprehensive knowledgebase for gene and protein expression across tissues, cell types, conditions, and multiple species. Since our last update, Expression Atlas has expanded substantially in both content and functionality, now comprising >4500 studies from 67 species, with increased proteomics coverage and updated Genotype-Tissue Expression (GTEx) tissue profiles. The resource also includes hundreds of single-cell RNA-seq experiments spanning 21 species, among them externally analysed community datasets such as Tabula Sapiens and GTEx single-nucleus profiles, allowing exploration of curated atlases while maintaining their original analytical framework. Key methodological advances include a new marker gene analysis module for bulk baseline experiments, alongside workflow updates that improve reproducibility. Expression Atlas data are integrated into EMBL-EBI resources such as Ensembl, UniProt, and Europe PMC and disseminated through collaboration with model organism communities such as FlyBase and Gramene. The resource also supports translational research through the European Diagnostic Transcriptomic Library and integration with the Open Targets platform. Future directions include modernizing analysis pipelines, enhancing programmatic access, and delivering AI-ready data formats, strengthening Expression Atlas as a findable, accessible, interoperable, and reusable (FAIR) community-driven resource for both fundamental and translational discovery.
Disease-Free Survival outcomes amongst genomic groups when applied to the 247 VHL mutated ccRCCs from the TCGA dataset
We introduce bia-binder (BioImage Archive Binder), an open-source, cloud-architectured, and web-based coding environment tailored to bioimage analysis that is freely accessible to all researchers. The service generates easy-to-use Jupyter Notebook coding environments hosted on EMBL-EBI's Embassy Cloud, which provides significant computational resources. The bia-binder architecture is free, open-source and publicly available for deployment. It features fast and direct access to images in the BioImage Archive, the Image Data Resource, and the BioStudies databases. We believe that this service can play a role in mitigating the current inequalities in access to scientific resources across academia. As bia-binder produces permanent links to compiled coding environments, we foresee the service to become widely-used within the community and enable exploratory research. bia-binder is built and deployed using helmsman and helm and released under the MIT licence. It can be accessed at binder.bioimagearchive.org and runs on any standard web browser.
The cross-species comparison of expression profiles uncovers functional similarities and differences between cell types and helps refine their evolutionary relationships. Current analysis strategies typically follow the ortholog conjecture, which posits that the expression of orthologous genes is most similar between species. However, the extent to which this holds true at different evolutionary distances is unknown. Here, we systematically explore the ortholog conjecture in comparative single-cell transcriptomics. We devise a robust analytical framework, GeneSpectra, to classify genes by expression specificity and distribution across cell types. Our analysis reveals that genes expressed ubiquitously across nearly all cell types exhibit strong conservation of this pattern across species, as do genes with high expression specificity. In contrast, genes within intermediate specificity fluctuate between classes. As expected, ortholog expression becomes more divergent with increased species distance. We also find an overall correlation between similarity in expression profiles and sequence conservation. Finally, our results allow identifying gene classes with the highest probability of expression pattern conservation that are most useful for cell type alignment between species. Calibrating reliance on the ortholog conjecture for individual genes, we thus provide a comprehensive framework for the comparative analysis of single-cell data.
Computational comparison of single cell expression profiles cross-species uncovers functional similarities and differences between cell types. Importantly, it offers the potential to refine evolutionary relationships based on gene expression. Current analysis strategies are limited by the strong hypothesis of ortholog conjecture, and lose expression information given by non-orthologs. To address this, we devise a novel analytical framework that redefines the analysis paradigm. This framework robustly classifies genes by expression specificity and distribution across cell types, allowing for a dataset-specific reassessment of the ortholog conjecture by evaluating the degree of ortholog class conservation. We utilise the gene classes to decode species effects on cross-species transcriptomics space, and compare sequence conservation with expression specificity similarity across different types of orthologs. We develop contextualised cell type similarity measurements while considering species-unique genes and non-one-to-one orthologs. Finally, we consolidate gene classification results into a knowledge graph, allowing hierarchical depiction of cell types and orthologous groups, and continuous integration of new data. ### Competing Interest Statement The authors have declared no competing interest.
Motivation:Cell-type deconvolution methods aim to infer cell composition from bulk transcriptomic data. The proliferation of developed methods coupled with inconsistent results obtained in many cases, highlights the pressing need for guidance in the selection of appropriate methods. Additionally, the growing accessibility of single-cell RNA sequencing datasets, often accompanied by bulk expression from related samples enable the benchmark of existing methods. Results:In this study, we conduct a comprehensive assessment of 31 methods, utilizing single-cell RNA-sequencing data from diverse human and mouse tissues. Employing various simulation scenarios, we reveal the efficacy of regression-based deconvolution methods, highlighting their sensitivity to reference choices. We investigate the impact of bulk-reference differences, incorporating variables such as sample, study and technology. We provide validation using a gold standard dataset from mononuclear cells and suggest a consensus prediction of proportions when ground truth is not available. We validated the consensus method on data from the stomach and studied its spillover effect. Importantly, we propose the use of the critical assessment of transcriptomic deconvolution (CATD) pipeline which encompasses functionalities for generating references and pseudo-bulks and running implemented deconvolution methods. CATD streamlines simultaneous deconvolution of numerous bulk samples, providing a practical solution for speeding up the evaluation of newly developed methods. Availability and implementation:https://github.com/Papatheodorou-Group/CATD_snakemake.
Expression Atlas (www.ebi.ac.uk/gxa) and its newest counterpart the Single Cell Expression Atlas (www.ebi.ac.uk/gxa/sc) are EMBL-EBI's knowledgebases for gene and protein expression and localisation in bulk and at single cell level. These resources aim to allow users to investigate their expression in normal tissue (baseline) or in response to perturbations such as disease or changes to genotype (differential) across multiple species. Users are invited to search for genes or metadata terms across species or biological conditions in a standardised consistent interface. Alongside these data, new features in Single Cell Expression Atlas allow users to query metadata through our new cell type wheel search. At the experiment level data can be explored through two types of dimensionality reduction plots, t-distributed Stochastic Neighbor Embedding (tSNE) and Uniform Manifold Approximation and Projection (UMAP), overlaid with either clustering or metadata information to assist users' understanding. Data are also visualised as marker gene heatmaps identifying genes that help confer cluster identity. For some data, additional visualisations are available as interactive cell level anatomograms and cell type gene expression heatmaps.
Cell type deconvolution methods can impute cell proportions from bulk transcriptomics data, revealing changes in disease progression or organ development. But benchmarking studies often use simulated bulk data from the same source as the reference, which limits its application scenarios. This study examines batch effects in deconvolution and introduces SCCAF-D, a computational workflow that ensures a Pearson Correlation Coefficient above 0.75 across simulated and real bulk data for various tissue types. Applied to non-alcoholic fatty liver disease, SCCAF-D unveils meaningful insights into changes in cell proportions during disease progression.
Melanoma is the deadliest form of skin cancer and develops from the melanocytes that are responsible for the pigmentation of the skin. The skin is also a highly regenerative organ, harboring a pool of undifferentiated melanocyte stem cells that proliferate and differentiate into mature melanocytes during regenerative processes in the adult. Melanoma and melanocyte regeneration share remarkable cellular features, including activation of cell proliferation and migration. Yet, melanoma considerably differs from the regenerating melanocytes with respect to abnormal proliferation, invasive growth, and metastasis. Thus, it is likely that at the cellular level, melanoma resembles early stages of melanocyte regeneration with increased proliferation but separates from the later melanocyte regeneration stages due to reduced proliferation and enhanced differentiation. Here, by exploiting the zebrafish melanocytes that can efficiently regenerate and be induced to undergo malignant melanoma, we unravel the transcriptome profiles of the regenerating melanocytes during early and late regeneration and the melanocytic nevi and malignant melanoma. Our global comparison of the gene expression profiles of melanocyte regeneration and nevi/melanoma uncovers the opposite regulation of a substantial number of genes related to Wnt signaling and transforming growth factor beta (TGF-β)/(bone morphogenetic protein) BMP signaling pathways between regeneration and cancer. Functional activation of canonical Wnt or TGF-β/BMP pathways during melanocyte regeneration promoted melanocyte regeneration but potently suppressed the invasiveness, migration, and proliferation of human melanoma cells in vitro and in vivo. Therefore, the opposite regulation of signaling mechanisms between melanocyte regeneration and melanoma can be exploited to stop tumor growth and develop new anti-cancer therapies.
Abstract Cell-type deconvolution methods aim to infer cell-type composition and the cell abundances from bulk transcriptomic data. The proliferation of currently developed methods, coupled with the inconsistent results obtained in many cases, highlights the pressing need for guidance in the selection of appropriate methods. Previous proposed tests have primarily been focused on simulated data and have seen limited application to actual datasets. The growing accessibility of systematic single-cell RNA sequencing datasets, often accompanied by bulk RNA sequencing from related or matched samples, makes it possible to benchmark the existing deconvolution methods more objectively. Here, we propose a comprehensive assessment of 29 available deconvolution methods, leveraging single-cell RNA-sequencing data from different tissues. We offer a new comprehensive framework to evaluate deconvolution across a wide range of simulation scenarios and we show that single-cell regression-based deconvolution methods perform well but their performance is highly dependent on the reference selection and the tissue type. We validate deconvolution results on a gold standard bulk PBMC dataset with well known cell-type proportions and suggest a novel methodology for consensus prediction of cell-type proportions for cases when ground truth is not available. Our study also explores the significant impact of various batch effects on deconvolution, including those associated with sample, study, and technology, which have been previously overlooked. The evaluation of cell-type prediction methods is provided in a modularised pipeline for reproducibility (https://github.com/Functional-Genomics/CATD_snakemake). Lastly, we suggest that the Critical Assessment of Transcriptomic Deconvolution (CATD) pipeline can be employed for the efficient, simultaneous deconvolution of hundreds of real bulk samples, utilising various references. We envision it to be used for speeding up the evaluation of newly published methods in the future and for systematic deconvolution of real samples.
Motivation: The nuclear pore complex (NPC) is the only passageway for macromolecules between nucleus and cytoplasm, and an important reference standard in microscopy: it is massive and stereotypically arranged. The average architecture of NPC proteins has been resolved with pseudoatomic precision, however observed NPC heterogeneities evidence a high degree of divergence from this average. Single-molecule localization microscopy (SMLM) images NPCs at protein-level resolution, whereupon image analysis software studies NPC variability. However, the true picture of this variability is unknown. In quantitative image analysis experiments, it is thus difficult to distinguish intrinsically high SMLM noise from variability of the underlying structure. Results: We introduce CIR4MICS ('ceramics', Configurable, Irregular Rings FOR MICroscopy Simulations), a pipeline that synthesizes ground truth datasets of structurally variable NPCs based on architectural models of the true NPC. Users can select one or more N- or C-terminally tagged NPC proteins, and simulate a wide range of geometric variations. We also represent the NPC as a spring-model such that arbitrary deforming forces, of user-defined magnitudes, simulate irregularly shaped variations. Further, we provide annotated reference datasets of simulated human NPCs, which facilitate a side-by-side comparison with real data. To demonstrate, we synthetically replicate a geometric analysis of real NPC radii and reveal that a range of simulated variability parameters can lead to observed results. Our simulator is therefore valuable to test the capabilities of image analysis methods, as well as to inform experimentalists about the requirements of hypothesis-driven imaging studies. Availability and implementation: Code: https://github.com/uhlmanngroup/cir4mics. Simulated data: BioStudies S-BSST1058.
The growing number of available single-cell gene expression datasets from different species creates opportunities to explore evolutionary relationships between cell types across species. Cross-species integration of single-cell RNA-sequencing data has been particularly informative in this context. However, in order to do so robustly it is essential to have rigorous benchmarking and appropriate guidelines to ensure that integration results truly reflect biology. Here, we benchmark 28 combinations of gene homology mapping methods and data integration algorithms in a variety of biological settings. We examine the capability of each strategy to perform species-mixing of known homologous cell types and to preserve biological heterogeneity using 9 established metrics. We also develop a new biology conservation metric to address the maintenance of cell type distinguishability. Overall, scANVI, scVI and SeuratV4 methods achieve a balance between species-mixing and biology conservation. For evolutionarily distant species, including in-paralogs is beneficial. SAMap outperforms when integrating whole-body atlases between species with challenging gene homology annotation. We provide our freely available cross-species integration and assessment pipeline to help analyse new data and develop new algorithms.