The scale of biological datasets now routinely exceeds system memory, making data access rather than model computation the primary bottleneck in training machine-learning models. This bottleneck is particularly acute in biology, where widely used community data formats must support heterogeneous metadata, sparse and dense assays, and downstream analysis within established computational ecosystems. Here we present annbatch, a mini-batch loader native to anndata that enables out-of-core training directly on disk-backed datasets. Across single-cell transcriptomics, microscopy and whole-genome sequencing benchmarks, annbatch increases loading throughput by up to an order of magnitude and shortens training from days to hours, while remaining fully compatible with the scverse ecosystem. Annbatch establishes a practical data-loading infrastructure for scalable biological AI, allowing increasingly large and diverse datasets to be used without abandoning standard biological data formats. Github: https://github.com/scverse/annbatch
Recent advances in multiplexed single-cell transcriptomics experiments are facilitating the high-throughput study of drug and genetic perturbations. However, an exhaustive exploration of the combinatorial perturbation space is experimentally unfeasible, so computational methods are needed to predict, interpret, and prioritize perturbations. Here, we present the compositional perturbation autoencoder (CPA), which combines the interpretability of linear models with the flexibility of deep-learning approaches for single-cell response modeling. CPA encodes and learns transcriptional drug responses across different cell type, dose, and drug combinations. The model produces easy-to-interpret embeddings for drugs and cell types, which enables drug similarity analysis and predictions for unseen dosage and drug combinations. We show that CPA accurately models single-cell perturbations across compounds, doses, species, and time. We further demonstrate that CPA predicts combinatorial genetic interactions of several types, implying that it captures features that distinguish different interaction programs. Finally, we demonstrate that CPA can generate in-silico 5,329 missing genetic combination perturbations (97.6% of all possibilities) with diverse genetic interactions. We envision our model will facilitate efficient experimental design and hypothesis generation by enabling in-silico response prediction at the single-cell level, and thus accelerate therapeutic applications using single-cell technologies.
nbproject is an open-source Python tool to help manage Jupyter notebooks with metadata, dependency, and integrity tracking. A draft-to-publish workflow creates more reproducible notebooks with context. There are a number of approaches to address reproducibility & manageability problems of computational R&D projects. nbproject complements - and should be combined with - approaches that are based on modularizing notebooks into pipelines, containerizing compute environments, or managing notebooks on centralized platforms.
readfcs loads data and metadata from Flow Cytometry Standard (FCS) files efficiently into general DataFrame and AnnData objects. Existing tools, by contrast, provide FCS parsers and data structures for specific downstream applications. readfcs allows to flexibly access data and metadata slots and offers a robust, tested implementation.
Cell biology is fundamentally limited in its ability to collect complete data on cellular phenotypes and the wide range of responses to perturbation. Areas such as computer vision and speech recognition have addressed this problem of characterizing unseen or unlabeled conditions with the combined advances of big data, deep learning, and computing resources in the past 5 years. Similarly, recent advances in machine learning approaches enabled by single-cell data start to address prediction tasks in perturbation response modeling. We first define objectives in learning perturbation response in single-cell omics; survey existing approaches, resources, and datasets (https://github.com/theislab/sc-pert); and discuss how a perturbation atlas can enable deep learning models to construct an informative perturbation latent space. We then examine future avenues toward more powerful and explainable modeling using deep neural networks, which enable the integration of disparate information sources and an understanding of heterogeneous, complex, and unseen systems.
Vaccines against SARS-CoV-2 have shown high efficacy, but immunocompromised participants were excluded from controlled clinical trials. We compared immune responses to the Pfizer/BioNTech mRNA vaccine in solid tumor patients (n=53) on active cytotoxic anti-cancer therapy to a control cohort (n=50) as an observational study. Using live SARS-CoV-2 assays, neutralizing antibodies were detected in 67% and 80% of cancer patients after the first and second immunizations, respectively, with a 3-fold increase in median titers after the booster. Similar trends were observed in serum antibodies against the receptor-binding domain (RBD) and S2 regions of Spike protein, and in IFNγ+ Spike-specific T cells. Yet the magnitude of each of these responses was diminished relative to the control cohort. We therefore quantified RBD- and Spike S1-specific memory B cell subsets as predictors of anamnestic responses to additional immunizations. After the second vaccination, Spike-specific plasma cell-biased memory B cells were observed in most cancer patients at levels similar to those of the control cohort after the first immunization. We initiated an interventional phase 1 trial of a third booster shot (NCT04936997); primary outcomes were immune responses with a secondary outcome of safety. After a third immunization, the 20 participants demonstrated an increase in antibody responses, with a median 3-fold increase in virus-neutralizing titers. Yet no improvement was observed in T cell responses at 1 week after the booster immunization. There were mild adverse events, primarily injection site myalgia, with no serious adverse events after a month of follow-up. These results suggest that a third vaccination improves humoral immunity against COVID-19 in cancer patients on active chemotherapy with no severe adverse events.
Summaryanndata is a Python package for handling annotated data matrices in memory and on disk (github.com/theislab/anndata), positioned between pandas and xarray. anndata offers a broad range of computationally efficient features including, among others, sparse data support, lazy operations, and a PyTorch interface.Statement of needGenerating insight from high-dimensional data matrices typically works through training models that annotate observations and variables via low-dimensional representations. In exploratory data analysis, this involvesiterativetraining and analysis using original and learned annotations and task-associated representations. anndata offers a canonical data structure for book-keeping these, which is neither addressed by pandas (McKinney, 2010), nor xarray (Hoyer & Hamman, 2017), nor commonly-used modeling packages like scikit-learn (Pedregosa et al., 2011).
Background T cell responses are tightly regulated and require a constant balance of signals during the different stages of their activation, expansion, and differentiation. As a result of chronic antigen exposure, T cells become exhausted in solid tumors, preventing them from controlling tumor growth. Methods We identified a transcriptional signature associated with T cell exhaustion in patients with melanoma and used our proprietary machine learning algorithms to predict molecules that would prevent T cell exhaustion and improve T cell function. Among the predictions, an orally available small molecule, Compound A, was highly predicted. Results Compound A was tested in an in vitro T cell Exhaustion assay and shown to prevent loss of proliferation and expression of immune checkpoint receptors. Transcriptionally, Compound A-treated cells looked indistinguishable from conventionally expanded, non-exhausted T cells. However, when assessed in a classical T cell activation assay, Compound A demonstrated dose dependent activity. At low dose, Compound A was immuno-stimulatory, allowing cells to divide further by preventing activation induced cell death. At higher doses, Compound A demonstrated immuno-suppressive activity preventing early CD69 upregulation and T cell proliferation. All together, these observations suggest that Compound A prevented exhaustion with a mechanism of action involving TCR signaling inhibition. While cessation of TCR signaling or rest has been recently associated with improved CAR-T efficacy by preventing or reversing exhaustion during the in vitro manufacturing phase, it is unclear if that mechanism would translate in vivo.Compound A was evaluated in the CT26 and MC38 syngeneic mouse models alongside anti-PD1. At low dose Compound A closely recapitulated anti-PD1 mediated cell behavior changes by scRNA-seq and flow cytometry in CT26 mice. At high dose, Compound A led to the accumulation of naive cells in the tumor microenvironment (TME) confirming the proposed mechanism of action. Low dose treatment was ineffective in MC38 mouse model but a pulsed treatment at high dose also recapitulated anti-PD1 activity in most animals. Importantly, we identified a new T cell population responding to anti-PD1 that was particularly increased in the MC38 mouse model; Compound A treatment also impacted this population. Conclusions These data confirm that mild TCR inhibition either suboptimal or fractionated can prevent exhaustion in vivo. However, this approach has a very limited window of activity between immuno-modulatory and immuno-suppressive effects, thereby limiting potential clinical benefit. Finally, these results demonstrate that our approach and platform was able to predict molecules that would prevent T cell exhaustion in vivo.
Vaccines against severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) have shown high efficacy, but immunocompromised participants were excluded from controlled clinical trials. In this study, we compared immune responses to the BNT162b2 mRNA Coronavirus Disease 2019 vaccine in patients with solid tumors (n = 53) who were on active cytotoxic anti-cancer therapy to a control cohort of participants without cancer (n = 50). Neutralizing antibodies were detected in 67% of patients with cancer after the first immunization, followed by a threefold increase in median titers after the second dose. Similar patterns were observed for spike protein-specific serum antibodies and T cells, but the magnitude of each of these responses was diminished relative to the control cohort. In most patients with cancer, we detected spike receptor-binding domain and other S1-specific memory B cell subsets as potential predictors of anamnestic responses to additional immunizations. We therefore initiated a phase 1 trial for 20 cancer cohort participants of a third vaccine dose of BNT162b2 ( NCT04936997 ); primary outcomes were immune responses, with a secondary outcome of safety. At 1 week after a third immunization, 16 participants demonstrated a median threefold increase in neutralizing antibody responses, but no improvement was observed in T cell responses. Adverse events were mild. These results suggest that a third dose of BNT162b2 is safe, improves humoral immunity against SARS-CoV-2 and could be immunologically beneficial for patients with cancer on active chemotherapy.
23 Recent advances in multiplexing single-cell transcriptomics across experiments are enabling the high24 throughput study of drug and genetic perturbations. However, an exhaustive exploration of the com25 binatorial perturbation space is experimentally unfeasible, so computational methods are needed to 26 predict, interpret and prioritize perturbations. Here, we present the Compositional Perturbation 27 Autoencoder (CPA), which combines the interpretability of linear models with the flexibility of 28 deep-learning approaches for single-cell response modeling. CPA encodes and learns transcriptional 29 drug response across different cell types, doses, and drug combinations. The model produces easy30 to-interpret embeddings for drugs and cell types, allowing drug similarity analysis and predictions 31 for unseen dosages and drug combinations. We show CPA accurately models single-cell perturba32 tions across compounds, dosages, species, and time. We further demonstrate that CPA predicts 33 combinatorial genetic interactions of several types, implying it captures features that distinguish 34 different interaction programs. Finally, we demonstrate CPA allows in-silico generation of 5,329 35 missing combinations (97.6% of all possibilities) with diverse genetic interactions. We envision our 36 model will facilitate efficient experimental design by enabling in-silico response prediction at the 37 single-cell level. 38
RNA velocity has opened up new ways of studying cellular differentiation in single-cell RNA-sequencing data. It describes the rate of gene expression change for an individual gene at a given time point based on the ratio of its spliced and unspliced messenger RNA (mRNA). However, errors in velocity estimates arise if the central assumptions of a common splicing rate and the observation of the full splicing dynamics with steady-state mRNA levels are violated. Here we present scVelo, a method that overcomes these limitations by solving the full transcriptional dynamics of splicing kinetics using a likelihood-based dynamical model. This generalizes RNA velocity to systems with transient cell states, which are common in development and in response to perturbations. We apply scVelo to disentangling subpopulation kinetics in neurogenesis and pancreatic endocrinogenesis. We infer gene-specific rates of transcription, splicing and degradation, recover each cell’s position in the underlying differentiation processes and detect putative driver genes. scVelo will facilitate the study of lineage decisions and gene regulation.
While generative models have shown great success in generating high-dimensional samples conditional on low-dimensional descriptors (learning e.g. stroke thickness in MNIST, hair color in CelebA, or speaker identity in Wavenet), their generation out-of-sample poses fundamental problems. The conditional variational autoencoder (CVAE) as a simple conditional generative model does not explicitly relate conditions during training and, hence, has no incentive of learning a compact joint distribution across conditions. We overcome this limitation by matching their distributions using maximum mean discrepancy (MMD) in the decoder layer that follows the bottleneck. This introduces a strong regularization both for reconstructing samples within the same condition and for transforming samples across conditions, resulting in much improved generalization. We refer to the architecture as \emph{transformer} VAE (trVAE). Benchmarking trVAE on high-dimensional image and tabular data, we demonstrate higher robustness and higher accuracy than existing approaches. In particular, we show qualitatively improved predictions for cellular perturbation response to treatment and disease based on high-dimensional single-cell gene expression data, by tackling previously problematic minority classes and multiple conditions. For generic tasks, we improve Pearson correlations of high-dimensional estimated means and variances with their ground truths from 0.89 to 0.97 and 0.75 to 0.87, respectively.
Existing methods for learning latent representations for single-cell RNA-seq data are based on autoencoders and factor models. However, representations learned by autoencoders are hard to interpret and representations learned by factor models have limited flexibility. Here, we introduce a framework for learning interpretable autoencoders based on regularized linear decoders. It decomposes variation into interpretable components using prior knowledge in the form of annotated feature sets obtained from public databases. Through this, it provides an alternative to enrichment techniques and factor models for the task of explaining observed variation with biological knowledge. Benchmarking our model on two single-cell RNA-seq datasets, we demonstrate how our model outperforms an existing factor model regarding scalability while maintaining interpretability.
Motivation While generative models have shown great success in sampling high-dimensional samples conditional on low-dimensional descriptors (stroke thickness in MNIST, hair color in CelebA, speaker identity in WaveNet), their generation out-of-distribution poses fundamental problems due to the difficulty of learning compact joint distribution across conditions. The canonical example of the conditional variational autoencoder (CVAE), for instance, does not explicitly relate conditions during training and, hence, has no explicit incentive of learning such a compact representation. Results We overcome the limitation of the CVAE by matching distributions across conditions using maximum mean discrepancy in the decoder layer that follows the bottleneck. This introduces a strong regularization both for reconstructing samples within the same condition and for transforming samples across conditions, resulting in much improved generalization. As this amount to solving a style-transfer problem, we refer to the model as transfer VAE (trVAE). Benchmarking trVAE on high-dimensional image and single-cell RNA-seq, we demonstrate higher robustness and higher accuracy than existing approaches. We also show qualitatively improved predictions by tackling previously problematic minority classes and multiple conditions in the context of cellular perturbation response to treatment and disease based on high-dimensional single-cell gene expression data. For generic tasks, we improve Pearson correlations of high-dimensional estimated means and variances with their ground truths from 0.89 to 0.97 and 0.75 to 0.87, respectively. We further demonstrate that trVAE learns cell-type-specific responses after perturbation and improves the prediction of most cell-type-specific genes by 65%. Availability and implementation The trVAE implementation is available via github.com/theislab/trvae. The results of this article can be reproduced via github.com/theislab/trvae_reproducibility.
Accurately modeling cellular response to perturbations is a central goal of computational biology. While such modeling has been based on statistical, mechanistic and machine learning models in specific settings, no generalization of predictions to phenomena absent from training data (out-of-sample) has yet been demonstrated. Here, we present scGen (https://github.com/theislab/scgen), a model combining variational autoencoders and latent space vector arithmetics for high-dimensional single-cell gene expression data. We show that scGen accurately models perturbation and infection response of cells across cell types, studies and species. In particular, we demonstrate that scGen learns cell-type and species-specific responses implying that it captures features that distinguish responding from non-responding genes and cells. With the upcoming availability of large-scale atlases of organs in a healthy state, we envision scGen to become a tool for experimental design through in silico screening of perturbation response in the context of disease and drug treatment.
Single-cell RNA-seq quantifies biological heterogeneity across both discrete cell types and continuous cell transitions. Partition-based graph abstraction (PAGA) provides an interpretable graph-like map of the arising data manifold, based on estimating connectivity of manifold partitions ( https://github.com/theislab/paga ). PAGA maps preserve the global topology of data, allow analyzing data at different resolutions, and result in much higher computational efficiency of the typical exploratory data analysis workflow. We demonstrate the method by inferring structure-rich cell maps with consistent topology across four hematopoietic datasets, adult planaria and the zebrafish embryo and benchmark computational performance on one million neurons.
Accurately modeling cellular response to perturbations is a central goal of computational biology. While such modeling has been proposed based on statistical, mechanistic and machine learning models in specific settings, no generalization of predictions to phenomena absent from training data (‘out-of-sample’) has yet been demonstrated. Here, we present scGen, a model combining variational autoencoders and latent space vector arithmetics for high-dimensional single-cell gene expression data. In benchmarks across a broad range of examples, we show that scGen accurately models dose and infection response of cells across cell types, studies and species. In particular, we demonstrate that scGen learns cell type and species specific response implying that it captures features that distinguish responding from non-responding genes and cells. With the upcoming availability of large-scale atlases of organs in healthy state, we envision scGen to become a tool for experimental design through in silico screening of perturbation response in the context of disease and drug treatment.
Flatworms of the species Schmidtea mediterranea are immortal-adult animals contain a large pool of pluripotent stem cells that continuously differentiate into all adult cell types. Therefore, single-cell transcriptome profiling of adult animals should reveal mature and progenitor cells. By combining perturbation experiments, gene expression analysis, a computational method that predicts future cell states from transcriptional changes, and a lineage reconstruction method, we placed all major cell types onto a single lineage tree that connects all cells to a single stem cell compartment. We characterized gene expression changes during differentiation and discovered cell types important for regeneration. Our results demonstrate the importance of single-cell transcriptome analysis for mapping and reconstructing fundamental processes of developmental and regenerative biology at high resolution.