Advances in single-cell technology have enabled the measurement of cell-resolved molecular states across a variety of cell lines and tissues under a plethora of genetic, chemical, environmental or disease perturbations. Current methods focus on differential comparison or are specific to a particular task in a multi-condition setting with purely statistical perspectives. The quickly growing number, size and complexity of such studies require a scalable analysis framework that takes existing biological context into account. Here we present pertpy, a Python-based modular framework for the analysis of large-scale single-cell perturbation experiments. Pertpy provides access to harmonized perturbation datasets and metadata databases along with numerous fast and user-friendly implementations of both established and novel methods, such as automatic metadata annotation or perturbation distances, to efficiently analyze perturbation data. As part of the scverse ecosystem, pertpy interoperates with existing single-cell analysis libraries and is designed to be easily extended.
Single-cell multiomic analysis of the epigenome, transcriptome, and proteome allows for comprehensive characterization of the molecular circuitry that underpins cell identity and state. However, the holistic interpretation of such datasets presents a challenge given a paucity of approaches for systematic, joint evaluation of different modalities. Here, we present Panpipes, a set of computational workflows designed to automate multimodal single-cell and spatial transcriptomic analyses by incorporating widely-used Python-based tools to perform quality control, preprocessing, integration, clustering, and reference mapping at scale. Panpipes allows reliable and customizable analysis and evaluation of individual and integrated modalities, thereby empowering decision-making before downstream investigations.
We developed ehrapy, an open-source Python software framework for the exploratory analysis of electronic health record data. Ehrapy handles various widely used data formats, preprocessing tasks such as imputation of missing data and bias detection, and offers tools for analyses including patient stratification, survival analysis, causal inference and trajectory inference.
With progressive digitalization of healthcare systems worldwide, large-scale collection of electronic health records (EHRs) has become commonplace. However, an extensible framework for comprehensive exploratory analysis that accounts for data heterogeneity is missing. Here we introduce ehrapy, a modular open-source Python framework designed for exploratory analysis of heterogeneous epidemiology and EHR data. ehrapy incorporates a series of analytical steps, from data extraction and quality control to the generation of low-dimensional representations. Complemented by rich statistical modules, ehrapy facilitates associating patients with disease states, differential comparison between patient clusters, survival analysis, trajectory inference, causal inference and more. Leveraging ontologies, ehrapy further enables data sharing and training EHR deep learning models, paving the way for foundational models in biomedical research. We demonstrate ehrapy's features in six distinct examples. We applied ehrapy to stratify patients affected by unspecified pneumonia into finer-grained phenotypes. Furthermore, we reveal biomarkers for significant differences in survival among these groups. Additionally, we quantify medication-class effects of pneumonia medications on length of stay. We further leveraged ehrapy to analyze cardiovascular risks across different data modalities. We reconstructed disease state trajectories in patients with severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) based on imaging data. Finally, we conducted a case study to demonstrate how ehrapy can detect and mitigate biases in EHR data. ehrapy, thus, provides a framework that we envision will standardize analysis pipelines on EHR data and serve as a cornerstone for the community.
MOTIVATION:Pangenome graphs offer a comprehensive way of capturing genomic variability across multiple genomes. However, current construction methods often introduce biases, excluding complex sequences or relying on references. The PanGenome Graph Builder (PGGB) addresses these issues. To date, though, there is no state-of-the-art pipeline allowing for easy deployment, efficient and dynamic use of available resources, and scalable usage at the same time. RESULTS:To overcome these limitations, we present nf-core/pangenome, a reference-unbiased approach implemented in Nextflow following nf-core's best practices. Leveraging biocontainers ensures portability and seamless deployment in High-Performance Computing (HPC) environments. Unlike PGGB, nf-core/pangenome distributes alignments across cluster nodes, enabling scalability. Demonstrating its efficiency, we constructed pangenome graphs for 1000 human chromosome 19 haplotypes and 2146 Escherichia coli sequences, achieving a two to threefold speedup compared to PGGB without increasing greenhouse gas emissions. AVAILABILITY AND IMPLEMENTATION:nf-core/pangenome is released under the MIT open-source license, available on GitHub and Zenodo, with documentation accessible at https://nf-co.re/pangenome/docs/usage.
Single-cell multiplexing techniques (cell hashing and genetic multiplexing) combine multiple samples, optimizing sample processing and reducing costs. Cell hashing conjugates antibody-tags or chemical-oligonucleotides to cell membranes, while genetic multiplexing allows to mix genetically diverse samples and relies on aggregation of RNA reads at known genomic coordinates. We develop hadge (hashing deconvolution combined with genotype information), a Nextflow pipeline that combines 12 methods to perform both hashing- and genotype-based deconvolution. We propose a joint deconvolution strategy combining best-performing methods and demonstrate how this approach leads to the recovery of previously discarded cells in a nuclei hashing of fresh-frozen brain tissue.
Targeted spatial transcriptomic methods capture the topology of cell types and states in tissues at single-cell and subcellular resolution by measuring the expression of a predefined set of genes. The selection of an optimal set of probed genes is crucial for capturing the spatial signals present in a tissue. This requires selecting the most informative, yet minimal, set of genes to profile (gene set selection) for which it is possible to build probes (probe design). However, current selections often rely on marker genes, precluding them from detecting continuous spatial signals or new states. We present Spapros, an end-to-end probe set selection pipeline that optimizes both gene set specificity for cell type identification and within-cell type expression variation to resolve spatially distinct populations while considering prior knowledge as well as probe design and expression constraints. We evaluated Spapros and show that it outperforms other selection approaches in both cell type recovery and recovering expression variation beyond cell types. Furthermore, we used Spapros to design a single-cell resolution in situ hybridization on tissues (SCRINSHOT) experiment of adult lung tissue to demonstrate how probes selected with Spapros identify cell types of interest and detect spatial variation even within cell types.
Recent advances in single-cell technologies have enabled high-throughput molecular profiling of cells across modalities and locations. Single-cell transcriptomics data can now be complemented by chromatin accessibility, surface protein expression, adaptive immune receptor repertoire profiling and spatial information. The increasing availability of single-cell data across modalities has motivated the development of novel computational methods to help analysts derive biological insights. As the field grows, it becomes increasingly difficult to navigate the vast landscape of tools and analysis steps. Here, we summarize independent benchmarking studies of unimodal and multimodal single-cell analysis across modalities to suggest comprehensive best-practice workflows for the most common analysis steps. Where independent benchmarks are not available, we review and contrast popular methods. Our article serves as an entry point for novices in the field of single-cell (multi-)omic analysis and guides advanced users to the most recent best practices.
ABSTRACTOrgan- and body-scale cell atlases have the potential to transform our understanding of human biology. To capture the variability present in the population, these atlases must include diverse demographics such as age and ethnicity from both healthy and diseased individuals. The growth in both size and number of single-cell datasets, combined with recent advances in computational techniques, for the first time makes it possible to generate such comprehensive large-scale atlases through integration of multiple datasets. Here, we present the integrated Human Lung Cell Atlas (HLCA) combining 46 datasets of the human respiratory system into a single atlas spanning over 2.2 million cells from 444 individuals across health and disease. The HLCA contains a consensus re-annotation of published and newly generated datasets, resolving under- or misannotation of 59% of cells in the original datasets. The HLCA enables recovery of rare cell types, provides consensus marker genes for each cell type, and uncovers gene modules associated with demographic covariates and anatomical location within the respiratory system. To facilitate the use of the HLCA as a reference for single-cell lung research and allow rapid analysis of new data, we provide an interactive web portal to project datasets onto the HLCA. Finally, we demonstrate the value of the HLCA reference for interpreting disease-associated changes. Thus, the HLCA outlines a roadmap for the development and use of organ-scale cell atlases within the Human Cell Atlas.
Idiopathic pulmonary fibrosis (IPF) remains a clinical challenge with several unmet needs. Evidences support monocytes as biomarkers of IPF progression. Yet, specific myeloid subtypes and their role in disease are unknown. Using multi-color flow cytometry, we analyzed the abundance of circulating myeloid subsets. We confirmed increases in classical monocytes and immunosuppressive myeloid cells, so called myeloid-derived suppressor cells (MDSC), in blood of IPF patients. To address whether MDSC suppression contributes to fibrosis, we co-cultured autologous MDSC with T cells to assess proliferation, exhaustion and Treg formation, which showed decreased proliferation of CD8+ and CD4+ T cells in IPF. We developed an in vitro model to assess exhaustion. Autologous co-cultures induced CD8+ T cell exhaustion (PD1, Lag3, Tim3, TNFα, INFγ), and de-novo Treg formation. We applied single-cell transcriptomics in magnetically-purified monocyte populations to dissect the circulating heterogeneity. We detected two CD14+ monocyte states enriched in publicly available MDSC gene signatures, further supporting the presence of circulating myeloid populations with immunosuppressive features in IPF. To address whether MDSC migrate into the lung, we performed 3D gel invasion assays confirming their invasiveness potential in IPF when compared to controls. Analysis of IPF atlas confirmed the presence of exhausted CD8+ T cells in IPF tissue. Taken together, immature suppressive monocytic subsets are expanded in the peripheral blood of IPF patients and induce an immunosuppressive environment. Further studies will determine their role in disease progression and therapeutic targetability.
Pulmonary fibrosis develops as a consequence of failed regeneration after injury. Analyzing mechanisms of regeneration and fibrogenesis directly in human tissue has been hampered by the lack of organotypic models and analytical techniques. In this work, we coupled ex vivo cytokine and drug perturbations of human precision-cut lung slices (hPCLS) with single-cell RNA sequencing and induced a multilineage circuit of fibrogenic cell states in hPCLS. We showed that these cell states were highly similar to the in vivo cell circuit in a multicohort lung cell atlas from patients with pulmonary fibrosis. Using micro-CT-staged patient tissues, we characterized the appearance and interaction of myofibroblasts, an ectopic endothelial cell state, and basaloid epithelial cells in the thickened alveolar septum of early-stage lung fibrosis. Induction of these states in the hPCLS model provided evidence that the basaloid cell state was derived from alveolar type 2 cells, whereas the ectopic endothelial cell state emerged from capillary cell plasticity. Cell-cell communication routes in patients were largely conserved in hPCLS, and antifibrotic drug treatments showed highly cell type-specific effects. Our work provides an experimental framework for perturbational single-cell genomics directly in human lung tissue that enables analysis of tissue homeostasis, regeneration, and pathology. We further demonstrate that hPCLS offer an avenue for scalable, high-resolution drug testing to accelerate antifibrotic drug development and translation.
With progressive digitalization of healthcare systems worldwide, large-scale collection of electronic health records (EHRs) has become commonplace. However, an extensible framework for comprehensive exploratory analysis that accounts for data heterogeneity is missing. Here, we introduce ehrapy, a modular open-source Python framework designed for exploratory end-to-end analysis of heterogeneous epidemiology and electronic health record data. Ehrapy incorporates a series of analytical steps, from data extraction and quality control to the generation of low-dimensional representations. Complemented by rich statistical modules, ehrapy facilitates associating patients with disease states, differential comparison between patient clusters, survival analysis, trajectory inference, causal inference, and more. Leveraging ontologies, ehrapy further enables data sharing and training EHR deep learning models paving the way for foundational models in biomedical research. We demonstrated ehrapys features in five distinct examples: We first applied ehrapy to stratify patients affected by unspecified pneumonia into finer-grained phenotypes. Furthermore, we revealed biomarkers for significant differences in survival among these groups. Additionally, we quantify medication-class effects of pneumonia medications on length of stay. We further leveraged ehrapy to analyze cardiovascular risks across different data modalities. Finally, we reconstructed disease state trajectories in SARS-CoV-2 patients based on imaging data. Ehrapy thus provides a framework that we envision will standardize analysis pipelines on EHR data and serve as a cornerstone for the community. ### Competing Interest Statement LH is an employee of LaminLabs. FJT consults for Immunai Inc., Singularity Bio B.V., CytoReason Ltd, and Omniscope Ltd, and has ownership interest in Dermagnostix GmbH and Cellarity. ### Funding Statement This work was supported by the German Center for Lung Research (DZL), the Helmholtz association and the CRC/TRR 359 Perinatal Development of Immune Cell Topology (PILOT). N.H. and F.J.T. acknowledge support from the German Federal Ministry of Education and Research (BMBF) (LODE, 031L0210A). Co-funded by the European Union (ERC, DeepCell - 101054957). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study used only openly available human data. See: Physionet provides access to the PIC database at https://physionet.org/content/picdb/1.1.0 for credentialed users. The BrixIA images are available at https://github.com/BrixIA/Brixia-score-COVID-19. The diabetic retinopathy dataset is available at https://www.kaggle.com/c/diabetic-retinopathy-detection/data. The data used in this study were obtained from the UK Biobank (www.ukbiobank.ac.uk). Access to the UK Biobank resource was granted under application number 49966. The data are available to researchers upon application to the UK Biobank in accordance with their data access policies and procedures. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Physionet provides access to the PIC database at https://physionet.org/content/picdb/1.1.0 for credentialed users. The BrixIA images are available at https://github.com/BrixIA/Brixia-score-COVID-19. The diabetic retinopathy dataset is available at https://www.kaggle.com/c/diabetic-retinopathy-detection/data. The data used in this study were obtained from the UK Biobank (www.ukbiobank.ac.uk). Access to the UK Biobank resource was granted under application number 49966. The data are available to researchers upon application to the UK Biobank in accordance with their data access policies and procedures.
Abstract Mass spectrometry has become an indispensable tool in the life sciences. The new major version 3 of the computational framework OpenMS provides significant advancements regarding open, scalable, and reproducible high-throughput workflows for proteomics, metabolomics, and oligonucleotide mass spectrometry. OpenMS makes analyses from emerging fields available to experimentalists, enhances computational workflows, and provides a reworked Python interface to facilitate access for bioinformaticians and data scientists.
The pulmonary extracellular matrix (ECM) provides structural integrity and essential mechanical properties such as elastic recoil, but also serves as an information-rich signaling template that instructs cell identity and activity in health and disease. The spatial distribution and specificity of ECM proteins to the lungs distinct tissue niches is mostly unknown. In this study, we use spatially resolved mass spectrometry-based proteomics to define global differences in the proteomic composition of anatomically and histologically defined regions of the distal human lung. We characterized the distal airway tree and associated arteries and veins from 12 human donors*, using both 3D surgical microdissection (n=5) and 2D laser-capture microdissection techniques (n=7). Our analysis identified 7594 proteins (including 446 matrisome proteins) from 3D dissections and 4898 proteins (including 341 matrisome proteins) from 2D dissections. We identified gradients of ECM composition along the proximal-distal axis of the airways and vessels and discovered a specific ECM and immune niche in the respiratory bronchioles (RB). Integrative analysis of the spatial proteomes with donor-matched snRNAseq data and the Human Lung Cell Atlas (HLCA) further suggests the presence of a novel RB-associated fibroblast state that can be found across donors also in public datasets. As proof of concept, we identify and validate a subset of fibroblasts characterized by an ECM expression program enriched in RB via interactive imaging method. Given the potential origin of chronic lung diseases in distal airways, our discovery warrants further investigation of this novel RB specific ECM niche.
MOTIVATION:Machine learning has shown extensive growth in recent years and is now routinely applied to sensitive areas. To allow appropriate verification of predictive models before deployment, models must be deterministic. Solely fixing all random seeds is not sufficient for deterministic machine learning, as major machine learning libraries default to the usage of nondeterministic algorithms based on atomic operations.RESULTS:Various machine learning libraries released deterministic counterparts to the nondeterministic algorithms. We evaluated the effect of these algorithms on determinism and runtime. Based on these results, we formulated a set of requirements for deterministic machine learning and developed a new software solution, the mlf-core ecosystem, which aids machine learning projects to meet and keep these requirements. We applied mlf-core to develop deterministic models in various biomedical fields including a single-cell autoencoder with TensorFlow, a PyTorch-based U-Net model for liver-tumor segmentation in computed tomography scans, and a liver cancer classifier based on gene expression profiles with XGBoost.AVAILABILITY AND IMPLEMENTATION:The complete data together with the implementations of the mlf-core ecosystem and use case models are available at https://github.com/mlf-core.
nbproject is an open-source Python tool to help manage Jupyter notebooks with metadata, dependency, and integrity tracking. A draft-to-publish workflow creates more reproducible notebooks with context. There are a number of approaches to address reproducibility & manageability problems of computational R&D projects. nbproject complements - and should be combined with - approaches that are based on modularizing notebooks into pipelines, containerizing compute environments, or managing notebooks on centralized platforms.
IPF is a lethal chronic lung disease characterized by progressive alveolar fibrogenesis. There are currently no effective therapies that halt or reverse IPF progression underlining the need for better disease models to facilitate drug testing. Human precision-cut lung slices (hPCLS) treated with a pro-fibrotic cytokine mix (fibrotic cocktail - FC) are a promising new ex vivo model of human lung fibrogenesis (Alsafadi et al. AJP Lung, 2017). Here we investigate at single-cell resolution which aspects of in vivo IPF pathogenesis can be modelled in hPCLS in order to apply drug mode of action screens. We performed single-cell RNA-seq of hPCLS treated with FC, FC+Nintedanib (clinically approved drug) or FC+CMP4 (novel drug candidate). First, we compared our ex vivo data against in vivo data from an integrated IPF cell atlas. Second, drug mode of actions were dissected systematically using Differential Gene Expression and Gene Set Enrichtment analysis. Analysis of 23,000 cells revealed that FC induces characteristic cell state shifts that are key cellular hallmarks of IPF lungs (Adams, Schupp et al. Sci Adv, 2020) including the induction of myofibroblasts, aberrant basaloid and ectopic endothelial cells. We identify drug specific effects on these IPF associated cell states and describe drug effects on specific cell-cell communication routes within the fibrogenic niche of human lung parenchyma. In summary, our study validates hPCLS as a powerful ex vivo model that recapitulates aspects of IPF pathogenesis and demonstrates its potential for scalable, high-resolution drug testing.
Interferon gamma has critical antiviral properties by inducing type-1 immunity. In addition important immunoregulatory functions have been described that may go beyond antiviral immunity. We used single cell transcriptomics to comparatively analyze wild-type (n=30) and interferon-gamma receptor knockout (IFNγR−/−, n=30) mice infected with murine gamma herpesvirus 68 (MHV-68). Lytic virus was cleared from the lungs within 14 days but with incompetent and delayed response in IFNγR−/− mice. The IFNγR−/− mice showed a type-2 bias of immune response and developed irreversible lung fibrosis with many features of IPF within 100 days after virus infection. We harvested the lungs for scRNA-seq at days 3/6/15/28/45 and 100 post infection, and sequenced a total of 70.000 single cells. Single cell analysis provided an unprecedented resolution of MHV68 and host transcriptomes at single cell level with the identification of the infected cell types. We analyzed cell type specific responses to the virus and resolved gene expression kinetics for >30 individual cell types, revealing prominent type-1 and type-2 immune polarization patterns in WT and IFNγR−/− mice, respectively. Cell-cell communication analysis revealed distinct cellular circuits associated with the evolution of immunopathology in IFNγR−/− mice. For instance, the discovery of an immune recruiting state of the alveolar type-2 pneumocyte overexpressing Ccl3 and Ccl9 chemokines and the evolution of pro-fibrotic macrophage states. Overall, we present time-resolved single cell data on virus-host interactions in the context of type-2 interferon immune modulation that reveals important functions of IFNγ beyond antiviral immunity.