Spatially-resolved expression profiling data has revolutionized biological research with multiple emerging clinical applications. Spatial transcriptomic assays are often jointly measured with histopathology imaging data, which is frequently used for diagnosing and staging various diseases. However, determining the extent to which the spatial transcriptomic and histopathology data represent overlapping or unique sources of variation is challenging, particularly given the myriad of factors influencing both, including expression variation, spatial context, tissue morphology, and batch effects. Here, we view this challenge as multi-modal disentanglement and develop an evaluation framework. We introduce SpatialDIVA, a disentanglement technique for jointly measured spatially resolved transcriptomics and histopathology data. We demonstrate that SpatialDIVA outperforms baseline techniques in disentangling salient factors of variation in curated pathologist-annotated multi-sample colorectal and pancreatic cancer cohorts. Further, SpatialDIVA removes batch effects from multi-modal data, allows for factor covariance analysis, and yields actionable biological insights through a novel conditional multi-modal generation method. The SpatialDIVA model, evaluation code, and datasets are available at https://github.com/hsmaan/SpatialDIVA. ### Competing Interest Statement The authors have declared no competing interest.
Generative pretrained models have achieved remarkable success in various domains such as language and computer vision. Specifically, the combination of large-scale diverse datasets and pretrained transformers has emerged as a promising approach for developing foundation models. Drawing parallels between language and cellular biology (in which texts comprise words; similarly, cells are defined by genes), our study probes the applicability of foundation models to advance cellular biology and genetic research. Using burgeoning single-cell sequencing data, we have constructed a foundation model for single-cell biology, scGPT, based on a generative pretrained transformer across a repository of over 33 million cells. Our findings illustrate that scGPT effectively distills critical biological insights concerning genes and cells. Through further adaptation of transfer learning, scGPT can be optimized to achieve superior performance across diverse downstream applications. This includes tasks such as cell type annotation, multi-batch integration, multi-omic integration, perturbation response prediction and gene network inference. Pretrained using over 33 million single-cell RNA-sequencing profiles, scGPT is a foundation model facilitating a broad spectrum of downstream single-cell analysis tasks by transfer learning.
The Iniquitate pipeline assessed the impacts of cell-type imbalance on single-cell RNA sequencing integration through perturbations to dataset balance. The results indicated that cell-type imbalance not only leads to loss of biological signal in the integrated space, but also can change the interpretation of downstream analyses after integration.
Existing RNA velocity estimation methods strongly rely on predefined dynamics and cell-agnostic constant transcriptional kinetic rates, assumptions often violated in complex and heterogeneous single-cell RNA sequencing (scRNA-seq) data. Using a graph convolution network, DeepVelo overcomes these limitations by generalizing RNA velocity to cell populations containing time-dependent kinetics and multiple lineages. DeepVelo infers time-varying cellular rates of transcription, splicing, and degradation, recovers each cell’s stage in the differentiation process, and detects functionally relevant driver genes regulating these processes. Application to various developmental and pathogenic processes demonstrates DeepVelo’s capacity to study complex differentiation and lineage decision events in heterogeneous scRNA-seq data.
Computational methods for integrating single-cell transcriptomic data from multiple samples and conditions do not generally account for imbalances in the cell types measured in different datasets. In this study, we examined how differences in the cell types present, the number of cells per cell type and the cell type proportions across samples affect downstream analyses after integration. The Iniquitate pipeline assesses the robustness of integration results after perturbing the degree of imbalance between datasets. Benchmarking of five state-of-the-art single-cell RNA sequencing integration techniques in 2,600 integration experiments indicates that sample imbalance has substantial impacts on downstream analyses and the biological interpretation of integration results. Imbalance perturbation led to statistically significant variation in unsupervised clustering, cell type classification, differential expression and marker gene annotation, query-to-reference mapping and trajectory inference. We quantified the impacts of imbalance through newly introduced properties—aggregate cell type support and minimum cell type center distance. To better characterize and mitigate impacts of imbalance, we introduce balanced clustering metrics and imbalanced integration guidelines for integration method users.
The incorporation of sequencing technologies in frontline and public health healthcare settings was vital in developing virus surveillance programs during the Coronavirus Disease 2019 (COVID-19) pandemic caused by transmission of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). However, increased data acquisition poses challenges for both rapid and accurate analyses. To overcome these hurdles, we developed the SARS-CoV-2 Illumina GeNome Assembly Line (SIGNAL) for quick bulk analyses of Illumina short-read sequencing data. SIGNAL is a Snakemake workflow that seamlessly manages parallel tasks to process large volumes of sequencing data. A series of outputs are generated, including consensus genomes, variant calls, lineage assessments and identified variants of concern (VOCs). Compared to other existing SARS-CoV-2 sequencing workflows, SIGNAL is one of the fastest-performing analysis tools while maintaining high accuracy. The source code is publicly available (github.com/jaleezyy/covid-19-signal) and is optimized to run on various systems, with software compatibility and resource management all handled within the workflow. Overall, SIGNAL illustrated its capacity for high-volume analyses through several contributions to publicly funded government public health surveillance programs and can be a valuable tool for continuing SARS-CoV-2 Illumina sequencing efforts and will inform the development of similar strategies for rapid viral sequence assessment.
A bstract Single-cell sequencing has emerged as a promising technique to decode cellular heterogeneity and analyze gene functions. With the high throughput of modern techniques and resulting large-scale sequencing data, deep learning has been used extensively to learn representations of individual cells for downstream tasks. However, most existing methods rely on fully connected networks and are unable to model complex relationships between both cell and gene representations. We hereby propose scFormer, a novel transformer-based deep learning framework to jointly optimize cell and gene embeddings for single-cell biology in an unsupervised manner. By drawing parallels between natural language processing and genomics, scFormer applies self-attention to learn salient gene and cell embeddings through masked gene modelling. scFormer provides a unified framework to readily address a variety of downstream tasks such as data integration, analysis of gene function, and perturbation response prediction. Extensive experiments using scFormer show state-of-the-art performance on seven datasets across the relevant tasks. The scFormer model implementation is available at https://github.com/bowang-lab/scFormer .
Macrophage colony stimulating factor-1 (CSF-1) plays a critical role in maintaining myeloid lineage cells. However, congenital global deficiency of CSF-1 (Csf1op/op) causes severe musculoskeletal defects that may indirectly affect hematopoiesis. Indeed, we show here that osteolineage-derived Csf1 prevented developmental abnormalities but had no effect on monopoiesis in adulthood. However, ubiquitous deletion of Csf1 conditionally in adulthood decreased monocyte survival, differentiation, and migration, independent of its effects on bone development. Bone histology revealed that monocytes reside near sinusoidal endothelial cells (ECs) and leptin receptor (Lepr)-expressing perivascular mesenchymal stromal cells (MSCs). Targeted deletion of Csf1 from sinusoidal ECs selectively reduced Ly6C- monocytes, whereas combined depletion of Csf1 from ECs and MSCs further decreased Ly6Chi cells. Moreover, EC-derived CSF-1 facilitated recovery of Ly6C- monocytes and protected mice from weight loss following induction of polymicrobial sepsis. Thus, monocytes are supported by distinct cellular sources of CSF-1 within a perivascular BM niche.
1 Abstract The introduction of RNA velocity in single-cell studies has opened new ways of examining cell differentiation and tissue development. Existing RNA velocity estimation methods rely on strong assumptions of predefined dynamics and cell-agnostic constant transcriptional kinetic rates, which are often violated in complex and heterogeneous single-cell RNA sequencing (scRNA-seq) data. To overcome these limitations, we propose DeepVelo, a novel method that estimates the cell-specific dynamics of splicing kinetics using Graph Convolution Networks (GCNs). DeepVelo generalizes RNA velocity to cell populations containing time-dependent kinetics and multiple lineages, which are common in developmental and pathological systems. We applied DeepVelo to disentangle multifaceted kinetics in the processes of dentate gyrus neurogenesis, pancreatic endocrinogenesis, and hindbrain development. The method infers time-varying cellular rates of transcription, splicing and degradation, recovers each cell’s stage in the underlying differentiation process, and detects functionally relevant driver genes regulating these processes. DeepVelo relaxes the constraints of previous techniques, facilitates the study of more complex differentiation and lineage decision events in heterogeneous scRNA-seq data, and is more computationally efficient than previous techniques.
The COVID-19 pandemic has highlighted the urgent need for the identification of new antiviral drug therapies for a variety of diseases. COVID-19 is caused by infection with the human coronavirus SARS-CoV-2, while other related human coronaviruses cause diseases ranging from severe respiratory infections to the common cold. We developed a computational approach to identify new antiviral drug targets and repurpose clinically-relevant drug compounds for the treatment of a range of human coronavirus diseases. Our approach is based on graph convolutional networks (GCN) and involves multiscale host-virus interactome analysis coupled to off-target drug predictions. Cell-based experimental assessment reveals several clinically-relevant drug repurposing candidates predicted by the in silico analyses to have antiviral activity against human coronavirus infection. In particular, we identify the MET inhibitor capmatinib as having potent and broad antiviral activity against several coronaviruses in a MET-independent manner, as well as novel roles for host cell proteins such as IRAK1/4 in supporting human coronavirus infection, which can inform further drug discovery studies.
Type I interferons (IFNs) are our first line of defense against virus infection. Recent studies have suggested the ability of SARS-CoV-2 proteins to inhibit IFN responses. Emerging data also suggest that timing and extent of IFN production is associated with manifestation of COVID-19 severity. In spite of progress in understanding how SARS-CoV-2 activates antiviral responses, mechanistic studies into wild-type SARS-CoV-2-mediated induction and inhibition of human type I IFN responses are scarce. Here we demonstrate that SARS-CoV-2 infection induces a type I IFN response in vitro and in moderate cases of COVID-19. In vitro stimulation of type I IFN expression and signaling in human airway epithelial cells is associated with activation of canonical transcriptions factors, and SARS-CoV-2 is unable to inhibit exogenous induction of these responses. Furthermore, we show that physiological levels of IFNα detected in patients with moderate COVID-19 is sufficient to suppress SARS-CoV-2 replication in human airway cells.
Genome sequencing of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is increasingly important to monitor the transmission and adaptive evolution of the virus. The accessibility of high-throughput methods and polymerase chain reaction (PCR) has facilitated a growing ecosystem of protocols. Two differing protocols are tiling multiplex PCR and bait capture enrichment. Each method has advantages and disadvantages but a direct comparison with different viral RNA concentrations has not been performed to assess the performance of these approaches. Here we compare Liverpool amplification, ARTIC amplification, and bait capture using clinical diagnostics samples. All libraries were sequenced using an Illumina MiniSeq with data analyzed using a standardized bioinformatics workflow (SARS-CoV-2 Illumina GeNome Assembly Line; SIGNAL). One sample showed poor SARS-CoV-2 genome coverage and consensus, reflective of low viral RNA concentration. In contrast, the second sample had a higher viral RNA concentration, which yielded good genome coverage and consensus. ARTIC amplification showed the highest depth of coverage results for both samples, suggesting this protocol is effective for low concentrations. Liverpool amplification provided a more even read coverage of the SARS-CoV-2 genome, but at a lower depth of coverage. Bait capture enrichment of SARS-CoV-2 cDNA provided results on par with amplification. While only two clinical samples were examined in this comparative analysis, both the Liverpool and ARTIC amplification methods showed differing efficacy for high and low concentration samples. In addition, amplification-free bait capture enriched sequencing of cDNA is a viable method for generating a SARS-CoV-2 genome sequence and for identification of amplification artifacts.
The ongoing COVID-19 pandemic is the greatest health-care challenge of this generation. Early viral genome sequencing studies of small cohorts have indicated the possibility of distinct severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) genotypes.1Tang X Wu C Li X et al.On the origin and continuing evolution of SARS-CoV-2.Natl Sci Rev. 2020; (published online March 3.)DOI:10.1093/nsr/nwaa036Crossref PubMed Scopus (1059) Google Scholar If these subtypes result in an altered virus tropism or pathogenesis in infected hosts, this could have immediate implications for vaccine design, drug development, and efforts to control the pandemic. Therefore, the genomic surveillance and characterisation of circulating viral strains is a high priority for research and development. To facilitate the epidemiological tracking of SARS-CoV-2, researchers worldwide have created various web-portals and tools, such as the Johns Hopkins University COVID-19 dashboard.2Dong E Du H Gardner L An interactive web-based dashboard to track COVID-19 in real time.Lancet Infect Dis. 2020; 20: 533-534Summary Full Text Full Text PDF PubMed Scopus (6279) Google Scholar An unprecedented effort to make COVID-19-related data accessible in near real-time has resulted in more than 25 000 publicly available genome sequences of SARS-CoV-2 on Global Initiative on Sharing All Influenza Data (GISAID).3Elbe S Buckland-Merrett G Data, disease and diplomacy: GISAID's innovative contribution to global health.Glob Chall. 2017; 1: 33-46Crossref PubMed Scopus (1055) Google Scholar Although platforms to survey epidemiological data are prevalent, tools that summarise publicly available viral genome data are scarce and those that are available do not offer users the ability to analyse in-house sequencing data. To address this gap, we have developed an accessible application, the COVID-19 Genotyping Tool (CGT). A video demonstration of CGT is available in appendix 1. CGT uses publicly deposited SARS-CoV-2 consensus genome sequences from the GISAID EpiCoV™ database,3Elbe S Buckland-Merrett G Data, disease and diplomacy: GISAID's innovative contribution to global health.Glob Chall. 2017; 1: 33-46Crossref PubMed Scopus (1055) Google Scholar and summarises relevant information through salient visualisations, including Uniform Manifold Approximation and Projection (UMAP), Minimum Spanning Trees (MST) of sequence networks, and allele frequencies of annotated high-prevalence non-synonymous Single-Nucleotide Polymorphisms (SNPs) within structural protein-coding genomic regions (envelope, membrane, nucleocapsid, and spike proteins; figure). New sequencing data from GISAID are added to CGT once a week. Currently, three metadata types can be overlaid on the visualisations: the region of sample collection, the country of sample collection, and the sample collection date with respect to the start of the pandemic (heuristically defined as Dec 1, 2019; figure; appendix 2). Users can upload post-assembly consensus FASTA sequences of SARS-CoV-2 for interactive analysis. The 3' and 5' untranslated regions of genomes are trimmed because of low sequence identity, caused by difficulties in their amplification and sequencing. After sequence alignment, the DNA distance is calculated using the Kimura-80 model of nucleotide substitution.4Kimura M A simple method for estimating evolutionary rates of base substitutions through comparative studies of nucleotide sequences.J Mol Evol. 1980; 16: 111-120Crossref PubMed Scopus (24466) Google Scholar After the calculation of the distance and annotation of SNPs, visualisations are reactively reprocessed with user-uploaded data. We use UMAP and network analysis instead of phylogenetics because of a faster computation time and ease of interpretability for large datasets. UMAP reduces the high-dimensional representation of DNA sequences to a two-dimensional embedding, indicating a sequence similarity based on the distance between points;5McInnes L Healy J Melville J UMAP: Uniform Manifold Approximation and Projection for dimension reduction.arXiv. 2018; (published online Feb 9.) (preprint).https://arxiv.org/abs/1802.03426Google Scholar that is, SARS-CoV-2 genomes. The MST of the SARS-CoV-2 genome network is a metric used to define a network subset such that all the nodes are connected while minimising distance.6Mamun A, Rajasekaran S. An efficient Minimum Spanning Tree algorithm. 2016 IEEE Symposium on Computers and Communication; Messina, Italy; June 27–30, 2016 (abstr 1047–52).Google Scholar MSTs have been used before in outbreak analysis to identify the most probable transmission events between hosts,7Spada E Sagliocca L Sourdis J et al.Use of the Minimum Spanning Tree model for molecular epidemiological investigation of a nosocomial outbreak of hepatitis C virus infection.J Clin Microbiol. 2004; 42: 4230-4236Crossref PubMed Scopus (39) Google Scholar and can therefore offer epidemiological insight into SARS-CoV-2 transmission. Lastly, heterogeneity within structural proteins is particularly notable in terms of the host immune response, and has direct implications for vaccine development. With this in mind, we present the most prevalent (based on minor allele frequency) non-synonymous SNPs in the genomic regions of structural proteins. Our results indicate that there are distinct viral isolate clusters for SARS-CoV-2 sequences uploaded to GISAID. Larger outbreak clusters and hubs from UMAP and MST probably reflect outbreak epicentres (figure). Smaller clusters might be indicative of isolated outbreaks, and singleton samples might be indicative of isolated cases (figure). The analysis of SNPs reveals variants with notable minor allele frequencies involving missense substitutions in structural protein-coding genome regions (figure). Our novel tool is accessible to those who might not be trained in bioinformatics and epidemiological analysis, and thus it serves as a platform for aiding in a pivotal aspect of the global research effort against the COVID-19 pandemic. The current limitations of CGT include the sequence input limit and processing time, both of which are because of the size of the public data that must be concurrently processed with user-input sequences (>25 000 GISAID sequences). Our team is working to continuously integrate optimisations to the CGT data processing pipeline through code parallelisation and the refinement of deployment infrastructure. Future releases of the application will aim to decrease the processing time, increase the user-input sequence limit, incorporate the input and processing of raw SARS-CoV-2 sequencing data (eg, FASTQ files), and add additional information related to sequence epidemiology, such as travel history. Complete documentation of the CGT analysis pipeline and application source-code are available in a GitHub repository. Our application does not store any user-uploaded sequence data on the server-side or client-side; CGT simply processes the data to create updated visualisations, and information does not persist after the user disconnects. SARS-CoV-2 genome sequence data and linked metadata from GISAID are not published on our website, as per the GISAID data usage policy. Up-to-date acknowledgments for the usage of GISAID uploaded sequences for analysis are available in the GitHub repository. Our team is grateful to all the researchers who have shared SARS-CoV-2 viral genome sequencing data on GISAID. We declare no competing interests. HMa, BW, and AGM conceptualised and designed the study. HMa and HMb collected and analysed the data. HMa, HMb, BW, AB, JAN, ARR, and AGM interpreted the initial results. HMa developed the application. HMa and ARR did the software testing. HMa and NK created the application documentation. HMa, HMb, ARR, AB, JAN, RAK, NK, SM, AGM, and BW tested the application. RAK, NK, SM, and AGM provided suggestions for application visualisations. HMa and HMb created the manuscript figures. HMa and BW wrote the manuscript. HMa, BW, HMb, AB, NK, SM, and AGM edited the manuscript. eyJraWQiOiI4ZjUxYWNhY2IzYjhiNjNlNzFlYmIzYWFmYTU5NmZmYyIsImFsZyI6IlJTMjU2In0.eyJzdWIiOiJiNzZmZjEwZGVjYTlmNDkwYThmMzc1OTJiMzcwZWI4NCIsImtpZCI6IjhmNTFhY2FjYjNiOGI2M2U3MWViYjNhYWZhNTk2ZmZjIiwiZXhwIjoxNzAyMzgwMzczfQ.g2yjYNk7xVxCyBbm4vi8Aa-0h1-odsfYeHUd0iQQ8gnXPmhtlWvlHRLiY2uXC2fduG_XljWdVEN2XnbRGfXcAJxR8_45TAz2DwvjLWmQhtI7CS6Tapk7LPFaH4UI2EJsplugk6HSE1m9X_ez7hGNkQnBvVofeq7IaE3yZ0y_deIqIovCZtDulhoxi4UU11fDZXLkd0MyaibigCcGNZPFgG-Y6AyPRGqhkYQ0FwIgKoiuciUG89rJXdhquM-S7nipapeTGwhaJFdDjey5DDhU84kWbXNS83KoNiN7P-VWS2SAJu4hUx074AFB7djpl_m3JFzBhmBpJr27wdl831uYWw Download .mp4 (84.41 MB) Help with .mp4 files Supplementary VideoThe COVID-19 Genotyping Tool tutorialA short demonstration on using the COVID-19 Genotyping Tool application, understanding its features, and navigating the website Download .pdf (.55 MB) Help with pdf files Supplementary appendix 2
The bone marrow niche factors that sustain monocytes – phagocytes with complex functions in the circulation and peripheral tissues – and the functionally important cells that produce them, are poorly defined. Here, we conditionally deleted macrophage colony stimulating factor 1 (Csf1) in adult mice and show that CSF1 regulated differentiation and survival of monocytes and their precursor cells independent of its effects on bone development. Specifically, CSF1 produced by sinusoidal endothelial cells (ECs) but not leukocytes or stromal cells selectively maintained abundance of nonclassical Ly6Cnegative monocytes. Ly6Chigh monocytes, on the other hand, required CSF1 produced by ECs as well as leptin receptor (Lepr)-expressing perivascular stromal cells. Importantly, osteoblast-derived CSF1 critically supported bone formation and hematopoiesis during early development but did not contribute to maintenance of monocytes in adulthood. Our findings reveal classical and non-classical monocytes receive support by distinct cellular sources of CSF1 within a perivascular bone marrow niche.
Objective: To determine the interobserver variability of the 2015 American Thyroid Association (ATA) thyroid guidelines and to evaluate the diagnostic accuracy of the guidelines in detecting thyroid cancer. Materials and methods: Sonographic patterns of 189 thyroid lesions were retrospectively analyzed by two radiologists according to the 2015 guidelines. The risk of malignancy was calculated for each pattern and compared with the published expected risk of malignancy. Results: The observed risk of malignancy for very low suspicion, low suspicion, intermediate suspicion and high suspicion patterns were 2%, 12.7%, 26.3% and 29.8% respectively. Interobserver agreement for final category assignment was moderate (kappa 0.518). Conclusion: The estimated risk of malignancy in the high suspicion pattern of the 2015 ATA thyroid biopsy guidelines appears to be less than stated. However, this needs further validation in a larger cohort study.
Unfortunately the article was published with a spell error in the co-author name "Hassan Maan". The correct co-author name should be "Hassaan Maan".
Two highly pathogenic human coronaviruses that cause severe acute respiratory syndrome (SARS) and Middle East respiratory syndrome (MERS) have evolved proteins that can inhibit host antiviral responses, likely contributing to disease progression and high case-fatality rates. SARS-CoV-2 emerged in December 2019 resulting in a global pandemic. Recent studies have shown that SARS-CoV-2 is unable to induce a robust type I interferon (IFN) response in human cells, leading to speculation about the ability of SARS-CoV-2 to inhibit innate antiviral responses. However, innate antiviral responses are dynamic in nature and gene expression levels rapidly change within minutes to hours. In this study, we have performed a time series RNA-seq and selective immunoblot analysis of SARS-CoV-2 infected lung (Calu-3) cells to characterize early virus-host processes. SARS-CoV-2 infection upregulated transcripts for type I IFNs and interferon stimulated genes (ISGs) after 12 hours. Furthermore, we analyzed the ability of SARS-CoV-2 to inhibit type I IFN production and downstream antiviral signaling in human cells. Using exogenous stimuli, we discovered that SARS-CoV-2 is unable to modulate IFNβ production and downstream expression of ISGs, such as IRF7 and IFIT1. Thus, data from our study indicate that SARS-CoV-2 may have evolved additional mechanisms, such as masking of viral nucleic acid sensing by host cells to mount a dampened innate antiviral response. Further studies are required to fully identify the range of immune-modulatory strategies of SARS-CoV-2.