Alternative polyadenylation (APA) is a key regulator of gene expression and cellular dynamics, yet systematic investigations of spatially resolved APA across diverse human tissues remain limited. Here, we developed SpatialAPA (https://github.com/Omicslab-Zhang/spatialAPA), a framework that benchmarks multiple APA identification methods and integrates spatial APA data with gene expression and cellular dynamics at spatial resolution. Applying SpatialAPA to 363 spatial transcriptomic datasets from 56 projects across 18 human tissues and 76 diseases, we constructed a spatially resolved APA atlas comprising 346,932 APA events across 52,175 genes. This atlas reveals organ-specific APA patterns and provides new insights into how APA regulates tissue homeostasis and disease progression beyond transcriptional control. To ensure cross-sample comparability, we applied batch correction, and spatial cell deconvolution was performed to uncover cell-type-specific dynamics and interactions. In triple-negative breast cancer, integrated spatial and single-cell analyses identified TSPAN8-positive epithelial subpopulations whose distinct APA regulation and transcriptional programs drive differentiation and malignant progression. To facilitate community access, we developed an online platform (http://www.biomedical-web.com/spatialAPAdb/home) for exploring APA, gene expression, and cellular dynamics in health and disease. Together, this study establishes the first comprehensive spatial APA atlas, providing a valuable resource and analytical framework for investigating molecular mechanisms and therapeutic targets.
Objectives.Reconstruction of visual perception from brain signals has emerged as a promising research topic. Electrocorticography (ECoG) is a kind of high-quality intracranial signal with good spatiotemporal resolution that offers some new opportunities. However, according to our knowledge, there are no studies to reconstruct the perceived images from human ECoG signals at present.Approach.We have conducted the pioneering work and developed a novel pipeline that integrates Talairach coordinate alignment masked autoencoders (TA-MAE) with denoising diffusion probabilistic models. Our approach exploits the spatiotemporal dynamics of human ECoG signals, enabling the restoration of details in high-resolution.Main results.Experiments show that our method outperforms the current state-of-the-art methods in terms of appearance, structure, signal-noise ratio, and semantic consistency. Additionally, our study indicated that unsupervised learning-based signal reconstruction outperforms manually annotated label-guided feature recognition in capturing the low-dimensional representation of brain signals, potentially facilitating the exploration of vision's intrinsic mechanisms.Significance.These results highlight the advantages of unsupervised decoding and provide a generalizable framework for human ECoG-based visual reconstruction.
BACKGROUND:The human lung is a highly complex organ characterised by extensive cellular heterogeneity, making it susceptible to a broad range of diseases. Single-cell transcriptomics has shed light on disease-specific cellular features, but previous studies have been fragmented, limiting a unified understanding of cellular mechanisms across various lung diseases. Furthermore, high-resolution reference atlases for lung cells are lacking, impeding the effective integration of spatial omics data and the exploration of shared pathogenic mechanisms. METHODS:We constructed uniLUNG, the most comprehensive single-cell RNA sequencing atlas of the human lung, by integrating 62 published datasets comprising 9.2 million cells from 1807 donors across health and 17 disease conditions. We leveraged this high-resolution cell atlas to seek distinct cell types across diverse lung pathologies. By integrating spatial transcriptomics, we also identified transitional cell populations in lung cancer, and provided new insights into tumour evolution and the associated microenvironment. FINDINGS:We present a comprehensive lung cell atlas, encompassing cellular data across major lung diseases and health states. Using this resource, we identified distinct cell populations, such as Lym-monocytes and T-like B cells, which are specifically enriched in certain lung diseases and linked to immune dysregulation. Furthermore, our spatially resolved multi-omics analysis revealed a transitional malignant subpopulation, NSCLC-like SCLC, which plays a key role in the transformation from non-small cell lung cancer (NSCLC) to small cell lung cancer (SCLC), driving tumour microenvironment remodelling. INTERPRETATION:We offered a high-resolution, cross-disease lung cell reference that uncovers distinct cell types and cellular transitions critical to disease progression and therapeutic resistance. This resource provides essential insights into lung disease mechanisms and has important implications for the development of targeted therapeutic strategies, particularly in the context of lung cancer. FUNDING:This work was supported by the grants of National Key R&D Program of China (2021YFF1200900 and 2021YFF1200903), National Natural Science Foundation of China (92474107), Guangdong Basic and Applied Basic Research Foundation of China (2022B1515120077), Major Project of Guangzhou National Laboratory of China (GZNL2024A01003), and Support Scheme of Guangzhou for Leading Talents in Innovation and Entrepreneurship of China (2020007).
Recent advances in single-cell multi-omics co-assays and spatiotemporal sequencing technologies have provided unprecedented opportunities for systematically characterizing cellular heterogeneity. However, severe sparsity and pronounced spatial heterogeneity—hallmark features of complex diseases and tumor microenvironment—remain major obstacles in deciphering cellular multi-omics data. Here, we present scPhoenix, a contrastive learning–based framework for single-cell cross-modality translation. scPhoenix adopts a two-stage Aux–Core strategy to disentangle modality-specific feature extraction from cross-modality feature interaction. Across diverse datasets, it preserves cellular heterogeneity during translation and demonstrates significant advantages for data with high sparsity. In addition, the framework integrates a contrastive learning framework with five effective data augmentation methods tailored to single-cell data. Moreover, scPhoenix’s design supports extensions to unpaired data training and spatial multi-omics translation, enabling robust performance in scenarios with high spatial heterogeneity. scPhoenix is freely available at . ### Competing Interest Statement To maximize the impact of this study, Sun Yat-sen University has submitted a patent application to the State Intellectual Property Office of Chia (SIPO). National Natural Science Foundation of China, 92474107 National Key R&D Program of China, 2021YFF1200903 Guangdong Basic and Applied Basic Research Foundation of China, 2022B1515120077 Major Project of Guangzhou National Laboratory of China, GZNL2024A01003
Recent advances in spatial multi-omics technologies provide unprecedented opportunities to integrate and interpret molecular features within the tissue microenvironment. Here we present SpatialFuser, the first unified deep learning framework for detailed molecular profiling within individual tissue sections and multiomics integrative analysis of cross-modality spatial data. SpatialFuser offers comprehensive tools and models for accurate spatial interpretation, robust cross-modality integration, and effective cross-slice alignment of spatial epigenomics, transcriptomics, proteomics, and metabolomics. Benchmarking results demonstrate SpatialFuser’s superior performance and reliability in spatial domain detection and consecutive slice alignment task compared to existing state-of-the-art methods. Applications to diverse datasets spanning various resolution and omics types further highlight SpatialFuser’s ability to capture precise molecular patterns, reveal developmental dynamics, and uncover fine-grained biological variation from complementary perspectives, offering a holistic view of cellular and tissue properties. The SpatialFuser framework is open-source and available at . ### Competing Interest Statement To maximize the impact of this study, Sun Yat-sen University has submitted a patent applications to the State Intellectual Property Office of Chia (SIPO). National Natural Science Foundation of China, 92474107 National Key R&D Program of China, 2021YFF1200903 Guangdong Basic and Applied Basic Research Foundation of China, 2022B1515120077 Major Project of Guangzhou National Laboratory of China, GZNL2024A01003
Human lung is a complex organ susceptible to various diseases. Single-cell transcriptomic studies provide rich data to targeting specific research questions. Here, we present uniLUNG, the largest lung transcriptomic cell atlas, comprising over 10 million cells across 20 disease states and healthy controls. We ensembled a universal hierarchical annotation framework and conducted a full benchmarking of data integration to define a standardized nomenclature and marker genes for lung cell types. Using uniLUNG, we identified Lym-monocyte and T-like B cells, new cell types in specific lung diseases, confirming their existence by comparing with external single-cell atlases. Additionally, we discovered the NSCLC-like SCLC subpopulation, a transitional malignant cell population associated with the transition from NSCLC to SCLC, which was validated and further characterized in spatial dimensions, revealing its complex role in tumour progression. Overall, uniLUNG represents a comprehensive range of human lung cell diversity, providing valuable data resources and a reliable foundation for lung single-cell research. ### Competing Interest Statement The authors have declared no competing interest.
The selection of embryos is a key for the success of in vitro fertilization (IVF). However, automatic quality assessment on human IVF embryos with optical microscope images is still challenging. In this study, we developed a clinical consensus-compliant deep learning approach, named Esava (Embryo Segmentation and Viability Assessment), to quantitatively evaluate the development of IVF embryos using optical microscope images. In total 551 optical microscope images of human IVF embryos of day-2 to day-3 were collected, preprocessed, and annotated. Using the Faster R-CNN model as baseline, our Esava model was constructed, refined, trained, and validated for precise and robust blastomere detection. A novel algorithm Crowd-NMS was proposed and employed in Esava to enhance the object detection and to precisely quantify the embryonic cells and their size uniformity. Additionally, an innovative GrabCut-based unsupervised module was integrated for the segmentation of blastomeres and embryos. Independently tested on 94 embryo images for blastomere detection, Esava obtained the high rates of 0.9940, 0.9121, and 0.9531 for precision, recall, and mAP respectively, and gained significant advances compared with previous computational methods. Intraclass correlation coefficients indicated the consistency between Esava and three experienced embryologists. Another test on 51 extra images demonstrated that Esava surpassed other tools significantly, achieving the highest average precision 0.9025. Moreover, it also accurately identified the borders of blastomeres with mIoU over 0.88 on the independent testing dataset. Esava is compliant with the Istanbul clinical consensus and compatible to senior embryologists. Taken together, Esava improves the accuracy and efficiency of embryonic development assessment with optical microscope images.
Untargeted metabolomic analysis using mass spectrometry provides comprehensive metabolic profiling, but its medical application faces challenges of complex data processing, high inter-batch variability, and unidentified metabolites. Here, we present DeepMSProfiler, an explainable deep-learning-based method, enabling end-to-end analysis on raw metabolic signals with output of high accuracy and reliability. Using cross-hospital 859 human serum samples from lung adenocarcinoma, benign lung nodules, and healthy individuals, DeepMSProfiler successfully differentiates the metabolomic profiles of different groups (AUC 0.99) and detects early-stage lung adenocarcinoma (accuracy 0.961). Model flow and ablation experiments demonstrate that DeepMSProfiler overcomes inter-hospital variability and effects of unknown metabolites signals. Our ensemble strategy removes background-category phenomena in multi-classification deep-learning models, and the novel interpretability enables direct access to disease-related metabolite-protein networks. Further applying to lipid metabolomic data unveils correlations of important metabolites and proteins. Overall, DeepMSProfiler offers a straightforward and reliable method for disease diagnosis and mechanism discovery, enhancing its broad applicability.
The human respiratory system is a complex and important system that can suffer a variety of diseases. Single-cell sequencing technologies, applied in many respiratory disease studies, have enhanced our ability in characterizing molecular and phenotypic features at a single-cell resolution. The exponentially increasing data from these studies have consequently led to difficulties in data sharing and analysis. Here, we present scMoresDB, a single-cell multi-omics database platform with extensive omics types tailored for human respiratory diseases. scMoresDB re-analyzes single-cell multi-omics datasets, providing a user-friendly interface with cross-omics search capabilities, interactive visualizations, and analytical tools for comprehensive data sharing and integrative analysis. Our example applications highlight the potential significance of BSG receptor in SARS-CoV-2 infection as well as the involvement of HHIP and TGFB2 in the development and progression of chronic obstructive pulmonary disease. scMoresDB significantly increases accessibility and utility of single-cell data relevant to human respiratory system and associated diseases.
Editor's noteA commentary on “Multi-omics single-cell data integration and regulatory inference with graph-linked embedding”
In the post COVID-19 era, new SARS-CoV-2 variant strains may continue emerging and long COVID is poised to be another public health challenge. Deciphering the molecular susceptibility of receptors to SARS-CoV-2 spike protein is critical for understanding the immune responses in COVID-19 and the rationale of multi-organ injuries. Currently, such systematic exploration remains limited. Here, we conduct multi-omic analysis of protein binding affinities, transcriptomic expressions, and single-cell atlases to characterize the molecular susceptibility of receptors to SARS-CoV-2 spike protein. Initial affinity analysis explains the domination of delta and omicron variants and demonstrates the strongest affinities between BSG (CD147) receptor and most variants. Further transcriptomic data analysis on 4100 experimental samples and single-cell atlases of 1.4 million cells suggest the potential involvement of BSG in multi-organ injuries and long COVID, and explain the high prevalence of COVID-19 in elders as well as the different risks for patients with underlying diseases. Correlation analysis validated moderate associations between BSG and viral RNA abundance in multiple cell types. Moreover, similar patterns were observed in primates and validated in proteomic expressions. Overall, our findings implicate important therapeutic targets for the development of receptor-specific vaccines and drugs for COVID-19.
BACKGROUND:Clinically, accurate pathological diagnosis is often challenged by insufficient tissue amounts and the unaffordability of additional immunohistochemical or genetic tests; thus, there is an urgent need for a universal approach to improve the subtyping of lung cancer without the above limitations. Here we aimed to develop a deep learning system to predict the immunohistochemistry (IHC) phenotype directly from whole-slide images (WSIs) to improve the subtyping of lung cancer from surgical resection and biopsy specimens.METHODS:A total of 1914 patients with lung cancer from three independent hospitals in China were enrolled for WSI-based immunohistochemical feature prediction system (WIFPS) development and validation.RESULTS:The WIFPS could directly predict the IHC status of nine subtype-specific biomarkers, including CK7, TTF-1, Napsin A, CK5/6, P63, P40, CD56, Synaptophysin, and Chromogranin A, achieving average areas under the curve (AUCs) of 0.912, 0.906, and 0.888 and overall diagnostic accuracies of 0.925, 0.941, and 0.887 in the validation datasets of total, external surgical resection specimens and biopsy specimens, respectively. The histological subtyping performance of the WIFPS remained comparable with that of general pathologists (GPs), with Cohen's kappa values ranging from 0.7646 to 0.8282. Furthermore, the WIFPS could be trained to not only predict the IHC status of anaplastic lymphoma kinase (ALK), programmed death-1 (PD-1), and programmed death ligand 1 (PD-L1), but also predict EGFR and KRAS mutation status, with AUCs from 0.525 to 0.917, as detected in separate populations.CONCLUSIONS:In this study, the WIFPS showed its proficiency as a useful complement to traditional histologic subtyping for integrated immunohistochemical spectrum prediction as well as potential in the detection of gene mutations.
Background Immune checkpoint blockade (ICB) therapy has revolutionized the treatment of lung squamous cell carcinoma (LUSC). However, a significant proportion of patients with high tumour PD-L1 expression remain resistant to immune checkpoint inhibitors. To understand the underlying resistance mechanisms, characterization of the immunosuppressive tumour microenvironment and identification of biomarkers to predict resistance in patients are urgently needed. Methods Our study retrospectively analysed RNA sequencing data of 624 LUSC samples. We analysed gene expression patterns from tumour microenvironment by unsupervised clustering. We correlated the expression patterns with a set of T cell exhaustion signatures, immunosuppressive cells, clinical characteristics, and immunotherapeutic responses. Internal and external testing datasets were used to validate the presence of exhausted immune status. Results Approximately 28 to 36% of LUSC patients were found to exhibit significant enrichments of T cell exhaustion signatures, high fraction of immunosuppressive cells (M2 macrophage and CD4 Treg), co-upregulation of 9 inhibitory checkpoints ( CTLA4 , PDCD1 , LAG3 , BTLA , TIGIT , HAVCR2 , IDO1 , SIGLEC7 , and VISTA ), and enhanced expression of anti-inflammatory cytokines (e.g. TGFβ and CCL18). We defined this immunosuppressive group of patients as exhausted immune class (EIC). Although EIC showed a high density of tumour-infiltrating lymphocytes, these were associated with poor prognosis. EIC had relatively elevated PD-L1 expression, but showed potential resistance to ICB therapy. The signature of 167 genes for EIC prediction was significantly enriched in melanoma patients with ICB therapy resistance. EIC was characterized by a lower chromosomal alteration burden and a unique methylation pattern. We developed a web application ( http://lilab2.sysu.edu.cn/tex & http://liwzlab.cn/tex ) for researchers to further investigate potential association of ICB resistance based on our multi-omics analysis data. Conclusions We introduced a novel LUSC immunosuppressive class which expressed high PD-L1 but showed potential resistance to ICB therapy. This comprehensive characterization of immunosuppressive tumour microenvironment in LUSC provided new insights for further exploration of resistance mechanisms and optimization of immunotherapy strategies.
Background Evidence has suggested that cytokine storms may be associated with T cell exhaustion (TEX) in COVID-19. However, the interaction mechanism between cytokine storms and TEX remains unclear. Methods With the aim of dissecting the molecular relationship of cytokine storms and TEX through single-cell RNA sequencing data analysis, we identified 14 cell types from bronchoalveolar lavage fluid of COVID-19 patients and healthy people. We observed a novel subset of severely exhausted CD8 T cells (Exh T_CD8) that co-expressed multiple inhibitory receptors, and two macrophage subclasses that were the main source of cytokine storms in bronchoalveolar. Results Correlation analysis between cytokine storm level and TEX level suggested that cytokine storms likely promoted TEX in severe COVID-19. Cell-cell communication analysis indicated that cytokines (e.g. CXCL10, CXCL11, CXCL2, CCL2, and CCL3) released by macrophages acted as ligands and significantly interacted with inhibitory receptors (e.g. CXCR3, DPP4, CCR1, CCR2, and CCR5) expressed by Exh T_CD8. These interactions formed the cytokine-receptor axes, which were also verified to be significantly correlated with cytokine storms and TEX in lung squamous cell carcinoma. Conclusions Cytokine storms may promote TEX through cytokine-receptor axes and be associated with poor prognosis in COVID-19. Blocking cytokine-receptor axes may reverse TEX. Our finding provides novel insights into TEX in COVID-19 and new clues for cytokine-targeted immunotherapy development.
DNA variants represent an important source of genetic variations among individuals. Next- generation sequencing (NGS) is the most popular technology for genome-wide variant calling. Third-generation sequencing (TGS) has also recently been used in genetic studies. Although many variant callers are available, no single caller can call both types of variants on NGS or TGS data with high sensitivity and specificity. In this study, we systematically evaluated 11 variant callers on 12 NGS and TGS datasets. For germline variant calling, we tested DNAseq and DNAscope modes from Sentieon, HaplotypeCaller mode from GATK and WGS mode from DeepVariant. All the four callers had comparable performance on NGS data and 30× coverage of WGS data was recommended. For germline variant calling on TGS data, we tested DNAseq mode from Sentieon, HaplotypeCaller mode from GATK and PACBIO mode from DeepVariant. All the three callers had similar performance in SNP calling, while DeepVariant outperformed the others in InDel calling. TGS detected more variants than NGS, particularly in complex and repetitive regions. For somatic variant calling on NGS, we tested TNscope and TNseq modes from Sentieon, MuTect2 mode from GATK, NeuSomatic, VarScan2, and Strelka2. TNscope and Mutect2 outperformed the other callers. A higher proportion of tumor sample purity (from 10 to 20%) significantly increased the recall value of calling. Finally, computational costs of the callers were compared and Sentieon required the least computational cost. These results suggest that careful selection of a tool and parameters is needed for accurate SNP or InDel calling under different scenarios.
As immunotherapy is evolving into an essential armamentarium against cancers, numerous translational studies associated with relevant biomarkers, targets, and clinical effects have been reported in recent years. However, a large amount of associated experimental data remains unexplored due to the difficulty in accessibility and utilization. Here, we established a comprehensive high-quality database for cancer immunotherapy called CanImmunother (http://www.biomedical-web.com/cancerit/) through manual curation on 4515 publications. CanImmunother contains 3267 experimentally validated associations between 218 cancer sub-types across 34 body parts and 484 immunotherapies with 642 biomarkers, 108 targets, and 121 control therapies. Each association was manually curated by professional curators, incorporated with valuable annotation and cross references, and assigned with an association score for prioritization. To help clinicians and researchers in identifying and discovering better cancer immunotherapy and their respective biomarkers and targets, CanImmunother offers user-friendly web applications including search, browse, excel table, association prioritization, and network visualization. CanImmunother presents a landscape of experimental cancer immunotherapy association data, serving as a useful resource to improve our insight and to facilitate further discovery of advanced immunotherapy options for cancer patients.
Microbes play important roles in human health and disease. The interaction between microbes and hosts is a reciprocal relationship, which remains largely under-explored. Current computational resources lack manually and consistently curated data to connect metagenomic data to pathogenic microbes, microbial core genes, and disease phenotypes. We developed the MicroPhenoDB database by manually curating and consistently integrating microbe-disease association data. MicroPhenoDB provides 5677 non-redundant associations between 1781 microbes and 542 human disease phenotypes across more than 22 human body sites. MicroPhenoDB also provides 696,934 relationships between 27,277 unique clade-specific core genes and 685 microbes. Disease phenotypes are classified and described using the Experimental Factor Ontology (EFO). A refined score model was developed to prioritize the associations based on evidential metrics. The sequence search option in MicroPhenoDB enables rapid identification of existing pathogenic microbes in samples without running the usual metagenomic data processing and assembly. MicroPhenoDB offers data browsing, searching, and visualization through user-friendly web interfaces and web service application programming interfaces. MicroPhenoDB is the first database platform to detail the relationships between pathogenic microbes, core genes, and disease phenotypes. It will accelerate metagenomic data analysis and assist studies in decoding microbes related to human diseases. MicroPhenoDB is available through http://www.liwzlab.cn/microphenodb and http://lilab2.sysu.edu.cn/microphenodb.
While variants of noncoding RNAs (ncRNAs) have been experimentally validated as a new class of biomarkers and drug targets, the discovery and interpretation of relationships between ncRNA variants and human diseases become important and challenging. Here we present ncRNAVar (http://www.liwzlab.cn/ncrnavar/), the first database that provides association data between validated ncRNA variants and human diseases through manual curation on 2650 publications and computational annotation. ncRNAVar contains 4565 associations between 711 human disease phenotypes and 3112 variants from 2597 ncRNAs. Each association was reviewed by professional curators, incorporated with valuable annotation and cross references, and designated with an association score by our refined score model. ncRNAVar offers web applications including association prioritization, network visualization, and relationship mapping. ncRNAVar, presenting a landscape of ncRNA variants in human diseases and a useful resource for subsequent software development, will improve our insight of relationships between ncRNA variants and human health.
BACKGROUND:Targeted therapy and immunotherapy put forward higher demands for accurate lung cancer classification, as well as benign versus malignant disease discrimination. Digital whole slide images (WSIs) witnessed the transition from traditional histopathology to computational approaches, arousing a hype of deep learning methods for histopathological analysis. We aimed at exploring the potential of deep learning models in the identification of lung cancer subtypes and cancer mimics from WSIs.METHODS:We initially obtained 741 WSIs from the First Affiliated Hospital of Sun Yat-sen University (SYSUFH) for the deep learning model development, optimization, and verification. Additional 318 WSIs from SYSUFH, 212 from Shenzhen People's Hospital, and 422 from The Cancer Genome Atlas were further collected for multi-centre verification. EfficientNet-B5- and ResNet-50-based deep learning methods were developed and compared using the metrics of recall, precision, F1-score, and areas under the curve (AUCs). A threshold-based tumour-first aggregation approach was proposed and implemented for the label inferencing of WSIs with complex tissue components. Four pathologists of different levels from SYSUFH reviewed all the testing slides blindly, and the diagnosing results were used for quantitative comparisons with the best performing deep learning model.RESULTS:We developed the first deep learning-based six-type classifier for histopathological WSI classification of lung adenocarcinoma, lung squamous cell carcinoma, small cell lung carcinoma, pulmonary tuberculosis, organizing pneumonia, and normal lung. The EfficientNet-B5-based model outperformed ResNet-50 and was selected as the backbone in the classifier. Tested on 1067 slides from four cohorts of different medical centres, AUCs of 0.970, 0.918, 0.963, and 0.978 were achieved, respectively. The classifier achieved high consistence to the ground truth and attending pathologists with high intraclass correlation coefficients over 0.873.CONCLUSIONS:Multi-cohort testing demonstrated our six-type classifier achieved consistent and comparable performance to experienced pathologists and gained advantages over other existing computational methods. The visualization of prediction heatmap improved the model interpretability intuitively. The classifier with the threshold-based tumour-first label inferencing method exhibited excellent accuracy and feasibility in classifying lung cancers and confused nonneoplastic tissues, indicating that deep learning can resolve complex multi-class tissue classification that conforms to real-world histopathological scenarios.
Additional nonanatomic prognostic factors beyond TNM categories supplement evidence-based TNM classifications. We aimed to refine TNM staging groups for Epstein-Barr virus (EBV)-related nasopharyngeal carcinoma (NPC) by incorporating the EBV DNA status.