Rapid advances in single-cell RNA sequencing (scRNA-seq) technology have enabled the investigation of gene expression changes at the single-cell level, particularly for elucidating the heterogeneity among cells and complex biological processes. This technique reveals subtle molecular differences within individual cells, thereby offering a unique viewpoint for the investigation of cell cycle progression, cellular differentiation, and disease pathogenesis. However, accurately identifying and analyzing cell cycle dynamics in scRNA-seq data remains challenging due to the complexity of the data and the subtle differences between cell states. To address this challenge, we developed the integrated Sinusoidal and Piecewise AutoEncoder (SPAE), an autoencoder-based piecewise linear model, for characterizing the cell cycle dynamics and cell states in scRNA-seq data. Compared with existing methods, SPAE demonstrates substantially improved accuracy and robustness in cell cycle characterization. Additionally, SPAE can accurately predict cancer cell cycle transitions and effectively facilitate the removal of cell cycle effects from gene expression data. SPAE is available for non-commercial use at https://github.com/YaJahn/SPAE.
Efficient and accurate diagnosis of Mycobacterium tuberculosis is crucial for tuberculosis control. However, traditional diagnostic methods rely on laboratory equipment, making them costly and time-consuming. To improve detection performance, a refined YOLOv5 framework embedding coordinate attention is introduced to strengthen feature representation in challenging environments. A large-scale image dataset of tuberculosis bacilli served as the foundation for training, validation, and evaluation. Model effectiveness was evaluated using precision, recall, and mean average precision at IoU thresholds of 0.5 and 0.5:0.95. Compared to the original YOLOv5 architecture, the enhanced version demonstrated notable improvements, it achieved a 2% increase in precision, 3% improvement in recall, 3.3% boost in mAP@0.5, and 1% enhancement in mAP@0.5:0.95. These advancements were achieved without compromising the model's real-time detection capabilities. The proposed method significantly improves detection accuracy and computational efficiency, providing a robust AI-driven solution for medical image analysis. It has the potential to be integrated into automated screening systems, facilitating early tuberculosis diagnosis and enhancing disease control efforts.
Spatial proteomics can visualize and quantify protein expression profiles within tissues at single-cell resolution. Although spatial proteomics can only detect a limited number of proteins compared to spatial transcriptomics, it provides comprehensive spatial information with single-cell resolution. By studying the spatial distribution of cells, we can clearly obtain the spatial context within tissues at multiple scales. Spatial context includes the spatial composition of cell types, the distribution of functional structures, and the spatial communication between functional regions, all of which are crucial for the patterns of cellular distribution. Here, we constructed a comprehensive spatial proteomics functional annotation knowledgebase, scProAtlas (https://relab.xidian.edu.cn/scProAtlas/#/), which is designed to help users comprehensively understand the spatial context within different tissue types at single-cell resolution and across multiple scales. scProAtlas contains multiple modules, including neighborhood analysis, proximity analysis and neighborhood network, to comprehensively construct spatial cell maps of tissues and multi-modal integration, spatial gene identification, cell-cell interaction and spatial pathway analysis to display spatial variable genes. scProAtlas includes data from eight spatial protein imaging techniques across 15 tissues and provides detailed functional annotation information for 17 468 394 cells from 945 region of interests. The aim of scProAtlas is to offer a new insight into the spatial structure of various tissues and provides detailed spatial functional annotation.
Drug resistance poses a significant challenge in cancer treatment. Despite the initial effectiveness of therapies such as chemotherapy, targeted therapy and immunotherapy, many patients eventually develop resistance. To gain deep insights into the underlying mechanisms, single-cell profiling has been performed to interrogate drug resistance at cell level. Herein, we have built the DRMref database (https://ccsm.uth.edu/DRMref/) to provide comprehensive characterization of drug resistance using single-cell data from drug treatment settings. The current version of DRMref includes 42 single-cell datasets from 30 studies, covering 382 samples, 13 major cancer types, 26 cancer subtypes, 35 treatment regimens and 42 drugs. All datasets in DRMref are browsable and searchable, with detailed annotations provided. Meanwhile, DRMref includes analyses of cellular composition, intratumoral heterogeneity, epithelial-mesenchymal transition, cell-cell interaction and differentially expressed genes in resistant cells. Notably, DRMref investigates the drug resistance mechanisms (e.g. Aberration of Drug's Therapeutic Target, Drug Inactivation by Structure Modification, etc.) in resistant cells. Additional enrichment analysis of hallmark/KEGG (Kyoto Encyclopedia of Genes and Genomes)/GO (Gene Ontology) pathways, as well as the identification of microRNA, motif and transcription factors involved in resistant cells, is provided in DRMref for user's exploration. Overall, DRMref serves as a unique single-cell-based resource for studying drug resistance, drug combination therapy and discovering novel drug targets.
The COVID-19 pandemic, caused by the coronavirus SARS-CoV-2, has resulted in the loss of millions of lives and severe global economic consequences. Every time SARS-CoV-2 replicates, the viruses acquire new mutations in their genomes. Mutations in SARS-CoV-2 genomes led to increased transmissibility, severe disease outcomes, evasion of the immune response, changes in clinical manifestations and reducing the efficacy of vaccines or treatments. To date, the multiple resources provide lists of detected mutations without key functional annotations. There is a lack of research examining the relationship between mutations and various factors such as disease severity, pathogenicity, patient age, patient gender, cross-species transmission, viral immune escape, immune response level, viral transmission capability, viral evolution, host adaptability, viral protein structure, viral protein function, viral protein stability and concurrent mutations. Deep understanding the relationship between mutation sites and these factors is crucial for advancing our knowledge of SARS-CoV-2 and for developing effective responses. To fill this gap, we built COV2Var, a function annotation database of SARS-CoV-2 genetic variation, available at http://biomedbdc.wchscu.cn/COV2Var/. COV2Var aims to identify common mutations in SARS-CoV-2 variants and assess their effects, providing a valuable resource for intensive functional annotations of common mutations among SARS-CoV-2 variants.
Fully understanding traditional Chinese medicines (TCMs) is still challenging because of the extreme complexity of their chemical components and mechanisms of action. The TCM Plant Genome Project aimed to obtain genetic information, determine gene functions, discover regulatory networks of herbal species, and elucidate the molecular mechanisms involved in the disease prevention and treatment, thereby accelerating the modernization of TCMs. A comprehensive database that contains TCM‐related information will provide a vital resource. Here, we present an integrative genome database of TCM plants (IGTCM) that contains 14,711,220 records of 83 annotated TCM‐related herb genomes, including 3,610,350 genes, 3,534,314 proteins and corresponding coding sequences, and 4,032,242 RNAs, as well as 1033 non‐redundant component records for 68 herbs, downloaded and integrated from the GenBank and RefSeq databases. For minimal interconnectivity, each gene, protein, and component was annotated using the eggNOG‐mapper tool and Kyoto Encyclopedia of Genes and Genomes database to acquire pathway information and enzyme classifications. These features can be linked across several species and different components. The IGTCM database also provides visualization and sequence similarity search tools for data analyses. These annotated herb genome sequences in IGTCM database are a necessary resource for systematically exploring genes related to the biosynthesis of compounds that have significant medicinal activities and excellent agronomic traits that can be used to improve TCM‐related varieties through molecular breeding. It also provides valuable data and tools for future research on drug discovery and the protection and rational use of TCM plant resources. The IGTCM database is freely available at http://yeyn.group:96/.
This study systematically analyses the mechanism of Spike(S) protein and bioinformatics research progress by mining domestic and foreign literatures. This paper summarizes the current research on use of bioinformatics tools by scientists at home and abroad to develop drugs and vaccines research of Spike protein. In addition, the existing Spike protein data resources are systematically organized. Related bioinformatics tools and databases are summarized, and the research prospects of Spike protein in this field are proposed. This review aims to provide systematic Spike protein research references for biologists,virologists and other researchers who study Spike proteins. It promotes the research of Spike protein mechanism and the development of effective drugs and vaccines.
Spike蛋白是I类融合糖蛋白,主要发挥糖蛋白的作用。尤其广泛存在于病毒中,在病毒入侵细胞时介导受体结合和膜融合,从而影响细胞的感染倾向。随着新型冠状病毒肺炎(Corona Virus Disease 2019, COVID-19)疫情爆发,生物学和毒理学家们利用生物信息学手段对Spike蛋白进行测序、注释等方法探索该蛋白的结构、功能和进化关系,发现S蛋白是病毒有效治疗、疫苗研发以及临床诊断的重要靶点。本研究通过挖掘国内外文献,系统地分析了S蛋白的作用机制和生物信息学方面的研究进展。总结分析了现有国内外与Spike蛋白相关的生物信息学工具及数据库,最后对病毒Spike蛋白在该领域的研究提出了展望。