Although the spatial characteristics within the tumor microenvironment of lung adenocarcinoma(LUAD)have been identified,the mechanisms by which these factors promote LUAD progression and immune evasion remain unclear.Using spatial transcriptomics and single-cell RNA-sequencing data from multi-regional LUAD biopsies consisting of tumor core,tumor edge,and normal area,we sought to delineate the spatial hetero-geneity and driving factors of cell colocalization.Two cancer cell sub-clusters(Cancer_c1 and Cancer_c2),associated with LUAD initiation and metastasis,respectively,exhibit distinct spatial distributions and immune cell colocalizations.In particular,Cancer_c1,enriched within the tumor core,could directly interact with B cells or indirectly recruit B cells through macrophages.Conversely,Cancer_c2 enriched within the tumor edge exhibits colocalization with CD8+T cells.Collectively,our work elucidates the spatial distribution of cancer cell subtypes and their interaction with immune cells in the core and edge of LUAD,providing insights for developing therapeutic strategies for cancer intervention.
Dear Editor, Drug repurposing is at the forefront of a transformative shift in computational methods driving new applications of approved or investigational drugs.1, 2 With the development of network pharmacology, repositioning algorithms for drug effects or drug targets are constantly expanding, but integrating multidimensional data to achieve precise repurposing is still a challenge.3-9 We focus on drug attribute characteristics and propose a scalable systematic paradigm. Using the Genomics of Drug Sensitivity in Cancer (GDSC) database for anti-tumor drugs, a integrated drug similarity network (iDSN) derived from different drug similarity networks (DSNs) based on chemical structure and drug target sequence data is constructed to infer potential drug pathways from drug properties and realise drug repurposing. Initially, we processed drug profile data by vectorizing it (Figure 1A). Based on chemical and pharmacological properties, we constructed two separate DSNs: chem-DSN and pharm-DSN. These were then merged into an iDSN using a nonlinear fusion algorithm called Similarity Network Fusion (SNF) (Figure 1B). To validate the iDSN's potential in therapeutic similarity, we utilized a spectral clustering model with seven gold-standard annotations from PubChem (Figure 1C). Downstream analysis was delineated across three dimensions for drug repurposing (Figure 1D): (1) identifying similar components within classes, amalgamating pharmacological mechanisms with pathway annotation; (2) establishing associations between drug network clusters and distinct biological pathways; (3) prioritizing higher-ranked drug pairs for drug repositioning. Employing spectral clustering, iDSN exhibited a more distinct clustering structure compared to the chem-DSN and a more evenly distributed structure than the pharm-DSN (Figures 2A and S1). With the advantage of framework transparency, pharmacological properties contribute more than chemical properties through the quantitative assessment in clustering (Figure 2B). In 11 clusters, pharmacological features accounted for over 70% of edge similarity, and four clusters were entirely determined by pharm-DSN. In comparison to single-property DSNs, iDSN demonstrated superior performance across all three metrics (Figure 2C). Evaluation using six diverse benchmark datasets of drug categories confirmed iDSN's stronger correlations with all benchmark annotations compared to single-property DSNs (Table S1). Comparing clustering performance among single-property DSNs, pharm-DSN displayed better interaction with cell line (ARI = .470) and pathway (ARI = .585) annotations (Figure 2D). Importantly, iDSN based on the cross-fusion network algorithm achieves higher performance on IC50 (ARI = .502, NMI = .53), indicating improved generalisation and accuracy through multi-feature fusion (Figures 2E and S2). Data contribution analysis within the IC50-based cluster of the three DSNs highlighted iDSN's predominant contribution (77.23%), while chem-DSN and pharm-DSN contributed less (9.83% and 12.94%, respectively), underscoring iDSN's dominance in the IC50-based network (Figure 2F). To validate the superior performance, our method was compared with state-of-the-art approaches, encompassing traditional machine learning, network propagation and matrix factorisation. The framework demonstrated a significantly higher value (SC = .58) compared to other methods. Similar results were observed with the NMI index using the IC50 dataset, where our framework outperformed in interactivity score (ARI = .512) (Table S2). To facilitate drug precision repositioning, we uncovered the drug preferences within each cluster for various molecular functions using the KEGG pathway and Gene Ontology (GO) annotations10 (Figure 3). Notably, certain highly similar drug pairs within clusters exhibited consistent downstream pathway annotations, indicating the potential for drug repositioning by leveraging common targets or similar cellular signalling pathways to achieve therapeutic effects. The results revealed that several drug clusters exhibited significant enrichment annotations on GO analysis, such as Cluster 4, Cluster 2, and Cluster 5 included in the histone deacetylation, ADP ribosylation and phosphorylation respectively based on biological process. Furthermore, we observed that some individual drug clusters met different KEGG enrichment pathways under secondary classification, but specific on GO enrichment analysis in biological processes, cell components or molecular functions. For instance, in Cluster 9, we observed enrichment in KEGG pathways related to cancer and cell growth and death, with emphasis on apoptosis and molecular functions associated with dimerisation in biological processes, which suggests that Cluster 9 may exhibit a pharmacodynamic pattern, potentially influencing protein dimerisation and participating in pathways related to cancer or cell growth and death through apoptosis regulation. Exploring drug pairs with high similarity in the iDSN reveals potential drug repositioning opportunities. For the top 100 similar drug pairs, 86% had consistent pathway annotations, confirming the reliability of our drug similarity calculations. However, some pairs with high similarity scores had different annotations, mainly linked to six pathways in four classification clusters (Figure 4A). For instance, a closely related set of drug pairs, including BMS-536924, BMS-754807, GSK1904529A, Linsitinib and NVP-ADW742, exhibited connections to annotations in the IGF1R signalling and RTK signalling pathways, suggesting potential shared targets or interactions with similar cellular signalling pathways. In a major cluster, drugs associated with the kinases pathway (KIN001-244) showed high similarity to drugs linked to the Metabolism pathway (BX-912 and OSU-03012) and the Mitosis pathway (MPS-1-IN-1), unveiling potential crosstalk for therapeutic strategies (Table S3). Among them, the drug pair with the highest similarity is CMK-LJI308 (ranked sixth), annotated in the kinase pathway and the PI3K/MTOR signalling pathway, respectively. Drug pairs with different annotation pathways in specific spectral clustering clusters indicate distinct subgroups with unique downstream pathway preferences. Some clusters may exhibit pathway-specific therapeutic effects, while others show divergent pathway orientations (Figure 4B). To explore the global repositioning associations, we aligned pathway annotations with clustering results and observed a well-balanced distribution of downstream pathways across drug clusters (Figure 4C). Within one cluster containing nine drugs, four were associated with the IGF1R signalling pathway, and the remaining drugs were linked to the RTK signalling pathway. In another cluster with 37 drugs, 14 were mapped to the RTK pathway, while the remaining drugs were connected to the kinase pathway. Furthermore, we independently analysed highly similar drug pairs within these clusters (Figure 4D). In conclusion, this scalable structure-derived framework offers fresh insights into deducing characteristic downstream pathways and repurposing drugs via common drug structural properties. With the accumulation of drug informatics data and the development of future drugs, we will continue to expand our data, to deepen our understanding of feature integration, and to further improve the algorithm's performance for new drug development. Zhaoman Wan performed the analysis and prepared the manuscript with the help of Yang Cao, Mingming Su, Xinlei Zhang, L.Y. and H.X. Aiping Wu, Peng Zhang and Taijiao Jiang supervised the studies, designed the analysis and revised the manuscript. All authors reviewed and approved the manuscript. M.S. and X.Z. are the co-founders of Beijing Cloudna Technology Co., Ltd., and the other authors declare no competing interests. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
OBJECTIVE:To facilitate the identification of related genes and candidate biomarkers for disorders of sex development (DSD), we present disorders of sex development atlas (http://dsd.geneworks.cn). Disorders of sex development are a spectrum of endocrine diseases with distinct mutations of genes or chromosomes, but several issues regarding their pathogenesis remain elusive. High-throughput methods have allowed genomic and transcriptomic analyses of DSD; however, these data are deposited in various repositories owing to a lack of integrated online resources.DESIGN:A descriptive study of a specialized gene discovery platform designed for DSD.SETTING:Publicly available DSD omics datasets and self-produced datasets.PATIENT(S):None.INTERVENTION(S):None.MAIN OUTCOME MEASURE(S):The gene ranking result, with detailed information based on DSD terms in a gene-disease association knowledge base, and results of differential gene expression and mutation analyses from omics datasets.RESULT(S):The disorders of sex development atlas maintains both a knowledgebase for ranking DSD candidate genes and a database for DSD-related omics data analysis and visualization. We included 4 dominant classes of DSD in the knowledgebase: 15 subclasses and 44 specific disease names. Construction of the knowledgebase was centered upon Phenolyzer, with add-on seed gene databases customized by DSD-related genes collected from MalaCards, GeneCards, and DisGeNET. For the database, 25 experimental datasets related to DSD were integrated, including 24 public datasets from Gene Expression Omnibus and Sequence Read Archive and 1 self-generated dataset. A total of 474 samples from 240 DSD samples were collected for the database.CONCLUSION(S):This platform provides a friendly interface that integrates flexible and comprehensive analysis tools for differential expression and gene mutations between the DSD groups and controls.
Ectopic Cushing's syndrome due to ectopic ACTH&CRH-secreting by pheochromocytoma is extremely rare and can be fatal if not properly diagnosed. It remains unclear whether a unique cell type is responsible for multiple hormones secreting. In this work, we performed single-cell RNA sequencing to three different anatomic tumor tissues and one peritumoral tissue based on a rare case with ectopic ACTH&CRH-secreting pheochromocytoma. And in addition to that, three adrenal tumor specimens from common pheochromocytoma and adrenocortical adenomas were also involved in the comparison of tumor cellular heterogeneity. A total of 16 cell types in the tumor microenvironment were identified by unbiased cell clustering of single-cell transcriptomic profiles from all specimens. Notably, we identified a novel multi-functionally chromaffin-like cell type with high expression of both POMC (the precursor of ACTH) and CRH, called ACTH+&CRH + pheochromocyte. We hypothesized that the molecular mechanism of the rare case harbor Cushing's syndrome is due to the identified novel tumor cell type, that is, the secretion of ACTH had a direct effect on the adrenal gland to produce cortisol, while the secretion of CRH can indirectly stimulate the secretion of ACTH from the anterior pituitary. Besides, a new potential marker (GAL) co-expressed with ACTH and CRH might be involved in the regulation of ACTH secretion. The immunohistochemistry results confirmed its multi-functionally chromaffin-like properties with positive staining for CRH, POMC, ACTH, GAL, TH, and CgA. Our findings also proved to some extent the heterogeneity of endothelial and immune microenvironment in different adrenal tumor subtypes.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Supervised learning methods are commonly applied in medical image analysis. However, the success of these approaches is highly dependent on the availability of large manually detailed annotated dataset. Thus an automatic refined segmentation of whole-slide image (WSI) is significant to alleviate the annotation workload of pathologists. But most of the current ways can only output a rough prediction of lesion areas and consume much time in each slide. In this paper, we propose a fast and refined cancer regions segmentation framework v3_DCNN, which first preselects tumor regions using a classification model Inception-v3 and then employs a semantic segmentation model DCNN for refined segmentation. Our framework can generate a dense likelihood heatmap with the 1/8 side of original WSI in 11.5 minutes on the Camelyon16 dataset, which saves more than one hour for each WSI compared with the initial DCNN model. Experimental results show that our approach achieves a higher FROC score 83.5% with the champion’s method of Camelyon16 challenge 80.7%. Based on v3 DCNN model, we further automatically produce heatmap of WSI and extract polygons of lesion regions for doctors, which is very helpful for their pathological diagnosis, detailed annotation and thus contributes to developing a more powerful deep learning model.
Yersinia enterocolitica is the most diverse species among the Yersinia genera and shows more polymorphism, especially for the non-pathogenic strains. Individual non-pathogenic Y. enterocolitica strains are wrongly identified because of atypical phenotypes. In this study, we isolated an unusual Y. enterocolitica strain LC20 from Rattus norvegicus. The strain did not utilize urea and could not be classified as the biotype. API 20E identified Escherichia coli; however, it grew well at 25 °C, but E. coli grew well at 37 °C. We analyzed the genome of LC20 and found the whole chromosome of LC20 was collinear with Y. enterocolitica 8081, and the urease gene did not exist on the genome which is consistent with the result of API 20E. Also, the 16 S and 23 SrRNA gene of LC20 lay on a branch of Y. enterocolitica. Furthermore, the core-based and pan-based phylogenetic trees showed that LC20 was classified into the Y. enterocolitica cluster. Two plasmids (80 and 50 k) from LC20 shared low genetic homology with pYV from the Yersinia genus, one was an ancestral Yersinia plasmid and the other was novel encoding a number of transposases. Some pathogenic and non-pathogenic Y. enterocolitica-specific genes coexisted in LC20. Thus, although it could not be classified into any Y. enterocolitica biotype due to its special biochemical metabolism, we concluded the LC20 was a Y. enterocolitica strain because its genome was similar to other Y. enterocolitica and it might be a strain with many mutations and combinations emerging in the processes of its evolution.
Bacteriophages and their hosts are continuously engaged in evolutionary competition. Here we isolated a lytic phage phiYe-F10 specific for Yersinia enterocolitica serotype O:3. We firstly described the phage receptor was regulated by DTDP-rhamnosyl transferase RfbF, encoded within the rfb cluster that was responsible for the biosynthesis of the O antigens. The deletion of DTDP-rhamnosyl transferase RfbF of wild type O:3 strain caused failure in phiYe-F10 adsorption; however, the mutation strain retained agglutination with O:3 antiserum; and complementation of its mutant converted its sensitivity to phiYe-F10. Therefore, DTDP-rhamnosyl transferase RfbF was responsible for the phage infection but did not affect recognition of Y. enterocolitica O:3 antiserum. Further, the deletions in the putative O-antigen biosynthesis protein precursor and outer membrane protein had no effect on sensitivity to phiYe-F10 infection. However, adsorption of phages onto mutant HNF10-ΔO-antigen took longer time than onto the WT, suggesting that deletion of the putative O-antigen biosynthesis protein precursor reduced the infection efficiency.
API 20E strip test, the standard for Enterobacteriaceae identification, is not sufficient to discriminate some Yersinia species for some unstable biochemical reactions and the same biochemical profile presented in some species, e.g. Yersinia ferderiksenii and Yersinia intermedia, which need a variety of molecular biology methods as auxiliaries for identification. The 16S rRNA gene is considered a valuable tool for assigning bacterial strains to species. However, the resolution of the 16S rRNA gene may be insufficient for discrimination because of the high similarity of sequences between some species and heterogeneity within copies at the intra-genomic level. In this study, for each strain we randomly selected five 16S rRNA gene clones from 768 Yersinia strains, and collected 3,840 sequences of the 16S rRNA gene from 10 species, which were divided into 439 patterns. The similarity among the five clones of 16S rRNA gene is over 99% for most strains. Identical sequences were found in strains of different species. A phylogenetic tree was constructed using the five 16S rRNA gene sequences for each strain where the phylogenetic classifications are consistent with biochemical tests; and species that are difficult to identify by biochemical phenotype can be differentiated. Most Yersinia strains form distinct groups within each species. However Yersinia kristensenii, a heterogeneous species, clusters with some Yersinia enterocolitica and Yersinia ferderiksenii/intermedia strains, while not affecting the overall efficiency of this species classification. In conclusion, through analysis derived from integrated information from multiple 16S rRNA gene sequences, the discrimination ability of Yersinia species is improved using our method.
High-throughput sequencing-based metagenomics has garnered considerable interest in recent years. Numerous methods and tools have been developed for the analysis of metagenomic data. However, it is still a daunting task to install a large number of tools and complete a complicated analysis, especially for researchers with minimal bioinformatics backgrounds. To address this problem, we constructed an automated software named MetaDP for 16S rRNA sequencing data analysis, including data quality control, operational taxonomic unit clustering, diversity analysis, and disease risk prediction modeling. Furthermore, a support vector machine-based prediction model for intestinal bowel syndrome (IBS) was built by applying MetaDP to microbial 16S sequencing data from 108 children. The success of the IBS prediction model suggests that the platform may also be applied to other diseases related to gut microbes, such as obesity, metabolic syndrome, or intestinal cancer, among others (http://metadp.cn:7001/).
Background: The data released by the 1000 Genomes Project contain an increasing number of genome sequences from different nations and populations with a large number of genetic variations. As a result, the focus of human genome studies is changing from single and static to complex and dynamic. The currently available human reference genome (GRCh37) is based on sequencing data from 13 anonymous Caucasian volunteers, which might limit the scope of genomics, transcriptomics, epigenetics, and genome wide association studies. Description: We used the massive amount of sequencing data published by the 1000 Genomes Project Consortium to construct the Virtual Chinese Genome Database (VCGDB), a dynamic genome database of the Chinese population based on the whole genome sequencing data of 194 individuals. VCGDB provides dynamic genomic information, which contains 35 million single nucleotide variations (SNVs), 0.5 million insertions/deletions (indels), and 29 million rare variations, together with genomic annotation information. VCGDB also provides a highly interactive user-friendly virtual Chinese genome browser (VCGBrowser) with functions like seamless zooming and real-time searching. In addition, we have established three population-specific consensus Chinese reference genomes that are compatible with mainstream alignment software. Conclusions: VCGDB offers a feasible strategy for processing big data to keep pace with the biological data explosion by providing a robust resource for genomics studies; in particular, studies aimed at finding regions of the genome associated with diseases.
Background: The data released by the 1000 Genomes Project contain an increasing number of genome sequences from different nations and populations with a large number of genetic variations. As a result, the focus of human genome studies is changing from single and static to complex and dynamic. The currently available human reference genome (GRCh37) is based on sequencing data from 13 anonymous Caucasian volunteers, which might limit the scope of genomics, transcriptomics, epigenetics, and genome wide association studies.Description: We used the massive amount of sequencing data published by the 1000 Genomes Project Consortium to construct the Virtual Chinese Genome Database (VCGDB), a dynamic genome database of the Chinese population based on the whole genome sequencing data of 194 individuals. VCGDB provides dynamic genomic information, which contains 35 million single nucleotide variations (SNVs), 0.5 million insertions/deletions (indels), and 29 million rare variations, together with genomic annotation information. VCGDB also provides a highly interactive user-friendly virtual Chinese genome browser (VCGBrowser) with functions like seamless zooming and real-time searching. In addition, we have established three population-specific consensus Chinese reference genomes that are compatible with mainstream alignment software.Conclusions: VCGDB offers a feasible strategy for processing big data to keep pace with the biological data explosion by providing a robust resource for genomics studies; in particular, studies aimed at finding regions of the genome associated with diseases.
Polypeptides containing ≤100 amino acid residues (AAs) are generally considered to be small proteins (SPs). Many studies have shown that some SPs are involved in important biological processes, including cell signaling, metabolism, and growth. SP generally has a simple domain and has an advantage to be used as model system to overcome folding speed limits in protein folding simulation and drug design. But SPs were once thought to be trivial molecules in biological processes compared to large proteins. Because of the constraints of experimental methods and bioinformatics analysis, many genome projects have used a length threshold of 100 amino acid residues to minimize erroneous predictions and SPs are relatively under-represented in earlier studies. The general protein discovery methods have potential problems to predict and validate SPs, and very few effective tools and algorithms were developed specially for SPs identification. In this review, we mainly consider the diverse strategies applied to SPs prediction and discuss the challenge for differentiate SP coding genes from artifacts. We also summarize current large-scale discovery of SPs in species at the genome level. In addition, we present an overview of SPs with regard to biological significance, structural application, and evolution characterization in an effort to gain insight into the significance of SPs.