Large-scale genomic initiatives like the UK Biobank have revolutionized our understanding of human disease. These studies typically assume that blood-derived DNA faithfully reflects an individual’s germline genome. However, this assumption is challenged by somatic mutations arising from processes like clonal hematopoiesis. Although standard bioinformatics pipelines employ variant allele frequency (VAF)-based filtering to mitigate such contamination, the efficacy of these approaches requires systematic evaluation. By systematically analyzing germline genome data from large cohorts through applications of mutational signatures, we revealed critical limitations in current filtering methodologies. We found that the mutational spectrum of rare “germline” variants is highly similar to that of somatic mutations. Furthermore, we uncovered that these variants show significant associations with phenotypes such as age, sex, and smoking status, established drivers of somatic mutagenesis. Notably, our multivariable regression models estimated that these somatic artifacts contribute to a substantial excess burden, such as 4.73 mutations per megabase (mut/Mb) in males compared to females, a magnitude exceeding the mutation burden of many cancers. Although the precise absolute size of this contamination may vary depending on specific pathologies and individual environmental exposures, this persistent somatic contamination introduces substantial confounder effects, posing a risk of spurious associations and reverse causality in genetic studies. Our work underscores the urgent reconsideration of two fundamental aspects of genomic research: (1) refinement of variant filtering strategies to better distinguish true germline variants from somatic contaminants, and (2) incorporation of somatic mutagenesis factors as essential covariates in study design. Our findings provide basic guidance for improving the accuracy and interpretability of large-scale genomic studies.
Cells are the fundamental units of life and exhibit significant diversity in structure, behavior, and function, known as cell heterogeneity. The advent and development of single-cell RNA sequencing (scRNA-seq) technology have provided a crucial data foundation for studying cellular heterogeneity. Currently, most computational methods based on scRNA-seq involve a sequential process of clustering followed by annotation. However, those clustering-based methods are susceptible to the selection of genes and clustering parameters, resulting in inaccuracies in cell annotation. To address this issue, we develop a flexible data-driven cell correction framework based on partially annotated scRNA-seq data. This framework employs a neighborhood purity strategy and global selection strategies to select the anchor cells. Then, it optimizes a prediction neural network model using a classification loss with a contrastive regularization term to correct the labels of the remaining cells. The validity of this correction framework is demonstrated through various assessments on real scRNA-seq datasets. Based on the correct labels of scRNA-seq data, we further assess the latest unsupervised clustering methods, thereby establishing a more objective benchmark to compare their performance.
Large-scale genomic initiatives like the UK Biobank (UKB) have revolutionized our understanding of human disease. These studies typically assume that blood-derived DNA faithfully reflects an individual’s germline genome. However, this assumption is challenged by somatic mutations arising from processes like clonal hematopoiesis. Although standard bioinformatics pipelines employ variant allele frequency (VAF)-based filtering to mitigate such contamination, the efficacy of these approaches requires systematic evaluation. By systematically analyzing whole-exome sequencing (WES) and whole-genome sequencing (WGS) data from large cohorts, including the UK Biobank, The Cancer Genome Atlas (TCGA), and the 1000 Genomes Project, we revealed critical limitations in current filtering methodologies. We found that the mutational spectrum of rare “germline” variants is highly similar to that of somatic mutations. Furthermore, we uncovered these variants show significant associations with phenotypes such as age, sex, and smoking status, established drivers of somatic mutagenesis. This persistent somatic contamination introduces substantial confounder effects, potentially generating spurious associations and reverse causality in genetic studies. Our work underscores the urgent reconsideration of two fundamental aspects of genomic research: (1) refinement of variant filtering strategies to better distinguish true germline variants from somatic contaminants, and (2) incorporation of somatic mutagenesis factors as essential covariates in study design. Our findings provide a basic framework for improving the accuracy and interpretability of large-scale genomic studies. ### Competing Interest Statement The authors have declared no competing interest. National Natural Science Foundation of China, 62025102, 82470373
The emergence of immune checkpoint inhibitors (ICIs) has significantly advanced cancer treatment. However, only 15-30% of the cancer patients respond to ICI treatment, which stimulates and enhances host immunity to eliminate tumor cells. ICI treatment is very expensive and has potential adverse reactions; therefore, it is crucial to develop a method which enables to accurately and rapidly assess a patient's suitability before ICI treatment. We complied germline whole-genome sequencing (WES) data of 37 melanoma patients who have been treated with ICIs and sequenced in our lab previously, and the WES data of other 700 ICI-treated cancer patients in public domain. Using these data, we proposed a novel double-channel attention neural network (DANN) model to predict cancer ICI-response and validate the predictions. DANN achieved a mean accuracy and AUC of 0.95 and 0.98, respectively, which outperformed traditional machine learning methods. Enrichment analysis of the DANN-identified genes indicated that cancer patients whose in-born genomic variants might mainly affect host immune system in a wide-ranging manner, and then affect ICI response. Finally, we found a set of 12 genes bearing genomic variants were significantly associated with cancer patient survivals after ICI treatment.
BACKGROUND:Tumour-infiltrating lymphocytes (TILs) are crucial for effective immune checkpoint blockade (ICB) therapy in solid tumours. However, ∼70% of these tumours exhibit poor lymphocyte infiltration, rendering ICB therapies less effective.METHODS:We developed a bioinformatics pipeline integrating multiple previously unconsidered factors or datasets, including tumour cell immune-related pathways, copy number variation (CNV), and single tumour cell sequencing data, as well as tumour mRNA-seq data and patient survival data, to identify targets that can potentially improve T cell infiltration and enhance ICB efficacy. Furthermore, we conducted wet-lab experiments and successfully validated one of the top-identified genes.FINDINGS:We applied this pipeline in solid tumours of the Cancer Genome Atlas (TCGA) and identified a set of genes in 18 cancer types that might potentially improve lymphocyte infiltration and ICB efficacy, providing a valuable drug target resource to be further explored. Importantly, we experimentally validated SUN1, which had not been linked to T cell infiltration and ICB therapy previously, but was one of the top-identified gene targets among 3 cancer types based on the pipeline, in a mouse colon cancer syngeneic model. We showed that Sun1 KO could significantly enhance antigen presentation, increase T-cell infiltration, and improve anti-PD1 treatment efficacy. Moreover, with a single-cell multiome analysis, we identified subgene regulatory networks (sub-GRNs) showing Stat proteins play important roles in enhancing the immune-related pathways in Sun1-KO cancer cells.INTERPRETATION:This study not only established a computational pipeline for discovering new gene targets and signalling pathways in cancer cells that block T-cell infiltration, but also provided a gene target pool for further exploration in improving lymphocyte infiltration and ICB efficacy in solid tumours.FUNDING:A full list of funding bodies that contributed to this study can be found in the Acknowledgements section.
Metastasis remains a major challenge in treating breast cancer. Breast tumors metastasize to organ-specific locations such as the brain, lungs, and bone, but why some organs are favored over others remains unclear. Breast tumors also show heterogeneity, plasticity, and distinct microenvironments. This contributes to treatment failure and relapse. The interaction of breast cancer cells with their metastatic microenvironment has led to the concept that primary breast cancer cells act as seeds, whereas the metastatic tissue microenvironment (TME) is the soil. Improving our understanding of this interaction could lead to better treatment strategies for metastatic breast cancer. Targeted treatments for different subtypes of breast cancers have improved overall patient survival, even with metastasis. However, these targeted treatments are based upon the biology of the primary tumor and often these patients’ relapse, after therapy, with metastatic tumors. The advent of immunotherapy allowed the immune system to target metastatic tumors. Unfortunately, immunotherapy has not been as effective in metastatic breast cancer relative to other cancers with metastases, such as melanoma. This review will describe the heterogeneic nature of breast cancer cells and their microenvironments. The distinct properties of metastatic breast cancer cells and their microenvironments that allow interactions, especially in bone and brain metastasis, will also be described. Finally, we will review immunotherapy approaches to treat metastatic breast tumors and discuss future therapeutic approaches to improve treatments for metastatic breast cancer.
BACKGROUND:In the past decade, single nucleotide variants (SNVs) have been identified as having a significant relationship with the development and treatment of diseases. Among them, prioritizing missense variants for further functional impact investigation is an essential challenge in the study of common disease and cancer. Although several computational methods have been developed to predict the functional impacts of variants, the predictive ability of these methods is still insufficient in the Mendelian and cancer missense variants.RESULTS:We present a novel prediction method called the disease-related variant annotation (DVA) method that predicts the effect of missense variants based on a comprehensive feature set of variants, notably, the allele frequency and protein-protein interaction network feature based on graph embedding. Benchmarked against datasets of single nucleotide missense variants, the DVA method outperforms the state-of-the-art methods by up to 0.473 in the area under receiver operating characteristic curve. The results demonstrate that the proposed method can accurately predict the functional impact of single nucleotide missense variants and substantially outperforms existing methods.CONCLUSIONS:DVA is an effective framework for identifying the functional impact of disease missense variants based on a comprehensive feature set. Based on different datasets, DVA shows its generalization ability and robustness, and it also provides innovative ideas for the study of the functional mechanism and impact of SNVs.
The Single-cell Assay for Transposase-Accessible Chromatin with high throughput sequencing (scATAC-seq) has gained increasing popularity in recent years, allowing for chromatin accessibility to be deciphered and gene regulatory networks (GRNs) to be inferred at single-cell resolution. This cutting-edge technology now enables the genome-wide profiling of chromatin accessibility at the cellular level and the capturing of cell-type-specific cis-regulatory elements (CREs) that are masked by cellular heterogeneity in bulk assays. Additionally, it can also facilitate the identification of rare and new cell types based on differences in chromatin accessibility and the charting of cellular developmental trajectories within lineage-related cell clusters. Due to technical challenges and limitations, the data generated from scATAC-seq exhibit unique features, often characterized by high sparsity and noise, even within the same cell type. To address these challenges, various bioinformatic tools have been developed. Furthermore, the application of scATAC-seq in plant science is still in its infancy, with most research focusing on root tissues and model plant species. In this review, we provide an overview of recent progress in scATAC-seq and its application across various fields. We first conduct scATAC-seq in plant science. Next, we highlight the current challenges of scATAC-seq in plant science and major strategies for cell type annotation. Finally, we outline several future directions to exploit scATAC-seq technologies to address critical challenges in plant science, ranging from plant ENCODE(The Encyclopedia of DNA Elements) project construction to GRN inference, to deepen our understanding of the roles of CREs in plant biology.
The association of neurogenesis and gliogenesis with glioma remains unclear. By conducting single-cell RNA-seq analyses on 26 gliomas, we reported their classification into primitive oligodendrocyte precursor cell (pri-OPC)-like and radial glia (RG)-like tumors and validated it in a public cohort and TCGA glioma. The RG-like tumors exhibited wild-type isocitrate dehydrogenase and tended to carry EGFR mutations, and the pri-OPC-like ones were prone to carrying TP53 mutations. Tumor subclones only in pri-OPC-like tumors showed substantially down-regulated MHC-I genes, suggesting their distinct immune evasion programs. Furthermore, the two subgroups appeared to extensively modulate glioma-infiltrating lymphocytes in distinct manners. Some specific genes not expressed in normal immune cells were found in glioma-infiltrating lymphocytes. For example, glial/glioma stem cell markers OLIG1/PTPRZ1 and B cell-specific receptors IGLC2/IGKC were expressed in pri-OPC-like and RG-like glioma-infiltrating lymphocytes, respectively. Their expression was positively correlated with those of immune checkpoint genes (e.g., LGALS33) and poor survivals as validated by the increased expression of LGALS3 upon IGKC overexpression in Jurkat cells. This finding indicated a potential inhibitory role in tumor-infiltrating lymphocytes and could provide a new way of cancer immune evasion.
At present, some methods have been proposed to solve the problem from the perspective of the chaos degree in gene functions or interaction network. However, errors in differentiation potency estimates arise if the scRNA-seq profile and underlying interaction network are disturbed by technique-induced or biological-induced noise. Thus, we proposed SPIDE, a single cell potency inference method based on local cell-specific network entropy. SPIDE constructs the weighted cell-specific network for each cell to preserve the heterogeneity of PPI network during differentiation, then estimates the entropy based on each network. The results show that SPIDE reveals better decreasing trends of cells’ differentiation potency than other state-of-the-art methods on most datasets. To conclude, our study provides a universal framework for cell entropy estimation with higher prediction accuracy and universal applicability.
Somatic mutational signatures (MSs) identified by genome sequencing play important roles in exploring the cause and development of cancer. Thus far, many such signatures have been identified, and some of them do imply causes of cancer. However, a major bottleneck is that we do not know the potential meanings (i.e. carcinogenesis or biological functions) and contributing genes for most of them. Here, we presented a computational framework, Gene Somatic Genome Pattern (GSGP), which can decipher the molecular mechanisms of the MSs. More importantly, it is the first time that the GSGP is able to process MSs from ribonucleic acid (RNA) sequencing, which greatly extended the applications of both MS analysis and RNA sequencing (RNAseq). As a result, GSGP analyses match consistently with previous reports and identify the etiologies for a number of novel signatures. Notably, we applied GSGP to RNAseq data and revealed an RNA-derived MS involved in deficient deoxyribonucleic acid mismatch repair and microsatellite instability in colorectal cancer. Researchers can perform customized GSGP analysis using the web tools or scripts we provide.
AbstractBackgroundAnti‐programmed death‐1 (PD‐1) immunotherapy has drastically improved survival for metastatic melanoma; however, 50% of patients have progression within 6 months despite treatment. In this study, we investigated host, and tumor factors for metastatic melanoma patients treated with anti‐PD‐1 immunotherapy.MethodsPatients treated with the anti‐PD‐1 immunotherapy between 2014 and 2017 were identified in Alberta, Canada. All patients had Stage IV melanoma. Patient characteristics, investigations, treatment, and clinical outcomes were obtained from electronic medical records.ResultsWe identified 174 patients treated with anti‐PD‐1 immunotherapy. At 37.1 months median follow‐up time 135 (77.6%) individuals had died and 150 (86.2%) had progressed. An elevated lactate dehydrogenase (LDH) had a response rate of 21.0% versus 41.0% for those with a normal LDH (p = 0.017). Host factors associated with worse median progression‐free survival (mPFS) and median overall survival (mOS) included liver metastases, >3 sites of disease, elevated LDH, thrombocytosis, neutrophilia, anemia, lymphocytopenia, and an elevated neutrophil/lymphocyte ratio. Primary ulcerated tumors had a worse mOS of 11.8 versus 19.3 months (p = 0.042). We identified four prognostic subgroups in advanced melanoma patients treated with anti‐PD‐1 therapy. (1) Normal LDH with <3 visceral sites, (2) normal LDH with ≥3 visceral sites, (3) LDH 1‐2x upper limit of normal (ULN), (4) LDH ≥2x ULN. The mPFS each group was 14.0, 6.5, 3.3, and 1.9 months, while the mOS for each group was 33.3, 15.7, 7.9, and 3.4 months.ConclusionOur study reports that host factors measuring the general immune function, markers of systemic inflammation, and tumor burden and location are the most prognostic for survival.
DNA methylation data-based precision tumor early diagnostics is emerging as state of the art technology, which could capture the signals of cancer occurrence 3∼5 years in advance and clinically more homogenous groups. At present, the sensitivity of early detection for many tumors is about 30%, which needs to be significantly improved. Nevertheless, based on the genome wide DNA methylation information, one could comprehensively characterize the entire molecular genetic landscape of the tumors and subtle differences among various tumors. With the accumulation of DNA methylation data, we need to develop high-performance methods that can model and consider more unbiased information. According to the above analysis, we have designed a self-attention graph convolutional network to automatically learn key methylation sites in a data-driven way for precision multi-tumor early diagnostics. Based on the selected methylation sites, we further trained a multi-class classification support vector machine. Large amount experiments have been conducted to investigate the performance of the computational pipeline. Experimental results demonstrated the effectiveness of the selected key methylation sites which are highly relevant for blood diagnosis.
Greetings and welcome to the fifth issue of IEEE Transactions on Computational Social Systems (TCSS) for 2023. This edition presents a collection of 55 diverse regular articles that illuminate various facets of the interaction between computer technology and society.
Background: It's critical to identify COVID-19 patients with a higher death risk at early stage to give them better hospitalization or intensive care. However, thus far, none of the machine learning models has been shown to be successful in an independent cohort. We aim to develop a machine learning model which could accurately predict death risk of COVID-19 patients at an early stage in other independent cohorts. Methods: We used a cohort containing 4711 patients whose clinical features associated with patient physiological conditions or lab test data associated with inflammation, hepatorenal function, cardiovascular function and so on to identify key features. To do so, we first developed a novel data preprocessing approach to clean up clinical features and then developed an ensemble machine learning method to identify key features. Results: Finally, we identified 14 key clinical features whose combination reached a good predictive performance of AUC 0.907. Most importantly, we successfully validated these key features in a large independent cohort containing 15,790 patients. Conclusions: Our study shows that 14 key features are robust and useful in predicting the risk of death in patients confirmed SARS-CoV-2 infection at an early stage, and potentially useful in clinical settings to help in making clinical decisions.
Background: Accurate target identification of small molecules and downstream target annotation are important in pharmaceutical research and drug development.Methods: We present TAIGET, a friendly and easy to operate graphical web interface, which consists of a docking module based on AutoDock Vina and LeDock, a target screen module based on a Bayesian–Gaussian mixture model (BGMM), and a target annotation module derived from >14,000 cancer-related literature works.Results: TAIGET produces binding poses by selecting ≤5 proteins at a time from the UniProt ID-PDB network and submitting ≤3 ligands at a time with the SMILES format. Once the identification process of binding poses is complete, TAIGET then screens potential targets based on the BGMM. In addition, three medical experts and 10 medical students curated associations among drugs, genes, gene regulation, cancer outcome phenotype, 2,170 cancer cell types, and 73 cancer types from the PubMed literature, with the aim to construct a target annotation module. A target-related PPI network can be visualized by an interactive interface.Conclusion: This online tool significantly lowers the entry barrier of virtual identification of targets for users who are not experts in the technical aspects of virtual drug discovery. The web server is available free of charge at http://www.taiget.cn/.
ABSTRACT Single-nucleotide polymorphism (SNPs) may cause the diverse functional impact on RNA or protein changing genotype and phenotype, which may lead to common or complex diseases like cancers. Accurate prediction of the functional impact of SNPs is crucial to discover the ‘influential’ (deleterious, pathogenic, disease-causing, and predisposing) variants from massive background polymorphisms in the human genome. Increasing computational methods have been developed to predict the functional impact of variants. However, predictive performances of these computational methods on massive genomic variants are still unclear. In this regard, we systematically evaluated 14 important computational methods including specific methods for one type of variant and general methods for multiple types of variants from several aspects; none of these methods achieved excellent (AUC ≥ 0.9) performance in both data sets. CADD and REVEL achieved excellent performance on multiple types of variants and missense variants, respectively. This comparison aims to assist researchers and clinicians to select appropriate methods or develop better predictive methods.
Fatty acids in crop seeds are a major source for both vegetable oils and industrial applications. Genetic improvement of fatty acid composition and oil content is critical to meet the current and future demands of plant-based renewable seed oils. Addressing this challenge can be approached by network modeling to capture key contributors of seed metabolism and to identify underpinning genetic targets for engineering the traits associated with seed oil composition and content. Here, we present a dynamic model, using an Ordinary Differential Equations model and integrated time-course gene expression data, to describe metabolic networks during Arabidopsis thaliana seed development. Through in silico perturbation of genes, targets were predicted in seed oil traits. Validation and supporting evidence were obtained for several of these predictions using published reports in the scientific literature. Furthermore, we investigated two predicted targets using omics datasets for both gene expression and metabolites from the seed embryo, and demonstrated the applicability of this network-based model. This work highlights that integration of dynamic gene expression atlases generates informative models which can be explored to dissect metabolic pathways and lead to the identification of causal genes associated with seed oil traits.
Continual reduction in sequencing cost and new generation sequencing (NGS) have expanded the accessibility of genome sequencing data for routine clinical applications. This has led to a new avenue in the cancer field as genomics and genetics testing are now being used daily in clinics. As such, germline variants are now being investigated for potential applications in the war against cancer. For many decades, similar to non-coding RNA, germline genetics was considered to be irrelevant for tumorigenesis and therefore, discarded. As of today, many studies have shown that the germline landscape has a major impact on cancer development and many other cancer-related phenotypes (i.e., drug resistance, recurrence, etc.). Combined with somatic information, germline genomes of patients could be incorporated in clinical decisions to allow for a better monitoring of cancer. Similarly, germline information could be used at diagnosis to inform on the likely course the disease will take as well as providing improvements in current treatment protocols. In this article, we will summarize the current germline cancer field and its origin. In addition, we will explore germline variants and their role in cancer diagnosis, prognosis, clinical and future applications.