Background: Polygenic risk scores (PRS) have proven valuable for disease risk prediction, but their predictive utility often remains limited because human traits result from complex interactions between environmental and genetic factors. Blood lipid levels are heritable and clinically important risk factors for cardiovascular disease, yet it remains unclear whether multi-omics integration can enhance lipid trait prediction beyond PRS alone. Methods: We first constructed single-omics scores, where gene expression, plasma protein, and plasma/serum metabolite levels were genetically predicted and weighted by effect sizes estimated via LASSO regression. Subsequently, we implemented two integration strategies to develop composite multi-omics risk scores (MoRS): step-MoRS, which integrates single-omics scores using stepwise regression, and Lasso-MoRS, which directly models all predicted features across omics layers using LASSO regression. Both approaches were evaluated across European, South Asian, and African ancestries within the UK Biobank. Results: MoRS-based methods consistently demonstrated superior predictive accuracy compared to PRS alone for four lipid traits across diverse populations. Notably, Lasso-MoRS prioritized key biomarkers with predictive utility complementary to genomic data. Conclusions: These findings confirm that integrating multi-omics biomarkers with genomic data significantly enhances lipid trait prediction across diverse ancestries, offering biological insights into the molecular regulation of lipid metabolism.
Breast cancer (BC) is one of the most common cancer types among women worldwide. Understanding the complex molecular and cellular characteristics of BC is crucial for advancing precision treatment. To enable more reliable and reproducible biological discoveries, it is critical to collect molecular data from diverse BC cohorts and establish an integrative, versatile analysis platform. Here, we present BCMA (Breast Cancer Molecular Atlas, http://lifeome.net/database/bcma/), a multi-scale, multi-omics BC database that encompasses 6 bulk multi-omics datasets and 9 single-cell transcriptomics datasets, collectively covering 5424 cases and 236,363 cells. The BCMA systemically characterizes the molecular features of BC, including gene mutations, copy number alterations, RNA expression, miRNA expression, DNA methylation, as well as clinical phenotypes and cell heterogeneity. Meanwhile, a user-friendly interface for gene-centered search is provided, achieving the clinical information statistics, genomic events analysis, differential multi-omics feature identification, functional enrichment analysis, survival analysis, co-expression analysis, as well as single-cell gene expression profiling and cell type annotation. This platform holds great potential to enhance the understanding of molecular characteristics underlying BC and to facilitate the identification of disease-associated biomarkers.
Single-cell resolution spatial transcriptomics (ST) provides a great opportunity to explore the complex cellular contents in different tissues. In this study, we introduce SpatialZoomer, a spectral graph-based method that applies a set of low-pass filters to efficiently extract spatial molecular features from ST data at multiple resolutions or scales, including the single cell scale, the niche scale with dozens of closely interacting cells, and the domain scale with spatially-organized cell contents. The corresponding "critical" scales can be automatically identified by partitioning a cross-scale similarity map. Results show that SpatialZoomer can identify disease-progression signals at specific scales in Alzheimer's Disease, and spatial context dependent cell subtypes in tumor microenvironment. The extracted multi-scale features also uncovered spatially heterogeneous niches in cancer. SpatialZoomer also takes advantage of high computational efficiency and low hardware requirements.
BACKGROUND:Glioblastoma (GBM) is the most malignant and highly recurrent brain tumor. Although over half of the GBM patients are elderly patients, the understanding of how aging affects GBM progression remains limited. METHODS:Clinical and genomic variation data of GBM patients from TCGA and CGGA databases were used for prognostic analysis. We collected single-cell transcriptome data of 88,908 cells from 13 primary GBM (pGBM) and 12 recurrent GBM (rGBM) patients. Age-related immune cells were identified through cell-cell communication and trajectory analysis. The results were validated by projecting the single-cell transcriptome profiles onto bulk data. Finally, we experimentally validated the results on syngeneic orthotopic models of younger and older mice. RESULTS:Prognostic analysis indicated the effect of age in pGBM patients is stronger than that in rGBM patients. Moreover, the mutational signatures in pGBM does not affect prognosis. Single-cell RNA sequencing analysis revealed age-related differences in immune cells between pGBM and rGBM, and identified microglia underwent significant cell state changes with aging only in pGBM patients. Next, we validated that high expression of HSPB1 in microglia from older pGBM patients is associated with poor prognosis. Finally, the syngeneic orthotopic model for aged mice exhibited more tumor invasion, with a shorter median survival time. Furthermore, the microglia within the tumor microenvironment (TME) of aged mice showed markedly high expression level of HSPB1. CONCLUSIONS:Our study highlights the crucial role of microglial aging in pGBM, reveals distinct age-related changes of immune cells in the TME between pGBM and rGBM, and offers valuable insights into clinical treatment strategies targeting elderly GBM patients.
Spatial transcriptome provides critical insights for inferring cell-cell communications. Based on the heuristic that physically adjacent cells are more likely to communicate with each other, we develop a computational framework that Infers the Cell-Cell Communication network (IC3) and subsequently identifies communication hotspots with rigorous error control. We demonstrate in simulations and real datasets that our method outperforms existing methods in accuracy and especially improves sensitivity to identify communications involving rare cell types. Applying IC3 in a mouse brain dataset, we recovered how cells communicate with each other to form a local structure that regulates the balance of the blood-brain barrier. These findings highlight that our method is effective in revealing tissue architecture and function from cellular communication networks.
Recent single-cell CRISPR screening experiments have combined the advances of genetic editing and single-cell technologies, leading to transcriptome-scale readouts of responses to perturbations at single-cell resolution. An outstanding question is how to efficiently identify heterogeneous causal effects of perturbations using these technologies. Here we present scCAPE, a tool designed to facilitate causal analysis of heterogeneous perturbation effects at the single-cell level. scCAPE disentangles perturbation effects from the inherent cell-state variations and provides nonparametric inferences of perturbation effects at single-cell resolution, permitting a range of downstream tasks including perturbation effect analysis, genetic interaction analysis, perturbation clustering and prioritizing. We benchmarked scCAPE through simulation studies and real datasets to evaluate its performance in characterizing latent confounding factors and accuracy in estimating heterogeneous perturbation effects. The application of scCAPE identified novel heterogeneous genetic interactions among erythroid differentiation drivers. For example, our analysis pinpointed the role of the synergistic interaction between CBL and CNN1 in the S phase. ### Competing Interest Statement The authors have declared no competing interest.
OBJECTIVES:To explore the risk factors associated with cow's milk protein allergy (CMPA) in infants.METHODS:This study was a multicenter prospective nested case-control study conducted in seven medical centers in Beijing, China. Infants aged 0-12 months were included, with 200 cases of CMPA infants and 799 control infants without CMPA. Univariate and multivariate logistic regression analyses were used to investigate the risk factors for the occurrence of CMPA.RESULTS:Univariate logistic regression analysis showed that preterm birth, low birth weight, birth from the first pregnancy, firstborn, spring birth, summer birth, mixed/artificial feeding, and parental history of allergic diseases were associated with an increased risk of CMPA in infants (P<0.05). Multivariate logistic regression analysis revealed that firstborn (OR=1.89, 95%CI: 1.14-3.13), spring birth (OR=3.42, 95%CI: 1.70-6.58), summer birth (OR=2.29, 95%CI: 1.22-4.27), mixed/artificial feeding (OR=1.57, 95%CI: 1.10-2.26), parental history of allergies (OR=2.13, 95%CI: 1.51-3.02), and both parents having allergies (OR=3.15, 95%CI: 1.78-5.56) were risk factors for CMPA in infants (P<0.05).CONCLUSIONS:Firstborn, spring birth, summer birth, mixed/artificial feeding, and a family history of allergies are associated with an increased risk of CMPA in infants.
Acknowledging the significance of information propagation and individual adaptive behavior has been regarded as an indispensable prerequisite for a complete understanding of epidemic spreading. Recent studies have widely considered the metapopulation model, where epidemics spread over a single layer of physical networks via individual mobility. However, these advances neglected the interventions of accompanied information and individual behavior response related to epidemics. In this article, we develop a coupled epidemic-information propagation model on multiplex metapopulation networks leveraging the microscopic Markov chain (MMC) approach, aiming to explore the spatiotemporal characteristics of epidemic spreading process. Taking the individual adaptive behavior into account, the stranding mechanism based on infection level and medical resources is introduced to capture the population size dynamics during individual mobility among different patches. Theoretical epidemic threshold is analytically derived under the improved framework. Extensive numerical simulations are performed to validate our theoretical analysis and further examine the impacts of information propagation and spreading parameters on epidemic threshold and steady-state prevalence. Our results indicate that both the scale of information diffusion and the specific configuration of spreading parameters can significantly suppress the epidemic prevalence. These findings shed a novel light on theoretical research and decision-making of coupled epidemic-information process in the spatiotemporal perspective.
TWAS have shown great promise in extending GWAS loci to a functional understanding of disease mechanisms. In an effort to fully unleash the TWAS and GWAS information, we propose MTWAS, a statistical framework that partitions and aggregates cross-tissue and tissue-specific genetic effects in identifying gene-trait associations. We introduce a non-parametric imputation strategy to augment the inaccessible tissues, accommodating complex interactions and non-linear expression data structures across various tissues. We further classify eQTLs into cross-tissue eQTLs and tissue-specific eQTLs via a stepwise procedure based on the extended Bayesian information criterion, which is consistent under high-dimensional settings. We show that MTWAS significantly improves the prediction accuracy across all 47 tissues of the GTEx dataset, compared with other single-tissue and multi-tissue methods, such as PrediXcan, TIGAR, and UTMOST. Applying MTWAS to the DICE and OneK1K datasets with bulk and single-cell RNA sequencing data on immune cell types showcases consistent improvements in prediction accuracy. MTWAS also identifies more predictable genes, and the improvement can be replicated with independent studies. We apply MTWAS to 84 UK Biobank GWAS studies, which provides insights into disease etiology.
Modeling the global dynamics of emerging infectious diseases (EIDs) like COVID-19 can provide important guidance in the preparation and mitigation of pandemic threats. While age-structured transmission models are widely used to simulate the evolution of EIDs, most of these studies focus on the analysis of specific countries and fail to characterize the spatial spread of EIDs across the world. Here, we developed a global pandemic simulator that integrates age-structured disease transmission models across 3,157 cities and explored its usage under several scenarios. We found that without mitigations, EIDs like COVID-19 are highly likely to cause profound global impacts. For pandemics seeded in most cities, the impacts are equally severe by the end of the first year. The result highlights the urgent need for strengthening global infectious disease monitoring capacity to provide early warnings of future outbreaks. Additionally, we found that the global mitigation efforts could be easily hampered if developed countries or countries near the seed origin take no control. The result indicates that successful pandemic mitigations require collective efforts across countries. The role of developed countries is vitally important as their passive responses may significantly impact other countries.
Single-cell chromatin accessibility sequencing (scCAS) technologies have enabled characterizing the epigenomic heterogeneity of individual cells. However, the identification of features of scCAS data that are relevant to underlying biological processes remains a significant gap. Here, we introduce a novel method Cofea, to fill this gap. Through comprehensive experiments on 5 simulated and 54 real datasets, Cofea demonstrates its superiority in capturing cellular heterogeneity and facilitating downstream analysis. Applying this method to identification of cell type-specific peaks and candidate enhancers, as well as pathway enrichment analysis and partitioned heritability analysis, we illustrate the potential of Cofea to uncover functional biological process.
Single-cell sequencing technologies, adopted extensively over the past decade, generate voluminous data that profile a myriad of cellular features. To harness the full potential of single-cell sequencing, it is crucial to develop powerful, efficient, and robust computational and statistical methods for analyzing the data and extracting meaningful insights. This Research Topic showcases five seminal papers that introduce novel analytical techniques and benchmark existing methods in this domain. One of the key areas of focus in recent years has been the analysis of single-cell ATAC sequencing (scATAC-seq) data. ScATAC-seq data measure chromatin accessibility in individual cells. Sophisticated statistical and computational techniques are required to analyze scATAC-seq data, specifically for clustering and trajectory reconstruction analysis, a common requirement for scRNA-seq data. The paper titled “Destin2: Integrative and Cross-Modality Analysis of Single-Cell Chromatin Accessibility Data” introduces a novel method designed for cross-modality dimension reduction, clustering, and trajectory reconstruction of single-cell ATAC-seq data. By integrating cellular-level epigenomic profiles from peak accessibility, motif deviation score, and pseudo-gene activity, Destin2 infers a sharedmanifold usingmultimodal input, followed by clustering or trajectory inference. The authors demonstrate the effectiveness of Destin2 through its application to experimental scATAC-seq datasets with both discretized cell types and transient cell states and by benchmarking against existing methods based on unimodal analyses. The findings illustrate that Destin2 corroborates and enhances existing techniques, establishing it as a valuable computational pipeline for scATAC-seq data analysis. Another noteworthy contribution in scATAC-seq analysis is presented in the paper “Benchmarking Automated Cell Type Annotation Tools for Single-Cell ATAC-seq Data.” This study evaluates the performance of five annotation methods for identifying cell types in scATAC-seq data. By assessing classification accuracy and scalability, the authors provide valuable guidance for selecting appropriate tools for cell type annotation. Using publicly available single-cell datasets from mouse and human tissues, including brain, lung, kidney, PBMC, and BMMC, the authors found Bridge integration as the most effective and robust method, impervious to alterations in data size, mislabeling rate, and sequencing depth.While OPEN ACCESS
Polygenic risk scores (PRS) calculated from genome-wide association studies (GWAS) of Europeans are known to have substantially reduced predictive accuracy in non-European populations, limiting their clinical utility and raising concerns about health disparities across ancestral populations. Here, we introduce a statistical framework named X-Wing to improve predictive performance in ancestrally diverse populations. X-Wing quantifies local genetic correlations for complex traits between populations, employs an annotation-dependent estimation procedure to amplify correlated genetic effects between populations, and combines multiple population-specific PRS into a unified score with GWAS summary statistics alone as input. Through extensive benchmarking, we demonstrate that X-Wing pinpoints portable genetic effects and substantially improves PRS performance in non-European populations, showing 14.1%–119.1% relative gain in predictive R 2 compared to state-of-the-art methods based on GWAS summary statistics. Overall, X-Wing addresses critical limitations in existing approaches and may have broad applications in cross-population polygenic risk prediction.
Heritability is a fundamental concept in genetic studies, measuring the genetic contribution to complex traits and bringing insights about disease mechanisms. The advance of high-throughput technologies has provided many resources for heritability estimation. Linkage disequilibrium (LD) score regression (LDSC) estimates both heritability and confounding biases, such as cryptic relatedness and population stratification, among single-nucleotide polymorphisms (SNPs) by using only summary statistics released from genome-wide association studies. However, only partial information in the LD matrix is utilized in LDSC, leading to loss in precision. In this study, we propose LD eigenvalue regression (LDER), an extension of LDSC, by making full use of the LD information. Compared to state-of-the-art heritability estimating methods, LDER provides more accurate estimates of SNP heritability and better distinguishes the inflation caused by polygenicity and confounding effects. We demonstrate the advantages of LDER both theoretically and with extensive simulations. We applied LDER to 814 complex traits from UK Biobank, and LDER identified 363 significantly heritable phenotypes, among which 97 were not identified by LDSC.
Recent advances in single-cell technologies have enabled the characterization of epigenomic heterogeneity at the cellular level. Computational methods for automatic cell type annotation are urgently needed given the exponential growth in the number of cells. In particular, annotation of single-cell chromatin accessibility sequencing (scCAS) data, which can capture the chromatin regulatory landscape that governs transcription in each cell type, has not been fully investigated. Here we propose EpiAnno, a probabilistic generative model integrated with a Bayesian neural network, to annotate scCAS data automatically in a supervised manner. We systematically validate the superior performance of EpiAnno for both intra- and inter-dataset annotation on various datasets. We further demonstrate the advantages of EpiAnno for interpretable embedding and biological implications via expression enrichment analysis, partitioned heritability analysis, enhancer identification, cis-coaccessibility analysis and pathway enrichment analysis. In addition, we show that EpiAnno has the potential to reveal cell type-specific motifs and facilitate scCAS data simulation. The investigation of single-cell epigenomics with technologies such as single-cell chromatin accessibility sequencing (scCAS) presents an opportunity to expand the understanding of gene regulation at the cellular level. The authors develop a probabilistic generative model to better characterize cell heterogeneity and accurately annotate the cell type of scCAS data.
Openness-weighted association study (OWAS) is a method that leverages the in silico prediction of chromatin accessibility to prioritize genome-wide association studies (GWAS) signals, and can provide novel insights into the roles of non-coding variants in complex diseases. A prerequisite to apply OWAS is to choose a trait-related cell type beforehand. However, for most complex traits, the trait-relevant cell types remain elusive. In addition, many complex traits involve multiple related cell types. To address these issues, we develop OWAS-joint, an efficient framework that aggregates predicted chromatin accessibility across multiple cell types, to prioritize disease-associated genomic segments. In simulation studies, we demonstrate that OWAS-joint achieves a greater statistical power compared to OWAS. Moreover, the heritability explained by OWAS-joint segments is higher than or comparable to OWAS segments. OWAS-joint segments also have high replication rates in independent replication cohorts. Applying the method to six complex human traits, we demonstrate the advantages of OWAS-joint over a single-cell-type OWAS approach. We highlight that OWAS-joint enhances the biological interpretation of disease mechanisms, especially for non-coding regions.
Abstract Motivation The interactions among various types of cells play critical roles in cell functions and the maintenance of the entire organism. While cell–cell interactions are traditionally revealed from experimental studies, recent developments in single-cell technologies combined with data mining methods have enabled computational prediction of cell–cell interactions, which have broadened our understanding of how cells work together, and have important implications in therapeutic interventions targeting cell–cell interactions for cancers and other diseases. Despite the importance, to our knowledge, there is no database for systematic documentation of high-quality cell–cell interactions at the cell type level, which hinders the development of computational approaches to identify cell–cell interactions. Results We develop a publicly accessible database, CITEdb (Cell–cell InTEraction database, https://citedb.cn/), which not only facilitates interactive exploration of cell–cell interactions in specific physiological contexts (e.g. a disease or an organ) but also provides a benchmark dataset to interpret and evaluate computationally derived cell–cell interactions from different tools. CITEdb contains 728 pairs of cell–cell interactions in human that are manually curated. Each interaction is equipped with structured annotations including the physiological context, the ligand–receptor pairs that mediate the interaction, etc. Our database provides a web interface to search, visualize and download cell–cell interactions. Users can search for cell–cell interactions by selecting the physiological context of interest or specific cell types involved. CITEdb is the first attempt to catalogue cell–cell interactions at the cell type level, which is beneficial to both experimental, computational and clinical studies of cell–cell interactions. Availability and implementation CITEdb is freely available at https://citedb.cn/ and the R package implementing benchmark is available at https://github.com/shanny01/benchmark. Supplementary information Supplementary data are available at Bioinformatics online.
BACKGROUND AND OBJECTIVES Cow's milk allergy (CMA) is the most common food allergy in young children. Previous studies have reported that single-nucleotide polymorphisms (SNPs) are associated with CMA. The extent to which SNPs contribute to the occurrence of CMA is unknown. The purpose of this study was to investigate the independent relevance of genetic predisposition to CMA in Chinese children. METHODS AND STUDY DESIGN 200 infants with CMA and 799 healthy controls aged 0-12 months were included. Five previously identified genetic variants (rs17616434, rs2069772, rs1800896, rs855791 and rs20541) were genotyped. Logistic regression was used to analyze the genetic associations or their interactions with a family history of allergy on CMA. RESULTS Among the five SNPs, only IL10 rs1800896 was significantly associated with CMA (odds ratio (OR) 1.60, p=0.042). Each 1-risk allele increase in the genetic risk score (GRS) was suggestively associated with an 11% higher risk of CMA (1.11: 0.99-1.27, p=0.069) and a 45% increased risk of CMA in the GRS high-risk group compared to the GRS low-risk group (1.45: 1.02-2.06, p=0.037). Furthermore, parental allergy also increased the risk of CMA among children (1.87: 1.46-2.39, p<0.001). Importantly, parental allergy exacerbated the genetic effect on the risk of CMA. CONCLUSIONS The rs1800896 variant in the IL-10 gene is associated with CMA in Chinese children. In addition, the GRS had an interaction with parental history of allergy, implying that genetic risk for CMA was exacerbated among those with parental history of allergy.
Exome sequencing on tens of thousands of parent-proband trios has identified numerous deleterious de novo mutations (DNMs) and implicated risk genes for many disorders. Recent studies have suggested shared genes and pathways are enriched for DNMs across multiple disorders. However, existing analytic strategies only focus on genes that reach statistical significance for multiple disorders and require large trio samples in each study. As a result, these methods are not able to characterize the full landscape of genetic sharing due to polygenicity and incomplete penetrance. In this work, we introduce EncoreDNM, a novel statistical framework to quantify shared genetic effects between two disorders characterized by concordant enrichment of DNMs in the exome. EncoreDNM makes use of exome-wide, summary-level DNM data, including genes that do not reach statistical significance in single-disorder analysis, to evaluate the overall and annotation-partitioned genetic sharing between two disorders. Applying EncoreDNM to DNM data of nine disorders, we identified abundant pairwise enrichment correlations, especially in genes intolerant to pathogenic mutations and genes highly expressed in fetal tissues. These results suggest that EncoreDNM improves current analytic approaches and may have broad applications in DNM studies.