Late-onset Alzheimer's disease (LOAD) is the most common cause of dementia in older adults, with specific genomic copy number variants (CNVs) implicated in its pathology. However, the aggregate burden of genome-wide CNVs in dementia and age-related neuropathologies is uncharacterized. This study investigated the association between genome-wide CNV scores (CNV-S) and dementia, as well as LOAD-related neuropathologies, in 1011 elderly participants (mean age 88.06) from two ongoing US-based longitudinal clinical-pathological cohort studies who were initially dementia-free and consented to brain donation upon death. Participants exhibited varying cognitive statuses at death (429 dementia, 258 mild cognitive impairment, 324 cognitively normal). We evaluated effects of (1) eight individual-level CNV-S based on gene loss intolerance and dosage sensitivity; (2) a single nucleotide polymorphism (SNP)-based LOAD polygenic score (PGS-LOAD) calculated using Bayesian continuous shrinkage; and (3) covariates (age, sex, and education). Outcomes included cognitive scores across 19 tests, clinical diagnoses of Alzheimer's disease or mild cognitive impairment, and four LOAD-related neuropathologies assessed postmortem. Analyses identified 4867 CNVs (3918 deletions, 949 duplications) mapped to 3211 genes. Higher deletion CNV-S were significantly associated with increased cerebrovascular pathologies (pLI: β = 0.14, 95% CI [0.08, 0.21]; LOEUF: β = 0.14, 95% CI [0.08, 0.20]; pHI: β = 0.15, 95% CI [0.08, 0.21]; binarized pHI: β = 0.14, 95% CI [0.08, 0.21]). Models predicting cerebral atherosclerosis that included deletion CNV-S significantly outperformed models based on only PGS-LOAD (R2 increase: 0.02). These findings suggest that genome-wide CNV burden, particularly deletions in dosage-sensitive genes, contributes to cerebrovascular pathology in aging. CNV-S may augment existing LOAD genetic risk models by capturing vascular pathways distinct from traditional SNP-based risk.
Background:Age is an independent prognostic factor in early-stage non-small cell lung cancer (NSCLC), yet the molecular differences between old and young patients and their contribution to disease progression remain unclear. We investigated age-related transcriptomic differences in early-stage lung adenocarcinoma (LUAD) and their association with recurrence. Methods:Tumor and adjacent normal lung tissue (NAT) from 126 stage I LUAD patients underwent bulk RNA sequencing to characterize age-related transcriptomic profiles. Differential expression and multiscale embedded gene co-expression network analysis (MEGENA) were used to identify age- and recurrence-associated modules. Pathways were annotated using Ingenuity Pathway Analysis. External confirmation was performed using TCGA (n=256) and TRACERx (n=83) cohorts. Results:Based on the cohort's median age, 60 patients were classified as old (>70 years) and 66 as young (≤70 years). In tumors, older patients with recurrence showed marked upregulation of cancer-associated, inflammatory, and extracellular matrix pathways compared with older patients without recurrence. In NAT samples, older patients with recurrence demonstrated upregulation of inflammatory and cancer-associated pathways-including phagosome formation, IL-17, IL-6, and Th2 signaling-that were absent or downregulated in young patients. MEGENA revealed a larger number of recurrence-associated co-expression modules in old versus young patients. These age-related patterns were highly conserved in both external cohorts across tumor and NAT samples. Conclusion:Aging in LUAD is associated with distinct cancer- and inflammation-related transcriptomic alterations that contribute to recurrence. Aging-related molecular signatures may improve risk stratification for early-stage lung cancer. Impact:Aging shapes tumor microenvironment transcriptomes in stage I LUAD, enabling improved relapse risk stratification after surgery.
Chronic pain represents a major health problem in the health care system. According to the CDC data brief in 2020, 20.4% of adults have chronic pain. There has been no promising therapy for chronic pain. Currently available treatments include medications such as nonsteroidal anti-inflammatory drugs, antiepileptic drugs, tricyclic antidepressants, corticosteroids, opioids, and cannabinoids, all of which may cause various negative side effects. Thus, there is an urgent need to develop novel, efficacious, and safe interventions for treating pain. Studies have shown that proinflammatory cytokines and chemokines make important contributions to the initiation and persistence of pain. We have found that C-C motif chemokine ligand 5 levels increased at day 14 post-spared nerve injury (SNI). This study was designed to investigate the effect of maraviroc (MVC), an FDA-approved CCR5 antagonist, on neuropathic pain in a mouse model of SNI. We found that MVC alleviated SNI-induced mechanical allodynia at 3, 7, and 14 days postinjury. MVC treatment also prevented SNI-mediated thermal hypersensitivity at 7 and 14 days postinjury in both male and female cohorts. SNI resulted in weight-bearing deficits, which were corrected by MVC administration in male mice. RNA sequencing analysis revealed that MVC rescued SNI-induced dysregulation of sex-specific canonical pathways in the spinal cord. Collectively, our findings showed that MVC could reduce neuropathic pain following peripheral nerve injury, providing a base for the repurposing of this FDA-approved human immunodeficiency virus drug as a pain reducer in clinical applications. SIGNIFICANCE STATEMENT: Spared nerve injury-induced neuropathic pain is associated with upregulation of the C-C motif chemokine ligand 5. Targeting the C-C motif chemokine ligand 5-CCR5 axis with FDA-approved maraviroc alleviated pain phenotype through modulating different pathways in male and female mice.
Background Blood-based biomarkers offer a promising non-invasive strategy for detecting disease-related changes and monitoring tissue and organ health, including brain function. While recent studies have leveraged blood transcriptomic data to predict gene expression in the brain, existing models generally suffer from poor accuracy, limiting their translational utility.Findings We present an integrative prediction system (IPS) that combines machine learning with network biology to predict region-specific brain gene expression from blood transcriptomic data. Our framework integrates global blood transcriptomic signals, co-expression network features, and inter-tissue gene-gene interaction data linking blood genes to their target genes in the brain. Applied to the Genotype-Tissue Expression cohort, IPS substantially outperforms existing approaches in both the number and accuracy of brain genes that can be reliably predicted from blood. Notably, immune-related blood genes emerged as key contributors to model performance, underscoring the systematic interplay between peripheral immune signaling and central nervous system.Conclusions These findings highlight the potential of blood-based transcriptomic models as scalable, non-invasive tools for studying brain function and developing diagnostic and prognostic biomarkers for neurological and psychiatric disorders.
Mild cognitive impairment (MCI) represents an initial phase of memory or other cognitive function decline and is viewed as an intermediary stage between normal aging and Alzheimer’s disease (AD), the most prevalent type of dementia. Individuals with MCI face a heightened risk of progressing to AD, and early detection of MCI can facilitate the prevention of such progression through timely interventions. Nonetheless, diagnosing MCI is challenging because its symptoms can be subtle and are easily missed. Using genomic data from blood samples has been proposed as a non-invasive and cost-efficient approach to build machine learning predictive models for assisting MCI diagnosis. However, these models often exhibit poor performance. In this study, we developed an XGBoost-based machine learning model with AUC (the Area Under the receiver operating characteristic Curve) of 0.9398 utilizing gene expression and copy number variation (CNV) data from patient blood samples. We demonstrated, for the first time, that data at a genome structure level such as CNVs could be as informative as gene expression data to classify MCI patients from normal controls. We identified 149 genomic features that are important for MCI prediction. Notably, these features are enriched in the pathways associated with neurodegenerative diseases, such as neuron development and G protein-coupled receptor activity. Overall, our study not only demonstrates the effectiveness of utilizing blood sample-based multi-omics for predicting MCI, but also provides insights into crucial molecular characteristics of MCI.
Understanding the molecular mechanisms underpinning diverse vaccination responses is critical for developing efficient vaccines. Molecular subtyping can offer insights into heterogeneous nature of responses and aid in vaccine design. We analyzed multi-omic data from 62 haemagglutinin seasonal influenza vaccine recipients (2019-2020), including transcriptomics, proteomics, glycomics, and metabolomics data collected pre-vaccination. We performed a subtyping analysis on the integrated data revealing five subtypes with distinct molecular signatures. These subtypes differed in the expression of pre-existing adaptive or innate immunity signatures, which were linked to significant variation in baseline immunoglobulin A (IgA) and hemagglutination inhibition (HAI) titer levels. It is worth noting that these differences persisted through day 28 post-vaccination, indicating the effect of initial immune state on vaccination response. These findings highlight the significance of interpersonal variation in baseline immune status as a crucial factor in determining the effectiveness of seasonal vaccines. Ultimately, incorporating molecular profiling could enable personalized vaccine optimization.
Understanding the molecular mechanisms that underpin diverse vaccination responses is a critical step toward developing efficient vaccines. Molecular subtyping approaches can offer valuable insights into the heterogeneous nature of responses and aid in the design of more effective vaccines. In order to explore the molecular signatures associated with the vaccine response, we analyzed baseline transcriptomics data from paired samples of whole blood, proteomics and glycomics data from serum, and metabolomics data from urine, obtained from influenza vaccine recipients (2019-2020 season) prior to vaccination. After integrating the data using a network-based model, we performed a subtyping analysis. The integration of multiple data modalities from 62 samples resulted in five baseline molecular subtypes with distinct molecular signatures. These baseline subtypes differed in the expression of pre-existing adaptive or innate immunity signatures, which were linked to significant variation across subtypes in baseline immunoglobulin A (IgA) and hemagglutination inhibition (HAI) titer levels. It is worth noting that these significant differences persisted through day 28 post-vaccination, indicating the effect of initial immune state on vaccination response. These findings highlight the significance of interpersonal variation in baseline immune status as a crucial factor in determining vaccine response and efficacy. Ultimately, incorporating molecular profiling could enable personalized vaccine optimization.
Inflammatory bowel disease (IBD) is a group of chronic digestive tract inflammatory conditions whose genetic etiology is still poorly understood. The incidence of IBD is particularly high among Ashkenazi Jews. Here, we identify 8 novel and plausible IBD-causing genes from the exomes of 4453 genetically identified Ashkenazi Jewish IBD cases (1734) and controls (2719). Various biological pathway analyses are performed, along with bulk and single-cell RNA sequencing, to demonstrate the likely physiological relatedness of the novel genes to IBD. Importantly, we demonstrate that the rare and high impact genetic architecture of Ashkenazi Jewish adult IBD displays significant overlap with very early onset-IBD genetics. Moreover, by performing biobank phenome-wide analyses, we find that IBD genes have pleiotropic effects that involve other immune responses. Finally, we show that polygenic risk score analyses based on genome-wide high impact variants have high power to predict IBD susceptibility.
Chronic stress induces changes in the periphery and the central nervous system (CNS) that contribute to neuropathology and behavioral abnormalities associated with psychiatric disorders. In this study, we examined the impact of peripheral and central inflammation during chronic social defeat stress (CSDS) in female mice. Compared to male mice, we found that female mice exhibited heightened peripheral inflammatory response and identified C-C motif chemokine ligand 5 (CCL5), as a stress-susceptibility marker in females. Blocking CCL5 signaling in the periphery promoted resilience to CSDS. In the brain, stress-susceptible mice displayed increased expression of C-C chemokine receptor 5 (CCR5), a receptor for CCL5, in microglia in the prefrontal cortex (PFC). This upregulation was associated with microglia morphological changes, their increased migration to the blood vessels, and enhanced phagocytosis of synaptic components and vascular material. These changes coincided with neurophysiological alterations and impaired blood-brain barrier (BBB) integrity. By blocking CCR5 signaling specifically in the PFC were able to prevent stress-induced physiological changes and rescue social avoidance behavior. Our findings are the first to demonstrate that stress-mediated dysregulation of the CCL5-CCR5 axis triggers excessive phagocytosis of synaptic materials and neurovascular components by microglia, resulting in disruptions in neurotransmission, reduced BBB integrity, and increased stress susceptibility. Our study provides new insights into the role of cortical microglia in female stress susceptibility and suggests that the CCL5-CCR5 axis may serve as a novel sex-specific therapeutic target for treating psychiatric disorders in females.
Gain-of-function (GOF) variants give rise to increased/novel protein functions whereas loss-of-function (LOF) variants lead to diminished protein function. Experimental approaches for identifying GOF and LOF are generally slow and costly, whilst available computational methods have not been optimized to discriminate between GOF and LOF variants. We have developed LoGoFunc, a machine learning method for predicting pathogenic GOF, pathogenic LOF, and neutral genetic variants, trained on a broad range of gene-, protein-, and variant-level features describing diverse biological characteristics. LoGoFunc outperforms other tools trained solely to predict pathogenicity for identifying pathogenic GOF and LOF variants and is available at https://itanlab.shinyapps.io/goflof/ .
Host genetic susceptibility is a key risk factor for severe illness associated with COVID-19. Despite numerous studies of COVID-19 host genetics, our knowledge of COVID-19-associated variants is still limited, and there is no resource comprising all the published variants and categorizing them based on their confidence level. Also, there are currently no computational tools available to predict novel COVID-19 severity variants. Therefore, we collated 820 host genetic variants reported to affect COVID-19 susceptibility by means of a systematic literature search and confidence evaluation, and obtained 196 high-confidence variants. We then developed the first machine learning classifier of severe COVID-19 variants to perform a genome-wide prediction of COVID-19 severity for 82,468,698 missense variants in the human genome. We further evaluated the classifier's predictions using feature importance analyses to investigate the biological properties of COVID-19 susceptibility variants, which identified conservation scores as the most impactful predictive features. The results of enrichment analyses revealed that genes carrying high-confidence COVID-19 susceptibility variants shared pathways, networks, diseases and biological functions, with the immune system and infectious disease being the most significant categories. Additionally, we investigated the pleiotropic effects of COVID-19-associated variants using phenome-wide association studies (PheWAS) in ~40,000 BioMe BioBank genotyped individuals, revealing pre-existing conditions that could serve to increase the risk of severe COVID-19 such as chronic liver disease and thromboembolism. Lastly, we generated a web-based interface for exploring, downloading and submitting genetic variants associated with COVID-19 susceptibility for use in both research and clinical settings (https://itanlab.shinyapps.io/COVID19webpage/). Taken together, our work provides the most comprehensive COVID-19 host genetics knowledgebase to date for the known and predicted genetic determinants of severe COVID-19, a resource that should further contribute to our understanding of the biology underlying COVID-19 susceptibility and facilitate the identification of individuals at high risk for severe COVID-19.
Epilepsy (EP) and congenital heart disease (CHD) are two apparently unrelated diseases that nevertheless display substantial mutual comorbidity. Thus, while congenital heart defects are associated with an elevated risk of developing epilepsy, the incidence of epilepsy in CHD patients correlates with CHD severity. Although genetic determinants have been postulated to underlie the comorbidity of EP and CHD, the precise genetic etiology is unknown. We performed variant and gene association analyses on EP and CHD patients separately, using whole exomes of genetically identified Europeans from the UK Biobank and Mount Sinai BioMe Biobank. We prioritized biologically plausible candidate genes and investigated the enriched pathways and other identified comorbidities by biological proximity calculation, pathway analyses, and gene-level phenome-wide association studies. Our variant- and gene-level results point to the Voltage-Gated Calcium Channels (VGCC) pathway as being a unifying framework for EP and CHD comorbidity. Additionally, pathway-level analyses indicated that the functions of disease-associated genes partially overlap between the two disease entities. Finally, phenome-wide association analyses of prioritized candidate genes revealed that cerebral blood flow and ulcerative colitis constitute the two main traits associated with both EP and CHD.
The human genetic dissection of clinical phenotypes is complicated by genetic heterogeneity. Gene burden approaches that detect genetic signals in case-control studies are underpowered in genetically heterogeneous cohorts. We therefore developed a genome-wide computational method, network-based heterogeneity clustering (NHC), to detect physiological homogeneity in the midst of genetic heterogeneity. Simulation studies showed our method to be capable of systematically converging genes in biological proximity on the background biological interaction network, and capturing gene clusters harboring presumably deleterious variants, in an efficient and unbiased manner. We applied NHC to whole-exome sequencing data from a cohort of 122 individuals with herpes simplex encephalitis (HSE), including 13 individuals with previously published monogenic inborn errors of TLR3-dependent IFN-α/β immunity. The top gene cluster identified by our approach successfully detected and prioritized all causal variants of five TLR3 pathway genes in the 13 previously reported individuals. This approach also suggested candidate variants of three reported genes and four candidate genes from the same pathway in another ten previously unstudied individuals. TLR3 responsiveness was impaired in dermal fibroblasts from four of the five individuals tested, suggesting that the variants detected were causal for HSE. NHC is, therefore, an effective and unbiased approach for unraveling genetic heterogeneity by detecting physiological homogeneity.
Identifying whether a given genetic mutation results in a gene product with increased (gain-of-function; GOF) or diminished (loss-of-function; LOF) activity is an important step toward understanding disease mechanisms because they may result in markedly different clinical phenotypes. Here, we generated an extensive database of documented germline GOF and LOF pathogenic variants by employing natural language processing (NLP) on the available abstracts in the Human Gene Mutation Database. We then investigated various gene- and protein-level features of GOF and LOF variants and applied machine learning and statistical analyses to identify discriminative features. We found that GOF variants were enriched in essential genes, for autosomal-dominant inheritance, and in protein binding and interaction domains, whereas LOF variants were enriched in singleton genes, for protein-truncating variants, and in protein core regions. We developed a user-friendly web-based interface that enables the extraction of selected subsets from the GOF/LOF database by a broad set of annotated features and downloading of up-to-date versions. These results improve our understanding of how variants affect gene/protein function and may ultimately guide future treatment options.
Over the last decade next generation sequencing (NGS) has been extensively used to identify new pathogenic mutations and genes causing rare genetic diseases. The efficient analyses of NGS data is not trivial and requires a technically and biologically rigorous pipeline that addresses data quality control, accurate variant filtration to minimize false positives and false negatives, and prioritization of the remaining genes based on disease genomics and physiological knowledge. This review provides a pipeline including all these steps, describes popular software for each step of the analysis, and proposes a general framework for the identification of causal mutations and genes in individual patients of rare genetic diseases.
Background Congenital heart disease (CHD) affects 1% of live births and is the most common birth defect. Although the genetic contribution to the CHD has been long suspected, it has only been well established recently. De novo variants are estimated to contribute to approximately 8% of sporadic CHD. Methods CHD is genetically heterogeneous, making pathway enrichment analysis an effective approach to explore and statistically validate CHD-associated genes. In this study, we performed novel gene and pathway enrichment analyses of high-impact de novo variants in the recently published whole-exome sequencing (WES) data generated from a cohort of CHD 2645 parent-offspring trios to identify new CHD-causing candidate genes and mutations. We performed rigorous variant- and gene-level filtrations to identify potentially damaging variants, followed by enrichment analyses and gene prioritization. Results Our analyses revealed 23 novel genes that are likely to cause CHD, including HSP90AA1, ROCK2, IQGAP1, and CHD4, and sharing biological functions, pathways, molecular interactions, and properties with known CHD-causing genes. Conclusions Ultimately, these findings suggest novel genes that are likely to be contributing to CHD pathogenesis.
Denatured proteins ire mostly partially folded and compact proteins. A statistical analysis on thermodynamic properties is presented to describe and characterize denatured proteins. Conformational free energy, energy, entropy and heat capacity expressions are derived using the Rotational Isomeric States model of polymer theory. The state space and the probabilities of each state are comprised faun a coil database. Properties for the denatured state are obtained for a sample set of proteins taken from the Protein Data Bank. Thermodynamic expressions of denatured state are derived.
RNA molecules are composed of modular architectural units that define their unique structural and functional properties. Characterization of these building blocks can help interpret RNA structure/function relationships. We present an RNA secondary structure motif and submotif library using dual graph representation and partitioning. Dual graphs represent RNA helices as vertices and loops as edges. Unlike tree graphs, dual graphs can represent RNA pseudoknots (intertwined base pairs). For a representative set of RNA structures, we construct dual graphs from their secondary structures, and apply our partitioning algorithm to identify non-separable subgraphs (or blocks) without breaking pseudoknots. We report 56 subgraph blocks up to nine vertices; among them, 22 are frequently occurring, 15 of which contain pseudoknots. We then catalog atomic fragments corresponding to the subgraph blocks to define a library of building blocks that can be used for RNA design, which we call RAG-3Dual, as we have done for tree graphs. As an application, we analyze the distribution of these subgraph blocks within ribosomal RNAs of various prokaryotic and eukaryotic species to identify common subgraphs and possible ancestry relationships. Other applications of dual graph partitioning and motif library can be envisioned for RNA structure analysis and design.