Disruption of protein-protein interactions (PPIs) is a major mechanism of a variant's deleterious effect. Computational tools are needed to assess such variants at scale, yet existing predictors rarely consider loss of specific interactions, particularly when variants perturb binding interfaces without significantly affecting protein stability. To address this problem, we present MutPred-PPI, a graph attention network that predicts interaction-specific (edgetic) effects of missense variants by operating on AlphaFold3-based protein complex contact graphs with protein language model embeddings imposed upon nodes. We systematically evaluated our model with stringent group cross-validation as well as benchmark data recently collected within the IGVF Consortium. MutPred-PPI outperformed all baseline methods across all evaluation criteria, achieving an AUC of 0.85 on seen proteins and 0.72 on previously unseen proteins in cross-validation, demonstrating strong generalizability despite scarce training data. To demonstrate biomedical relevance, we applied MutPred-PPI to variants from ClinVar, HGMD, COSMIC, gnomAD, and two de novo neurodevelopmental disorder-linked datasets. Disease-associated variants from ClinVar and HGMD showed strong enrichment for both quasi-null and edgetic effects, whereas population variants from gnomAD increasingly preserved interactions with higher allele frequencies. Notably, we observed a strong edgetic disruption signature in highly recurrent cancer variants from both the full COSMIC dataset and a subset of variants from oncogenes. Recurrent tumor suppressor gene variants and autism spectrum disorder-associated variants exhibited moderate quasi-null enrichment, whilst neurodevelopmental disorder-linked variants showed a weak edgetic disruption signature. These results indicate distinct PPI perturbation mechanisms across disease types and show that MutPred-PPI captures functionally relevant molecular effects of pathogenic variants.
Predicting the effects of genetic variants and assessing prediction performance are key computational tasks in genomic medicine. It has been shown that well-calibrated variant effect predictors can be reliably used as evidence towards establishing pathogenicity (or benignity) of missense variants, thereby rendering these variants suitable for use in (or exclusion from) the genetic diagnosis of rare Mendelian conditions. However, most predictors have been trained or calibrated on data that may not be sufficiently representative to lead to similar performance across all genetic ancestries. This raises questions about the responsible deployment of these tools to improve human health. To better understand the utility of computational predictors, we set out to assess their ancestry-specific performance in terms of accuracy and evidence strength according to the ACMG/AMP guidelines. First, we determined that the expected count of rare variants in an individual's genome and the allele frequency distribution of these variants are the key confounders when evaluating a predictor's performance across different genetic ancestries. Second, we found that a predictor's accuracy itself inversely correlates with the allele frequency of the rare variant. After stratifying according to allele frequency, we show that established methods for predicting the pathogenicity of missense variants have comparable performance levels across major ancestry groups. Our results therefore support the wide deployment of such models in the context of genetic diagnosis and related applications.
New World monkeys (Platyrrhini), a highly diverse primate lineage endemic to Mexico, Central and South America, include howler monkeys (Alouatta spp.), one of the most ecologically successful genera in the region. They have evolved key adaptations including a specialized folivorous diet and distinctive howling calls that carry over long distances in tropical rainforests, conferring critical survival advantages. However, the genetic underpinnings of these adaptations—specifically, the genes involved in fiber digestion, fermentation, and metabolic adaptation for folivory, and those regulating vocalization for howling—remain poorly understood. Elucidating these mechanisms is important for understanding howler monkey adaptive evolution. To elucidate the genetic underpinnings of these adaptive traits, we generate a high-quality genome of the Guyanan red howler monkey (Alouatta macconnelli) through high-fidelity long-read sequencing. Leveraging this assembly, we conduct extensive population genomic analyses to reconstruct the phylogeny and historical introgression events across Alouatta lineages. Furthermore, we systematically identify genetic changes within protein-coding genes and regulatory regions that underlie the adaptive evolution of digestive shifts, changes in energy metabolism, and the enlarged hyoid bone in howler monkeys. Notably, we identify a lineage-specific duplication of FBP1 (fructose-1,6-bisphosphatase 1). Experiments show that subfunctionalized paralogs have lost catalytic activity but retain AMP binding, which reduces AMP-mediated FBPase inhibition, consequently augmenting gluconeogenic capacity. This innovation provides sustained gluconeogenic capacity utilizing short-chain fatty acids from leaf fermentation, securing metabolic homeostasis in howler monkeys. Together, our findings reveal novel mechanisms underlying the adaptive evolution of folivory and howling in howler monkeys, and provide the first genomic reconstruction of their phylogeographic history.
Variants with intermediate functional effects-neither fully disruptive nor functionally neutral-represent an under-recognized source of genetic complexity and define a functional gray zone that complicates variant classification. Here, we address this issue using GT>GC (+2T>C) 5' splice-site variants as a tractable model, as approximately 15%-18% of such substitutions retain variable levels of residual wild-type (WT) transcript. Using residual WT transcript as a quantitative functional readout, we first revisited disease-associated GT>GC variants previously shown to retain substantial WT transcript, including SPINK1 c.194+2T>C, HBB c.315+2T>C, and BRCA2 c.8331+2T>C, illustrating how intermediate splicing effects complicate clinical interpretation across distinct genes and disease contexts. We then performed a locus-wide assessment of all 26 theoretically possible GT>GC substitutions in CFTR, integrating SpliceAI delta donor-loss scores with classifications from expert-curated databases. Minigene splicing analyses of four selected CFTR variants, together with full-length and minigene analyses of a BAP1 GT>GC variant with conflicting clinical interpretations, revealed heterogeneous and context-dependent splicing outcomes, underscoring both inter-assay variability and the inherent limitations of commonly used splicing assay systems. Collectively, our findings indicate that GT>GC variants capable of generating appreciable residual WT transcript exemplify a broader class of intermediate-effect alleles that expose the limitations of both computational prediction and experimental assessment. These observations highlight the need for classification frameworks that incorporate quantitative functional data and better capture the continuum of variant effects.
Abstract As genomic sequencing evolves beyond rare disease diagnostics toward population screening and precision medicine, clinical variant interpretation is increasingly challenged by variants whose clinical consequences depend on biological context. Current frameworks, including the ACMG/AMP guidelines, generally assign a single classification to each variant regardless of inheritance state or genetic context, potentially failing to communicate context-dependent clinical consequences. Here, we address this issue using loss-of-function variants in LPL as a uniquely informative model system in which residual physiological LPL activity can be directly quantified in vivo . By systematically integrating published biallelic LPL genotypes, physiological measurements, functional studies, and clinical phenotypes, we identified a biologically meaningful transition at approximately 10% residual physiological LPL activity. Activity below this level was predominantly associated with classical childhood-onset familial chylomicronemia syndrome (FCS), whereas higher activity was associated with phenotypic attenuation and modifier-dependent clinical expression. Furthermore, heterozygous loss-of-function variants exhibited an estimated penetrance of 5–7% for severe hypertriglyceridemia. We therefore propose a context-dependent framework in which biallelic complete- or near-complete loss-of-function genotypes are interpreted as causative for FCS, whereas heterozygous variants are interpreted as predisposing to severe hypertriglyceridemia while retaining recognition of FCS carrier status. Together, our findings demonstrate that clinical variant interpretation should integrate available biological context—including, where relevant, allelic configuration, residual biological function, and penetrance—rather than rely on the intrinsic molecular consequence of the variant alone. More broadly, this framework provides a conceptual model for interpreting variants across the continuum from Mendelian disease to genetic predisposition in the era of precision medicine.
RareGPS is a machine-learning framework prioritizing drug targets for rare and uncommon diseases, integrating 11 genetic, clinical, and experimental evidence sources. It uses the full distribution of genetic associations across allele-frequency bins in an allelic-series model. Across 161 phenotypes, RareGPS outperforms existing resources for predicting drug indications and clinical trial progression; top 1% targets show 58-fold higher likelihood of advancing from nonindicated to phase IV and 8-fold from phase I to IV versus the middle 50%. We validated RareGPS using prescriptome analyses in two million patients and an independent literature evaluation tool (AMELIE). We publish predictions for 3,021,965 gene-phenotype pairs.
Purpose Classification of DNA sequence data requires the implementation of the American College of Medical Genetics and Genomics (ACMG) standards and guidelines. Therefore, automated tools have been developed. However, these tools often lack robust and up-to-date methodologies. This study reports on the development of a new tool and examines its performance for diagnostic and research purposes. Methods The automated ACMG-based variant classifier (AAVC) presented here computationally analyzes sequence variants following the ACMG guidelines, the Clinical Genome Resource specifications and a novel framework by leveraging large public databases and in silico prediction tools. Results AAVC demonstrated high concordance (94.39%) with the Food and Drug Administration recognized variant classifications, outperforming currently available tools. It classified 55% of the variants of uncertain significance in clinical variation into clinically significant categories. We identified, in the Turkish Variome, 215 novel pathogenic, likely pathogenic, or variants of uncertain significance high variants in the secondary finding genes and revealed that 1 in 10 individuals carried an actionable genotype. Conclusion AAVC constitutes a robust framework for the accurate classification of human germline sequence diversity is available at https://aavc.bilkent.edu.tr/, offering a highly accurate, rapid, and up-to-date platform for clinical laboratories and research groups to automatically interpret sequence variants.
The 5 ' untranslated region (5 ' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5 ' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5 ' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5 ' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5 ' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.
Congenital heart disease (CHD) is the most common congenital anomaly and a leading cause of infant morbidity and mortality. Despite extensive exploration of the monogenic causes of CHD over the last decades, ∼55% of cases still lack a molecular diagnosis. Investigating digenic interactions, the simplest form of oligogenic interactions, using high-throughput sequencing data can elucidate additional genetic factors contributing to the disease. Here, we conducted a comprehensive analysis of digenic interactions in CHD by utilizing a large CHD trio exome sequencing cohort, comprising 3,910 CHD and 3,644 control trios. We extracted pairs of presumably deleterious rare variants observed in CHD-affected and unaffected children but not in a single parent. Burden testing of gene pairs derived from these variant pairs revealed 29 nominally significant gene pairs. These gene pairs showed a significant enrichment for known CHD genes (p < 1.0 × 10-4) and exhibited a shorter average biological distance to known CHD genes than expected by chance (p = 3.0 × 10-4). Utilizing three complementary biological relatedness approaches including network analyses, biological distance calculations, and candidate gene prioritization methods, we prioritized 10 final gene pairs that are likely to underlie CHD. Analysis of bulk RNA-sequencing data showed that these genes are highly expressed in the developing embryonic heart (p < 1 × 10-4). In conclusion, our findings suggest the potential role of digenic interactions in CHD pathogenesis and provide insights into unresolved molecular diagnoses. We suggest that the application of the digenic approach to additional disease cohorts will significantly enhance genetic discovery rates.
Parkinson's disease (PD) is a devastating neurodegenerative disorder with growing prevalence worldwide and, as yet, no effective treatment. Drug repurposing is invaluable for detecting novel PD therapeutics. Here, we compiled gene expression data from 1231 healthy human brain samples and 357 samples across tissues, ethnicities, brain regions, Braak stages, and disease status. By integrating them with multiple-source genomic data, we found a PD-associated gene co-expression module, and its alignment with the CMAP database successfully identified drug candidates. Among these, meclofenoxate hydrochloride (MH) and sodium phenylbutyrate (SP) are indicated to be able to prevent mitochondrial destruction, reduce lipid peroxidation, and protect dopamine synthesis. MH was validated to prevent neuronal death and synaptic damage, improve motor function, and reduce anhedonic and depressive-like behaviors of PD mice. The interaction of MH with a PD-related protein, sigma1, was confirmed experimentally. Thus, our findings support that MH potentially ameliorates PD by interacting with sigma1.
The widely used American College of Medical Genetics and Genomics (ACMG)/Association for Molecular Pathology (AMP) variant classification system is inherently limited by its binary categorization of variants as “pathogenic” or “benign,” failing to account for the full spectrum of variant effects within the complex genetic architecture of human disease. Although various refinements have been proposed, a framework that adequately captures this continuum remains to be established. To address this limitation, we conducted an in-depth analysis of SPINK1 variants associated with chronic pancreatitis (CP), a disorder ranging from Mendelian to environmentally influenced forms. We collated and reviewed SPINK1 variants identified in both genome-wide association studies (GWASs) and non-GWASs. Focusing on predicted loss-of-function (LoF) and experimentally characterized variants, we demonstrated through aggregation analysis that complete- or near-complete-LoF SPINK1 variants cause autosomal-dominant disease with moderate penetrance (∼55%). This finding establishes a critical baseline for comprehensively deciphering the genetic complexity underlying SPINK1-related CP. Concentrating on two well-characterized partial-LoF (hypomorphic) variants, c.194+2T>C and c.-4141G>T (enhancer), we present converging evidence for a distinct variant category that neither aligns with the ACMG/AMP binary classifications nor fits the recently proposed “risk alleles” category. Although some variants remain classified as variants of uncertain significance (VUSs), we propose a refined classificatory framework that integrates “risk,” “predisposing,” and “pathogenic” variants to accommodate the full spectrum of clinically relevant SPINK1 variants. This refined framework is expected to serve as a model for variant interpretation beyond SPINK1, providing insights into the issue of “missing heritability” and fostering further exploration of variant effects and genetic complexity across different contexts of human disease.
BACKGROUND:The post-anal tail is a common physical feature of vertebrates including mammals. Although it exhibits rich phenotypic diversity, its development has been evolutionarily conserved as early as the embryonic period. Genes participating in embryonic tail morphogenesis have hitherto been widely explored on the basis of experimental discovery, whereas the associated cis-regulatory elements (CREs) have not yet been systematically investigated for vertebrate/mammalian tail development. RESULTS:Here, utilizing high-throughput sequencing schemes pioneered in mice, we profiled the dynamic transcriptome and CREs marked by active histone modifications during embryonic tail morphogenesis. Temporal and spatial disparity analyses revealed the genes specific to tail development and their putative CREs, which facilitated the identification of novel molecular expression features and potential regulatory influence of non-coding loci including long non-coding RNA (lncRNA) genes and CREs. Moreover, these identified sets of multi-omics data supply genetic clues for understanding the regulatory effects of relevant signaling pathways (such as Fgf, Wnt) dominating embryonic tail morphogenesis. CONCLUSIONS:Our work brings new insights and provides exploitable fundamental datasets for the elucidation of the complex genetic mechanisms responsible for the formation of the vertebrate/mammalian tail.
Many drug failures in clinical trials are due to inadequate safety profiles. We developed an in-silico side effect genetic priority score (SE-GPS) that leverages human genetic evidence to inform side effect risk for a given drug target. We construct the SE-GPS in the Open Target dataset using post-marketing side effect data, externally test it in OnSIDES using side effects reported from drug labels and then generate a SE-GPS for 19,422 protein coding genes and 502 phecodes, of which 1.7% had a SE-GPS > 0. To consider drug mechanism, we incorporated the direction of genetic effect into a directional version of the score called the SE-GPS-DOE. We observe that restricting to at least two lines of genetic evidence conferred a 2.3- and 2.5-fold increased risk in side effects in Open Targets and OnSIDES respectively, with increased enrichments in severe drugs. We make all predictions publicly available in a web portal.
Regular, systematic, and independent assessments of computational tools that are used to predict the pathogenicity of missense variants are necessary to evaluate their clinical and research utility and guide future improvements. The Critical Assessment of Genome Interpretation (CAGI) conducts the ongoing Annotate-All-Missense (Missense Marathon) challenge, in which missense variant effect predictors (also called variant impact predictors) are evaluated on missense variants added to disease-relevant databases following the prediction submission deadline. Here we assess predictors submitted to the CAGI 6 Annotate-All-Missense challenge, predictors commonly used in clinical genetics, and recently developed deep learning methods. We examine performance across a range of settings relevant for clinical and research applications, focusing on different subsets of the evaluation data as well as high-specificity and high-sensitivity regimes. Our evaluations reveal notable advances in current methods relative to older, well-cited tools in the field. While meta-predictors tend to outperform their constituent individual predictors, several newer individual predictors perform comparably to commonly used meta-predictors. Predictor performance varies between high-specificity and high-sensitivity regimes, highlighting that different methods may be optimal for different use cases. We also characterize two potential sources of bias. Predictors that incorporate allele frequency as a predictive feature tend to have reduced performance when distinguishing pathogenic variants from very rare benign variants, and predictors trained on pathogenicity labels from curated variant databases often inherit gene-level label imbalances. Our findings help illuminate the clinical and research utility of modern missense variant effect predictors and identify potential areas for future development.
Combining genotype and phenotype data promises to greatly increase the value of macaque as biomedical models for human disease. Here we launch the Macaque Biobank project by deeply sequencing 919 captive Chinese rhesus macaques (CRM) while assessing 52 phenotypic traits. Genomic analyses reveal the captive CRMs are a mixture of multiple wild sources and exhibit significantly lower mutational load than their Indian counterparts. We identify hundreds of loss-of-function variants linked to human inherited disease and drug targets, and at least seven exert significant effects on phenotypes using forward genomic screens. Genome-wide association analyses reveal 30 independent loci associated with phenotypic variations. Using reverse genomic approaches, we identify DISC1 (p.Arg517Trp) as a genetic risk factor for neuropsychiatric disorders, with macaques carrying this deleterious allele exhibiting impairments in working memory and cortical architecture. This study demonstrates the potential of macaque cohorts for the investigation of genotype-phenotype relationships and exploring potential spontaneous models of human genetic disease.
SMARCB1 is a core unit of the BAF chromatin remodelling complex and its functional impairment interferes with the self-renewal and pluripotency of stem cells, lineage commitment, cellular identity and differentiation. SMARCB1 is also an important tumour suppressor gene and somatic SMARCB1 pathogenic variants (PVs) have been detected in 5
Missense variants play a key role in the diagnosis of genetic disorders and in disease risk prediction. Existing methods focus primarily on the prediction of variant effects in terms of their deleteriousness, without taking into account the disease-specific context, and are therefore limited in terms of their utility in real-world diagnosis and decision making. Here, we introduce di sease-specific va riant pathogenicity prediction (DIVA), a novel deep learning framework that directly predicts specific disease types alongside the probability of deleteriousness for missense variants. Our approach integrates information from two different modalities - protein sequence and disease-related textual annotations - encoded using two pre-trained language models and optimized within a contrastive learning paradigm designed to align variants with relevant diseases in the learned representation space. Our results demonstrate that DIVA outperforms baselines and provides accurate disease predictions with high relevance to clinically curated disease annotations for missense variants. Variant deleteriousness prediction is enhanced by incorporating AlphaMissense scores through learnable weights derived from protein function annotations, which additionally boosts DIVA ' s ability to accurately classify deleterious variants. Our work provides new insights into variant pathogenicity prediction with awareness of disease specificity, addressing a hitherto unmet need in relation to clinical variant interpretation.
Background Understanding the genetics underlying cancer development and progression is the most important goal of biomedical research to improve patient survival rates. Recently, researchers have proposed computationally combining the mutational burden with biological networks as a novel means to identify cancer driver genes. However, these approaches treated all mutations as having the same functional impact on genes and incorporated gene-gene interaction networks without considering tissue specificity, which may have hampered our ability to identify novel cancer drivers. Methods We have developed a framework, DGAT-cancer that integrates the predicted pathogenicity of somatic mutation in cancers and germline variants in the healthy population, with topological networks of gene expression in tumor tissues, and the gene expression levels in tumor and paracancerous tissues in predicting cancer drivers. These features were filtered by an unsupervised approach, Laplacian selection, and those selected were combined by Hotelling and Box-Cox transformations to score genes. Finally, the scored genes were subjected to Gibbs sampling to determine the probability that a given gene is a cancer driver. Results This method was applied to nine types of cancer, and achieved the best area under the precision-recall curve compared to three commonly used methods, leading to the identification of 571 novel cancer drivers. One of the top genes, EEF1A1 was experimentally confirmed as a cancer driver of glioma. Knockdown of EEF1A1 led to a ~ 41-50% decrease in glioma size and improved the temozolomide sensitivity of glioma cells. Conclusion By combining the pathogenic status of mutational spectra in tumors alongside the spectrum of variation in the healthy population, with gene expression in both tumors and paracancerous tissues, DGAT-cancer has significantly improved our ability to detect novel cancer driver genes.
Aims Many studies indicated use of diabetes medications can influence the electrocardiogram (ECG), which remains the simplest and fastest tool for assessing cardiac functions. However, few studies have explored the role of genetic factors in determining the relationship between the use of diabetes medications and ECG trace characteristics (ETC). Methods Genome-wide association studies (GWAS) were performed for 168 ETCs extracted from the 12-lead ECGs of 42,340 Europeans in the UK Biobank. The genetic correlations, causal relationships, and phenotypic relationships of these ETCs with medication usage, as well as the risk of cardiovascular diseases (CVDs), were estimated by linkage disequilibrium score regression (LDSC), Mendelian randomization (MR), and regression model, respectively. Results The GWAS identified 124 independent single nucleotide polymorphisms (SNPs) that were study-wise and genome-wide significantly associated with at least one ETC. Regression model and LDSC identified significant phenotypic and genetic correlations of T-wave area in lead aVR (aVR_T-area) with usage of diabetes medications (ATC code: A10 drugs, and metformin), and the risks of ischemic heart disease (IHD) and coronary atherosclerosis (CA). MR analyses support a putative causal effect of the use of diabetes medications on decreasing aVR_T-area, and on increasing risk of IHD and CA. Conclusion Patients taking diabetes medications are prone to have decreased aVR_T-area and an increased risk of IHD and CA. The aVR_T-area is therefore a potential ECG marker for pre-clinical prediction of IHD and CA in patients taking diabetes medications.
Studies have shown that drug targets with human genetic support are more likely to succeed in clinical trials. Hence, a tool integrating genetic evidence to prioritize drug target genes is beneficial for drug discovery. We built a genetic priority score (GPS) by integrating eight genetic features with drug indications from the Open Targets and SIDER databases. The top 0.83%, 0.28% and 0.19% of the GPS conferred a 5.3-, 9.9- and 11.0-fold increased effect of having an indication, respectively. In addition, we observed that targets in the top 0.28% of the score were 1.7-, 3.7- and 8.8-fold more likely to advance from phase I to phases II, III and IV, respectively. Complementary to the GPS, we incorporated the direction of genetic effect and drug mechanism into a directional version of the score called the GPS with direction of effect. We applied our method to 19,365 protein-coding genes and 399 drug indications and made all results available through a web portal.