The advent of high-throughput sequencing technologies has revolutionized cardiovascular genetics, generating vast amounts of data. A major challenge in the genomic era is the efficient and accurate identification of causative genetic variants associated with cardiovascular disease. Artificial intelligence (AI)-driven computational approaches offer a powerful solution by enabling the automation of variant classification, improving consistency and reproducibility, and enhancing predictive accuracy. These methods may guide not only variant classification, variant effect prediction, and clinical prioritization, but also open the door to data-driven precision medicine, enabling more accurate and individualized diagnoses and targeted therapeutic strategies tailored to each patient’s unique genetic and clinical profile. This state-of-the-art review provides a structured overview of AI methodologies applied to cardiovascular genomics, beginning with rule-based expert systems that encode standardized guidelines for consistent variant interpretation. Next, we examine machine learning approaches capable of identifying complex patterns in annotated multimodal clinical and multi-omic datasets. The role of deep learning algorithms is highlighted for their ability to extract features from high-dimensional, unstructured data relevant to cardiovascular disease. In addition, the potential of generative AI is explored, including applications in synthetic data generation, variant impact prediction, and automated summarization of biomedical literature. Despite advances, several challenges remain, including data heterogeneity, the need for explainable AI models to elucidate the decision mechanisms, and the complexity linked to the integration of AI-based tools into clinical workflows. Addressing these issues requires interdisciplinary collaboration among clinicians, geneticists, data scientists, and bioinformaticians to ensure the effective translation of AI-generated insights into clinical practice. This review aims to provide a comprehensive perspective on the evolving role of AI in cardiovascular genomics and its implications for advancing precision medicine.
Abstract Motivation Standardized phenotypic descriptions are essential for accurate diagnosis, yet clinicians and researchers face challenges in manually extracting and mapping phenotypes from scientific literature or patient clinical records to the Human Phenotype Ontology. Recent advances in deep learning offer new opportunities for automation. We developed PhenoXtract, a novel phenotype extraction approach that combines Large Language Models and Knowledge Graph embedding. PhenoXtract is a multistep pipeline that takes clinical descriptions as input, extracts candidate phenotype entities using large language models, and maps them to terms from an enriched version of the Human Phenotype Ontology, processed as a knowledge graph. Results Evaluation against expert-curated ground-truth datasets show a recall of 0.70 and precision of 0.85 for PhenoXtract, demonstrating concordance with manually extracted phenotypes, with a computation time of 10-20 seconds for each text analyzed. Moreover, PhenoXtract surpasses rule-based and deep learning-based state-of-the-art tools in two out of the three ground-truth datasets evaluated. These results suggest that hybrid approaches combining Large Language Models and Knowledge Graph embeddings represent a promising direction for automated clinical phenotyping at scale. Contact sberardelli@engenome.com
AIMS:The current diagnostic approach to inherited cardiac conditions (ICCs) is primarily focused on the analysis of single-nucleotide variants (SNVs) and small insertions/deletions (InDels). However, as recommended for other inherited diseases, the analysis of Copy Number Variants (CNVs) should be equally considered. In cardiology, the diagnostic contribution of CNVs remains insufficiently studied, with limited evidence and no standardized recommendations for analytical workflows. This study aims to assess the prevalence of pathogenic CNVs in an Italian multicentre cohort of ICC patients, providing recommendations for integrating CNV analysis into routine genetic testing. METHODS AND RESULTS:A total of 203 ICC probands with prior negative or inconclusive results for SNVs and InDels testing were included. The tertiary bioinformatic analysis was performed by eVai enGenome software, and putative CNVs were validated by Multiplex Ligation-dependent Probe Amplification (MLPA) assays. MLPA confirmed only one of the five initially suspected CNVs, identifying a deletion in MYBPC3 gene classified as pathogenic according to ACMG/AMP guidelines. This result confirmed a limited diagnostic yield of 0.50%. Considering the substantial costs, time constraints, and specialized expertise required, we propose a strategy to prioritize selected ICC patients for CNV analysis. CONCLUSION:Evidence from this real-world cohort suggests the incorporation of CNV analysis as a second-tier test in patients with ICC and specific clinical or molecular indications.
The digenic inheritance hypothesis holds the potential to enhance diagnostic yield in rare diseases. Computational approaches capable of accurately interpreting and prioritizing digenic combinations of variants based on the proband’s phenotypes and family information can provide valuable assistance during the diagnostic process. We developed diVas, a hypothesis-driven machine learning approach that interprets genomic variants across different gene pairs. DiVas demonstrates strong performance in both classifying and prioritizing causative digenic combinations of rare variants within the top positions across 11 cases with the complete list of variants available (73% sensitivity and a median ranking of 3). Furthermore, it achieves a sensitivity of 0.81 when applied to 645 published causative digenic combinations. Additionally, diVas leverages explainable artificial intelligence to elucidate the digenic disease mechanism for predicted positive pairs.
Background: Structural variants (SVs) play a significant role in gene function and are implicated in numerous human diseases. With advances in sequencing technologies, identifying SVs through whole-genome sequencing (WGS) has become a key area of research. However, variability in SV detection persists due to the wide range of available tools and the absence of standardized methodologies. Methods: We assessed the accuracy of SV detection across various short-read (srWGS) and long-read (lrWGS) sequencing technologies—including Illumina short reads, PacBio long reads, and Oxford Nanopore Technologies (ONT) long reads—using deletion calls from the HG002 benchmark dataset. We examined how variables such as variant calling algorithms, reference genome choice, alignment strategies, and sequencing coverage influence SV detection performance. Results: DRAGEN v4.2 delivered the highest accuracy among ten srWGS callers tested. Notably, leveraging a graph-based multigenome reference improved SV calling in complex genomic regions. Moreover, we proved that combining minimap2 with Manta achieved performance comparable to DRAGEN for srWGS. For PacBio lrWGS data, Sniffles2 outperformed the other two tested tools. For ONT lrWGS, alignment with minimap2—among four aligners tested—consistently led to the best results. At up to 10× coverage, Duet achieved the highest accuracy, while at higher coverages, Dysgu yielded the best results. Conclusions: These results show for the first time that alignment software choice significantly impacts SV calling from srWGS, with results comparable to commercial solutions. For lrWGS, the performance depends on the technology and coverage.
Chromoanagenesis events consist of complex chromosome rearrangements with multiple breakpoints in one or few chromosomes. Mechanisms of chromoanagenesis are split into three major groups: chromothripsis, chromoanasynthesis and chromoplexy. This study aims to delineate a chromoanagenesis event at the level of chromosome 22 in an individual showing obesity and borderline cognitive performance as major disturbances. The proband and his parents were subjected to conventional karyotyping, CGH array and whole genomic sequencing (WGS). By conventional karyotyping a "de novo" pericentric inversion of chromosome 22 was identified. CGH array identified several imbalances (either deletions or duplications) in the long arm of chromosome 22; the largest is a 4.5 Mb duplication at 22q12.1-22q1.3. The detection of extensive duplications would suggest the occurrence of a chromoanasynthesis event. WGS, in addition to the structural alterations identified by karyotyping and CGH array, revealed two translocations from chromosome 22 to chromosomes 6 and 21 as well as a heterozygous pathogenetic variant of ALMS1 gene; the latter could have contributed to the obesity of our patient. The pericentric inversion induces loss of initial part of TCF20 gene including the 5' regulatory region and the first, noncoding, exon. Heterozygous loss-of-function mutations of TCF20 gene have been found in patients with autism spectrum disorder or intellectual disability, some of them presenting obesity. It is, therefore, possible that disruption of TCF20 gene structure would contribute to a fraction of the patient's phenotype.
Background: Sudden death is the leading cause of mortality in medically refractory epilepsy. Middle-aged persons with epilepsy (PWE) are under investigated regarding their mortality risk and burden of cardiovascular disease (CVD). Methods: Using UK Biobank, we identi fi ed 7786 (1.6%) participants with diagnoses of epilepsy and 6,171,803 person -years of follow-up (mean 12.30 years, standard deviation 1.74); 566 patients with previous histories of stroke were excluded. The 7220 PWE comprised the study cohort with the remaining 494,676 without epilepsy as the comparator group. Prevalence of CVD was determined using validated diagnostic codes. Cox proportional hazards regression was used to assess all -cause mortality and sudden death risk. Results: Hypertension, coronary artery disease, heart failure, valvular heart disease, and congenital heart disease were more prevalent in PWE. Arrhythmias including atrial fi brillation/ fl utter (12.2% vs 6.9%; P < 0.01), bradyarrhythmias (7.7% vs 3.5%; P < 0.01), conduction defects (6.1% vs 2.6%; P < 0.01), and ventricular arrhythmias (2.3% vs 1.0%; P < 0.01), as well as cardiac implantable electric devices (4.6% vs 2.0%; P < 0.01) were more prevalent in PWE. PWE had higher adjusted all -cause mortality (hazard ratio [HR], 3.9; 95% con fi dence interval [CI], 3.01-3.39), and sudden death -speci fi c mortality (HR, 6.65; 95% CI, 4.53-9.77); and were almost 2 years younger at death (68.1 vs 69.8; P < 0.001). Conclusions: Middle-aged PWE have increased all -cause and sudden death -speci fi c mortality and higher burden of CVD including arrhythmias and heart failure. Further work is required to elucidate mechanisms underlying all -cause mortality and sudden death risk in PWE of middle age, to identify prognostic biomarkers and develop preventative therapies in PWE.
MOTIVATION:In the modern era of genomic research, the scientific community is witnessing an explosive growth in the volume of published findings. While this abundance of data offers invaluable insights, it also places a pressing responsibility on genetic professionals and researchers to stay informed about the latest findings and their clinical significance. Genomic variant interpretation is currently facing a challenge in identifying the most up-to-date and relevant scientific papers, while also extracting meaningful information to accelerate the process from clinical assessment to reporting. Computer-aided literature search and summarization can play a pivotal role in this context. By synthesizing complex genomic findings into concise, interpretable summaries, this approach facilitates the translation of extensive genomic datasets into clinically relevant insights. RESULTS:To bridge this gap, we present VarChat (varchat.engenome.com), an innovative tool based on generative AI, developed to find and summarize the fragmented scientific literature associated with genomic variants into brief yet informative texts. VarChat provides users with a concise description of specific genetic variants, detailing their impact on related proteins and possible effects on human health. In addition, VarChat offers direct links to related scientific trustable sources, and encourages deeper research. AVAILABILITY AND IMPLEMENTATION:varchat.engenome.com.
BACKGROUND:A major obstacle faced by families with rare diseases is obtaining a genetic diagnosis. The average "diagnostic odyssey" lasts over five years and causal variants are identified in under 50%, even when capturing variants genome-wide. To aid in the interpretation and prioritization of the vast number of variants detected, computational methods are proliferating. Knowing which tools are most effective remains unclear. To evaluate the performance of computational methods, and to encourage innovation in method development, we designed a Critical Assessment of Genome Interpretation (CAGI) community challenge to place variant prioritization models head-to-head in a real-life clinical diagnostic setting. METHODS:We utilized genome sequencing (GS) data from families sequenced in the Rare Genomes Project (RGP), a direct-to-participant research study on the utility of GS for rare disease diagnosis and gene discovery. Challenge predictors were provided with a dataset of variant calls and phenotype terms from 175 RGP individuals (65 families), including 35 solved training set families with causal variants specified, and 30 unlabeled test set families (14 solved, 16 unsolved). We tasked teams to identify causal variants in as many families as possible. Predictors submitted variant predictions with estimated probability of causal relationship (EPCR) values. Model performance was determined by two metrics, a weighted score based on the rank position of causal variants, and the maximum F-measure, based on precision and recall of causal variants across all EPCR values. RESULTS:Sixteen teams submitted predictions from 52 models, some with manual review incorporated. Top performers recalled causal variants in up to 13 of 14 solved families within the top 5 ranked variants. Newly discovered diagnostic variants were returned to two previously unsolved families following confirmatory RNA sequencing, and two novel disease gene candidates were entered into Matchmaker Exchange. In one example, RNA sequencing demonstrated aberrant splicing due to a deep intronic indel in ASNS, identified in trans with a frameshift variant in an unsolved proband with phenotypes consistent with asparagine synthetase deficiency. CONCLUSIONS:Model methodology and performance was highly variable. Models weighing call quality, allele frequency, predicted deleteriousness, segregation, and phenotype were effective in identifying causal variants, and models open to phenotype expansion and non-coding variants were able to capture more difficult diagnoses and discover new diagnoses. Overall, computational models can significantly aid variant prioritization. For use in diagnostics, detailed review and conservative assessment of prioritized variants against established criteria is needed.