In rare disease diagnosis, described genotype-phenotype associations are evaluated first. In the absence of strong evidence, WES and WGS provide hundred to million other genetic variants, most poorly annotated, that need to be prioritized. While several in silico approaches leverage existing gene-disease knowledge to predict novel associations, doing so in isolation can hide how different genes are represented across other predictions. We hypothesize that a global perspective, accounting for differences in the knowledge accumulated in the gene collections, can refine predictions. Using a network-based algorithm, we explored functional neighborhoods of known disease-associated genes to predict novel candidates for over 200 rare and other Mendelian diseases. A global analysis of gene and protein family behavior across predictions identified genes and functions broadly associated with multiple conditions, 192 genes linked to a single disease and 251 genes functionally associated with specific classes of genetic diseases. These findings are integrated into a gene-disease specificity score, aimed at enhancing variant prioritization and guiding geneticists in advancing candidate genes toward functional validation.
Prostate cancer (PCa) is one of the most heritable common malignancies, yet data on germline genetic susceptibility in the Spanish population remain limited, and testing indications vary across guidelines. We aimed to characterize germline pathogenic variants (PVs) in Spanish PCa patients and identify predictors of carrier status. We retrospectively analyzed 360 PCa patients referred for genetic counseling. Multigene panel testing of 13 cancer predisposition genes was performed. PV frequencies were compared with internal pseudo-controls and public population controls. Associations with age at diagnosis, clinical risk group, and family history (FH) were evaluated. Overall, 6.4% of patients carried a germline PV. BRCA2 and ATM were the most frequently mutated genes (1.67% each), showing significant enrichment compared with all control groups. TP53 and PALB2 were also significantly associated with PCa risk. High-risk or metastatic disease, age < 60 years, and FH of hereditary breast/ovarian or Lynch syndrome-spectrum tumors were the best predictors for the presence of PV. Our findings suggest an unexpected major role for TP53 and PALB2 as key PCa susceptibility genes, reinforce the role of BRCA2 and ATM, and question the role of BRCA1. Genetic testing may be most informative when focused on younger patients with non-low-risk disease and relevant FH.
Biallelic variants in GTF3C3 have been recently associated with an autosomal recessive form of syndromic intellectual disability, with only 16 patients reported to date. Using whole exome sequencing, we identified previously unreported biallelic GTF3C3 variants in a new patient, who also represents the oldest reported individual so far. We detail the genotype and phenotype of this case and provide a comprehensive review of all published patients to better define the clinical features and their frequency. The most consistent manifestations across the cohort include global developmental delay, intellectual disability, epilepsy, central nervous system malformations, and additional neurological abnormalities. This case, together with the literature review, further delineates the phenotype associated with GTF3C3.
ABSTRACT Rare diseases (RDs) pose a major diagnostic challenge. Genetic and phenotypic heterogeneity, incomplete knowledge of disease mechanisms, and limitations in variant clinical interpretation leave many patients without a molecular diagnosis. Meanwhile, the growing volume of genomic data generated in clinical practice offers an opportunity to develop data-driven methodologies for exploring disease mechanisms and supporting the reanalysis of unsolved cases. In this study, we aggregated real-world genomic data from 11,084 unrelated patients clinically classified into 122 diseases. We built a multi-disease genomic variant frequency database (FJD-DB), which enabled the development of variant and gene-disease association scores by means of case-control subcohort comparisons across 32 disease groups. Functional enrichment analyses were then used to highlight disease-associated protein domains, pathways, biological processes, and phenotypes. Finally, the resulting knowledge was integrated into a data-driven framework for the guided reanalysis of unsolved RD patients applied to Inherited Retinal Dystrophies (IRD) patients as first use case. The FJD-DB contained over 45 million unique variants, including ∼184,000 potentially pathogenic variants. We identified disease-associated pathogenic variants and highlighted both established and candidate disease genes. We identified 179 protein domains, 124 Human Phenotype Ontology terms, 79 Reactome pathways, and 72 Gene Ontology biological processes significantly enriched across multiple diseases, revealing disease-specific functional signatures. Integration of disease-associated variant, gene, and functional signals enabled the development of a data-driven framework for guided reanalysis of unsolved RD cases. Applied to over 1,100 unsolved IRD cases, this data-driven reanalysis yielded clinically relevant findings in 32 patients, including three molecular diagnoses, 25 possibly solved cases, and four cases in which prioritized variants highlighted four candidate genes. We illustrate how aggregated real-world genomic data can be leveraged to identify disease-associated molecular signals generating novel biological hypotheses. A unified analytical framework provides a scalable strategy for knowledge discovery and guided reanalysis, facilitating the identification of overlooked and potentially novel genetic causes of RDs.
Viral isolates consist of complex mixtures of many variants that are termed mutant spectra, distributions, clouds or swarms. They include information on virus behavior that is not captured by consensus sequences. Here, we describe experimental procedures, a bioinformatics pipeline, and calculations to characterize the mutant spectrum of RNA viruses. The data are obtained through the high-resolution MiSeq Illumina ultra-deep sequencing platform. We describe protocols for SARS-CoV-2 patients' isolates and laboratory populations, which can be adapted to other RNA viral pathogens. Precautions for sample handling to avoid cross-contaminations and controls for mutation and deletion detection reliability are also outlined.
The broad genetic heterogeneity of neurodevelopmental disorders (NDDs) makes their molecular diagnosis particularly challenging. In this context, Whole-Exome Sequencing (WES), specifically in a trio-based design, is a powerful strategy due to its ability to detect de novo variants, which are a major contributor to NDDs. However, its clinical implementation is often limited by its associated cost. In this study, we applied a sequential diagnostic workflow to a cohort of 221 individuals with syndromic NDDs and prior negative results from targeted sequencing. The workflow integrates initial solo-WES, followed by a second-tier trio-WES using pooled parental DNA (trio pooled-WES). Overall, this workflow achieved a diagnostic yield of 20.98% and led to the identification of 13 novel candidate genes. The pooling strategy was optimized and validated, demonstrating that trio pooled-WES retains the main advantages of conventional trio-WES while substantially reducing sequencing costs. These results support its implementation as a clinically applicable approach for the genetic diagnosis of NDDs.
Abstract Motivation Allele frequency (AF) is central to clinical variant classification under ACMG/AMP guidelines. Public reference databases offer broad ancestry coverage, but local ancestries, rare-disease enrichment, and institutional case distributions are often underrepresented, so cohort-derived AF is a valuable complement. Computing accurate AF from institutional cohorts is nonetheless error-prone: even successive versions of the same capture kit cover substantially different target regions, and naive methods inflate the allele number (AN) at positions not shared by all kits, deflating AF and biasing ACMG frequency evidence toward pathogenic categories. Results We present AFQuery, a bitmap-indexed AF engine that computes capture-aware, ploidy-aware allele frequencies from pre-indexed Roaring Bitmaps in ∼ 14 ms per point query (∼34 ms for 1-Mbp region queries), independently of cohort size up to 50,000 samples. In simulated mixed-technology cohorts, capture-aware AN reduced AF mean absolute error 8–13-fold and removed the systematic bias toward pathogenic ACMG categories, yielding 10–45-fold fewer spurious pathogenic-evidence calls. Availability AFQuery is freely available under the MIT licence at https://github.com/babelomics/afquery .
Background: Prostate cancer (PCa) is one of the most heritable common malignancies, yet data on germline genetic susceptibility in the Spanish population remain limited and testing indications vary across guidelines. We aimed to characterize germline pathogenic variants (PVs) in Spanish PCa patients and identify predictors of carrier status. Methods: We retrospectively analyzed 360 PCa patients referred for genetic counseling. Multigene panel testing of 13 cancer predisposition genes was performed. PV frequencies were compared with internal pseudo-controls and public population controls. Associations with age at diagnosis, clinical risk group, and family history (FH) were evaluated. Results: Overall, 6.4% of patients carried a germline PV. BRCA2 and ATM were the most frequently mutated genes (1.67% each), showing significant enrichment compared with all control groups. PALB2 and TP53 were also significantly associated with PCa risk. High-risk or metastatic disease, age < 60 years, and FH of hereditary breast/ovarian or Lynch syndrome-spectrum tumors were the best predictors for the presence of PV. Conclusions: Our findings support BRCA2 and ATM as key PCa susceptibility genes and suggest an unexpected major role for PALB2 and TP53 . Genetic testing may be most informative when focused on younger patients with non-low-risk disease and relevant FH.
Interstitial lung disease (ILD) is a major manifestation of systemic autoimmune rheumatic diseases (SARD). Unmet needs in this population are the early identification of patients at risk of functional deterioration and tailoring therapies on a pathophysiological basis. The study of microRNA is being applied to trace pathogenetic mechanisms and as predictive tool in complex diseases. We have performed bulk small-RNA sequencing in sera from 24 patients with SARD-associated ILD and 5 healthy subjects, in order to identify differentially expressed molecules in the patients and in disease subgroups. Variables of the study included disease duration, radiographic patterns, clinical diagnosis, alterations in nailfold capillaroscopy, functional status, levels of Krebs von den Lungen 6 (KL-6), presence of rheumatoid factor, and outcomes over the following 18-month period. The patients showed 13 differentially expressed miRNA, out of which Let-7i-5p associated with higher KL-6 levels, whereas miR-151a-5p was increased during early disase and miR-483-5p was higher in patients with microvascular involvement. The latter subgroup displayed a specific signature, characterized by the up-regulation of the miR-320 and miR-10 families. Also to underscore was the value of miR-223, miR-142-5p, miR-145-5p, miR-23a-5p, miR-29a-3p, miR-320c, miR-320d and miR-10a-3p in forecasting a progressive phenotype. Functional analysis pointed to an up-regulation of profibrotic miRNA as early biomarkers of poor outcome. On the whole, our data uncover for the first time miRNomes associated to SARD-ILD in an exporatory approach. Whether or not the miRNA shown up in this study applies to the whole population of SARD-ILD warrants confirmation in further research.
RNA virus populations consist of complex and dynamic mutant spectra in which most individual genomes differ in one or more positions from the other genomes of the same population. This behavior, known as quasispecies dynamics, applies to SARS-CoV-2 which exhibits intrahost genetic and functional heterogeneity while evolving at a high rate in the human population. In the present study, we describe a remarkable reduction in mutant spectrum complexity (intrahost viral genome heterogeneity) in SARS-CoV-2 isolates of late relative to early COVID-19 waves, as they reached Madrid (Spain) from 2020 until 2022. In contrast, the consensus (average) sequence of the corresponding isolates displayed a continuing divergence from the initial Wuhan-Hu-1 virus as the pandemic advanced. The mutant spectrum complexity developed upon replication in Vero E6 cells of the isolates from the first and sixth COVID-19 waves, as well as of biological clones retrieved from them, was similar. Therefore, the mutant spectrum complexity reduction observed in vivo was not due to an increased accuracy of the viral replicative machinery, but rather to other factors related to viral epidemiology or pathogenesis. Such possible factors and their implications for viral trait modifications in the course of a viral pandemic are discussed. The results establish that mutant spectrum complexity of genetically variable viruses can be an epidemiologically evolvable trait.
AbstractBackgroundAtherosclerosis is a chronic inflammatory disease characterized by the accumulation of lipids and leukocytes within the arterial wall. By studying the aortic transcriptome of atherosclerosis‐prone apolipoprotein E (ApoE−/−) mice, we aimed to identify novel players in the progression of atherosclerosis.MethodsRNA‐Seq analysis was performed on aortas from ApoE−/− and wild‐type mice. AnxA8 expression was assessed in human and mice atherosclerotic tissue and healthy aorta. ApoE−/− mice lacking systemic AnxA8 (ApoE−/−AnxA8−/−) were generated to assess the effect of AnxA8 deficiency on atherosclerosis. Bone marrow transplantation (BMT) was also performed to generate ApoE−/− lacking AnxA8 specifically in bone marrow‐derived cells. Endothelial‐specific AnxA8 silencing in vivo was performed in ApoE−/− mice. The functional role of AnxA8 was analysed in cultured murine cells.ResultsRNA‐Seq unveiled AnxA8 as one of the most significantly upregulated genes in atherosclerotic aortas of ApoE−/− compared to wild‐type mice. Moreover, AnxA8 was upregulated in human atherosclerotic plaques. Germline deletion of AnxA8 decreased the atherosclerotic burden, the size and volume of atherosclerotic plaques in the aortic root. Plaques of ApoE−/−AnxA8−/− were characterized by lower lipid and inflammatory content, smaller necrotic core, thicker fibrous cap and less apoptosis compared with those in ApoE−/−AnxA8+/+. BMT showed that hematopoietic AnxA8 deficiency had no effect on atherosclerotic progression. Oxidized low‐density lipoprotein (ox‐LDL) increased AnxA8 expression in murine aortic endothelial cells (MAECs). In vitro experiments revealed that AnxA8 deficiency in MAECs suppressed P/E‐selectin and CD31 expression and secretion induced by ox‐LDL with a concomitant reduction in platelet and leukocyte adhesion. Intravital microscopy confirmed the reduction in leukocyte and platelet adhesion in ApoE−/−AnxA8−/− mice. Finally, endothelial‐specific silencing of AnxA8 decreased atherosclerosis progression.ConclusionOur findings demonstrate that AnxA8 promotes the progression of atherosclerosis by modulating endothelial−leukocyte interactions. Interventions capable of reducing AnxA8 expression in endothelial cells may delay atherosclerotic plaque progression.Key points This study shows that AnxA8 is upregulated in aorta of atheroprone mice and in human atherosclerotic plaques. Germline AnxA8 deficiency reduces platelet and leukocyte recruitment to activated endothelium as well as atherosclerotic burden, plaque size, and macrophage accumulation in mice. AnxA8 regulates oxLDL‐induced adhesion molecules expression in aortic endothelial cells. Our data strongly suggest that AnxA8 promotes disease progression through regulation of adhesion and influx of immune cells to the intima. Endothelial specific silencing of AnxA8 reduced atherosclerosis progression. Therapeutic interventions to reduce AnxA8 expression may delay atherosclerosis progression.
Advances in whole-genome sequencing (WGS) have significantly enhanced our ability to detect genomic variants underlying inherited diseases. In this study, we performed long-read WGS on 24 patients with inherited retinal dystrophies (IRDs) to validate the utility of nanopore sequencing in detecting genomic variations. We confirmed the presence of all previously detected variants and demonstrated that this approach allows for the precise refinement of structural variants (SVs). Furthermore, we could perform genotype phasing by sequencing only the probands, confirming that the variants were inherited in trans. Moreover, nanopore sequencing enables the detection of complex variants, such as transposon insertions and structural rearrangements. This comprehensive assessment illustrates the power of long-read sequencing in capturing diverse forms of genomic variation and in improving diagnostic accuracy in IRDs.
ABSTRACT Defective genomes are part of SARS-CoV-2 quasispecies. High-resolution, ultra-deep sequencing of bulk RNA from viral populations does not distinguish RNA mutations, insertions, and deletions in viable genomes from those in defective genomes. To quantify SARS-CoV-2 infectious variant progeny, virus from four individual plaques (biological clones) of a preparation of isolate USA-WA1/2020, formed on Vero E6 cell monolayers, was subjected to further biological cloning to yield 9 second-generation and 15 third-generation sub-clones. Consensus genomic sequences of the biological clones and sub-clones included an average of 2.8 variations per viable genome, relative to the consensus sequence of the parental USA-WA1/2020 virus. This value is 6.5-fold lower than the estimates for biological clones of other RNA viruses such as bacteriophage Qβ, foot-and-mouth disease virus, or hepatitis C virus in cell culture. The mutant spectrum complexity of the nsp12 (polymerase)- and spike (S)-coding region was unique in the progeny of each of 10 third-generation sub-clones; they shared 2.4% of the total of 164 different mutations and deletions scored in the 3,719 genomic residues that were screened. The presence of minority out-of-frame deletions revealed the ease of defective genome production from an individual infectious genome. Several low-frequency point mutations and deletions were clade-discordant in that they were not typical of USA-WA1/2020 but served to define the consensus sequences of future SARS-CoV-2 clades. Implications for SARS-CoV-2 adaptability and COVID-19 control of the viable genome heterogeneity and the generation of complex mutant spectra from individual genomes are discussed. IMPORTANCE Sequencing of biological clones is a means to identify mutations, insertions, and deletions located in viable genomes. This distinction is particularly important for viral populations, such as those of SARS-CoV-2, that contain large proportions of defective genomes. By sequencing biological clones and sub-clones, we quantified the heterogeneity of the viable complement of USA-WA1/2020 to be lower than exhibited by other RNA viruses. This difference may be due to a reduced mutation rate or to limited tolerance of the large coronavirus genome to incorporate mutations and deletions and remain functional or a combination of both influences. The presence of clade-discordant residues in the progeny of individual biological sub-clones suggests limitations in the occupation of sequence space by SARS-CoV-2. However, the complex and unique mutant spectra that are rapidly generated from individual genomes suggest an aptness to confront selective constraints.
THRBencodes thyroid hormone receptor β which produces two human isoforms (TRβ1 and TRβ2) by alternative splicing. The first THRB variant associated with autosomal dominant macular dystrophy (ADMD), NM_001354712.2:c.283+1G>A, was recently described. This study aims to refine the ophthalmologic phenotype, report a novel THRB variant, and investigate the impact of these splicing variants at the protein level. THRBvariants were identified by re-analysis of next-generation sequencing data from the FJD database. Family segregation was performed using Sanger sequencing. Clinical data were collected from self-reported ophthalmic history questionnaires and ophthalmic exams. Functional splicing test was performed by in vitrominigene approach. We identified 12 patients with ADMD from 3 families carrying variants in THRB. Two families carried the variant NM_001354712.2:c.283+1G>A, and one the novel variant NM_001354712.2:c.283G>A. Patients exhibited common ophthalmologic findings with disruption of subfoveal ellipsoid layers, and variable onset of symptoms. Splicing assays showed complete exon 5 skipping or a 6 bp deletion in both variants. Our results support the association of THRBwith ADMD. The high intra-familial variability could be influenced by phenotype modifiers. Aberrant TRβ1/TRβ2 proteins could lead to a gain-of-function mechanism. Including THRB in inherited retinal dystrophy genetic panels could enhance diagnoses and clinical patient management.
Anti-epidermal growth factor receptor (EGFR) therapies are the most recommended first-line treatment for RAS/RAF wild-type unresectable metastatic colorectal cancer (CRC) according to the European Society for Medical Oncology guidelines. However, primary resistance renders this treatment ineffective for almost 40% of patients. Our previous work identified Aurora kinase A (AURKA) as a key resistance driver through non-canonical, Hippo-independent Yes-associated protein 1 (YAP1) activation. However, the role of the other main Hippo coactivator, transcriptional coactivator with PDZ-binding motif (TAZ), in this resistance mechanism remains unexplored. By integrating preclinical in vitro and in vivo models, including cell lines and patient-derived xenografts, with RNA sequencing, we investigated the impact of TAZ overexpression in cetuximab resistance driven by the AURKA/YAP1 axis. Our findings reveal that TAZ overexpression sustains YAP1-mediated resistance and stemness. Even under YAP1 suppression, TAZ-overexpressing cells remain unresponsive to anti-EGFR therapies, whereas dual YAP1/TAZ silencing restores sensitivity. Treatment with alisertib, a phase III AURKA inhibitor, simultaneously destabilizes YAP1 and TAZ, restoring anti-EGFR efficacy by suppressing stemness. Transcriptomic analyses further show that AURKA inhibition and dual YAP1/TAZ suppression disrupt stem-like traits and reveal transcriptional deregulations affecting nucleotide metabolism. These findings demonstrate that AURKA orchestrates YAP1/TAZ crosstalk, which is crucial for driving stemness and resistance to anti-EGFR therapies, highlighting AURKA inhibitors as a promising strategy to enhance anti-EGFR therapies in metastatic CRC.
The synergy between SWI/SNF and Polycomb Repressive Complexes (PRCs) has been a topic of extensive research. The demonstrated heightened preclinical sensitivity of SWI/SNF gene mutations to EZH2 inhibition together with the accelerated approval of Tazemetostat for selected EZH2 mutant tumors by the FDA, has significantly expanded our understanding of chromatin remodeling complexes and their vulnerabilities. Here we report a durable response in a patient with PBRM1 mutated cholangiocarcinoma treated with Tazemetostat. We provide hypothesis-generating mechanistic findings into the relationship between PBRM1 mutations and EZH2 inhibition and the interplay between the SWI/SNF and PRC2 complexes.
Background/Objectives: The COVID-19 pandemic resulted in 675 million cases and 6.9 million deaths by 2022. Despite substantial declines in case fatalities following widespread vaccination campaigns, the threat of future coronavirus outbreaks remains a concern. Current treatments for COVID-19 have been repurposed from existing therapies for other infectious and non-infectious diseases. Emerging evidence suggests a role for genetic factors in both susceptibility to SARS-CoV-2 infection and response to treatment. However, comprehensive studies correlating clinical outcomes with genetic variants are lacking. The main aim of our study is the identification of host genetic biomarkers that predict the clinical outcome of COVID-19 pharmacological treatments. Methods: In this study, we present findings from GWAS and candidate gene and pathway enrichment analyses leveraging diverse patient samples from the Spanish Coalition to Unlock Research of Host Genetics on COVID-19 (SCOURGE), representing patients treated with immunomodulators (n = 849), corticoids (n = 2202), and the combined cohort of both treatments (n = 2487) who developed different outcomes. We assessed various phenotypes as indicators of treatment response, including survival at 90 days, admission to the intensive care unit (ICU), radiological affectation, and type of ventilation. Results: We identified significant polymorphisms in 16 genes from the GWAS and candidate gene studies (TLR1, TLR6, TLR10, CYP2C19, ACE2, UGT1A1, IL-1α, ZMAT3, TLR4, MIR924HG, IFNG-AS1, ABCG1, RBFOX1, ABCB11, TLR5, and ANK3) that may modulate the response to corticoid and immunomodulator therapies in COVID-19 patients. Enrichment analyses revealed overrepresentation of genes involved in the innate immune system, drug ADME, viral infection, and the programmed cell death pathways associated with the response phenotypes. Conclusions: Our study provides an initial framework for understanding the genetic determinants of treatment response in COVID-19 patients, offering insights that could inform precision medicine approaches for future epidemics.