MDGA2 encodes a membrane-associated protein that is critical for regulating glutamatergic synapse development, modulating neuroligins (Nlgns), and maintaining excitatory-inhibitory synaptic balance. While MDGA2 functions have been extensively studied in murine and cellular models, its association with human developmental disorders has yet to be established. Through exome sequencing, we identified seven distinct homozygous loss-of-function variants in MDGA2 in nine individuals from seven consanguineous families, all presenting with developmental and epileptic encephalopathy (DEE). Clinically, these individuals exhibited a consistent phenotype including infantile hypotonia, severe neurodevelopmental delay, intractable seizures, along with distinct dysmorphic features. Neuroimaging findings included delayed/incomplete myelination, early-onset brain atrophy, white-matter thinning, basal ganglia volume loss, and small hippocampi. Functional studies of three representative nonsense variants revealed impaired MDGA2 membrane trafficking, disrupted Nlgn1 interaction, and perturbed MDGA2-mediated excitatory synaptic functions in mammalian expression systems and cultured hippocampal neurons. Our findings support the involvement of MDGA2 in a subtype of autosomal-recessive DEE. This not only underscores a loss-of-function pathogenic mechanism but also highlights the previously unrecognized role of MDGA2 in human synaptic development and regulation, significantly expanding our understanding of the genetic architecture of DEEs.
Transcriptomics enables comprehensive, multiplexed characterization of cellular states, yet prevailing methods typically require cell fixation or lysis, precluding longitudinal analysis of RNA expression in living cells. Here, we present non-destructive transcriptomics by vesicular export (NTVE), a platform for multi-time-point monitoring of RNA expression dynamics in living cells. Stabilized RNA reporter barcodes can be selectively packaged and exported from cells via virus-like particles (VLPs) bearing bioorthogonal affinity handles for convenient multichannel tracking of co-cultured cells. Using an engineered poly(A)-binding protein adapter, NTVE exports endogenous transcripts from inducible human and murine cell lines with high concordance to conventional lysate-derived RNA-seq. NTVE captures transcriptome changes in response to genetic and chemical perturbations within the same cells over time using standard sequencing workflows. NTVE can further be equipped with fusogens to deliver mRNA-encoded effectors or ribonucleoprotein gene editors from sender cells, activating gene reporters in co-cultured recipient cells. We demonstrate the utility of NTVE for monitoring hiPSC differentiation through daily non-destructive transcriptomic profiling of lineage-specific marker dynamics.
OBJECTIVE:Genomic sequencing leaves >50% of dystonia-affected individuals without a diagnosis. Where DNA-oriented approaches remain insufficient, integrating multiomics is essential to advance genome interpretation. Herein, we incorporated RNA sequencing (RNA-seq) data from 167 patients with dystonia across a range of ages and presentations. METHODS:We leveraged an RNA-seq analysis pipeline, focused on the identification of expression and splicing aberrations, on RNA-seq from skin biopsies. The recruited patients had early-onset dystonia in 85.0%, non-focal dystonia in 92.2%, and coexisting features in 76.0%. Thirty-six patient samples with pre-identified variants (36/167, 21.6%) and 131 samples with no previously prioritized diagnostic candidates from genomic sequencing (131/167, 78.4%) were evaluated. RESULTS:We found that >80% of dystonia-associated genes were detected by fibroblast RNA-seq. Expression and splicing aberration analyses produced a manageable number of significant RNA defects affecting dystonia-associated genes. The approach was especially successful in validating pathogenic effects of loss-of-function variants, with disease-relevant RNA-underexpression detected for 66.7% (10/15). Studying aberrant expression and splicing in the context of other pre-identified variant types yielded relevant results in 28.6% (6/21 samples). We obtained a 6.9% (9/131) diagnostic uplift for patients without prior candidates, all of whom exhibited combined dystonia with autosomal recessive inheritance. The new diagnoses from RNA-seq and genomic reanalysis were based on previously neglected splice-region (3/9) and deep(er) intronic (6/9) variants. For the observed events, integration of new machine-learning scores predicted corresponding aberrant gene expression in the brain. INTERPRETATION:Fibroblast-based RNA-seq in our selected cohort improved variant interpretation and offered a modest yield in patients without prior candidate variants. ANN NEUROL 2026;99:1363-1378.
RNA-binding proteins (RBPs) orchestrate a complex combinatorial regulatory "code" that governs RNA splicing, stability, localization, and translation. Learning the relationship between RNA sequences and these processes is a central challenge in genomics. Foundation models, notably RNA language models, have emerged as the dominant approach, learning general-purpose representations from unlabeled sequence at scale. While RNA language models have demonstrated impressive performance across a broad range of downstream tasks, they generally learn from sequence reconstruction objectives alone, lacking direct connections to the regulatory principles that govern RNA function. Here we introduce Parnet, an RNA foundation model trained directly and exclusively on experimental CLIP-seq data. Parnet is a multi-task foundation model trained end-to-end on 223 eCLIP-seq experiments spanning 150 RBPs to predict base-resolution RBP binding profiles directly from RNA sequence. This CLIP-seq pretraining strategy departs fundamentally from the masked-language-modeling paradigm, anchoring learned RNA representations directly in measured protein-RNA interactions rather than sequence statistics. Parnet substantially outperforms its single-task predecessor RBPNet in binding profile and motif recovery, generalizes to unseen cell types and iCLIP data, and recapitulates position-dependent splicing regulation. Frozen Parnet embeddings, without task-specific fine-tuning, match or exceed the performance of both task-specific tools, as well as larger self-supervised RNA and genomic language models across diverse downstream tasks, including RNA biotype classification, lncRNA chromatin localization, translational efficiency, splice-site recognition, intron retention, and non-coding variant effect prediction. Importantly, Parnet remains mechanistically interpretable, tracing predictions back to the specific RBPs and motifs that drive them. These results establish the RBP interactome as a compact, functionally sufficient, and interpretable basis for foundation model pretraining in RNA biology.
RNA sequencing (RNA-seq) provides a powerful complement to DNA sequencing for uncovering pathogenic defects affecting gene expression and splicing in individuals with genetically undiagnosed rare disorders. However, as large rare disease consortia adopt RNA-seq, challenges arise due to cohort heterogeneity, variability in tissues and sample sizes, and differences in interpretation practices. Here, we present a harmonized analytical and interpretation framework developed by the pan-European Solve-RD consortium to address these challenges. We analyzed 521 RNA-seq samples from whole blood, fibroblasts, muscle and peripheral blood mononuclear cells collected across more than 30 clinics and five European Reference Networks. Aberrant expression and splicing events were identified using OUTRIDER and FRASER 2.0 and analysed through a standardized four-level scoring framework that encompassed RNA-seq outlier reliability, phenotype relevance, variant mechanism, and segregation evidence, captured in structured reports for interpretation. Regular meetings, and collaborative Solvathon workshops were used to evaluate variant pathogenicity. This effort resulted in molecular diagnoses for 19 families out of 248 (7.7%) for whom DNA analyses had been inconclusive. Furthermore, three cases diagnosed using DNA analyses were confirmed, and 49 candidate events and five novel candidate disease genes were identified in the remaining families. Our results demonstrate the feasibility and impact of large-scale, standardized RNA-seq analysis in a transnational research setting. This framework provides a model for other international initiatives such as the Undiagnosed Diseases Network and ERDERA, paving the way for broader clinical implementation of transcriptome-based rare disease diagnostics. ### Competing Interest Statement V.A.Y., F.B., and C.M., are founders, shareholders and managing directors of OmicsDiscoveries. The other authors declare no competing interests. ### Clinical Trial NCT03491280 ### Funding Statement The SolveRD project has received funding from the European Union Horizon 2020 research and innovation programme under grant agreement No 779257. SolveRD research is supported (not financially) by ERN ITHACA (project ID no. 101085231), ERN RND (project ID no. 101155994), ERN EURO NMD (project ID no. 101156434), ERN EpiCARE (project ID 101156811), and ERN RITA (project ID 101155878). All ERNs are cofunded by the European Union within the framework of the Third Health Programme. ERDERA has received funding from the European Union Horizon Europe research and innovation programme under grant agreement 101156595. Views and opinions expressed are those of the author(s) only and do not necessarily reflect those of the European Union or any other granting authority, who cannot be held responsible for them. VAY, RL, CM and JG were supported by the Deutsche Forschungsgemeinschaft (German Research Foundation) via the project NFDI 1/1 GHGA German Human Genome Phenome Archive(441914366). The TUM IT infrastructure was cofunded via the Deutsche Forschungsgemeinschaft (German Research Foundation, project ID 461264291). BEA was supported by the predoctoral program Joan Oro of the Secretary of Universities and Research of the Department of Research and Universities of the Government of Catalonia with codes 2024 FI1 00075 and 2025 FI2 00075, cofinanced by the European Union. HM was supported by the Wellcome Trust grant 220906/Z/20/Z and UCL Global Engagement Fund scheme (2022/23 GEF project). HL receives support from the Canadian Institutes of Health Research (CIHR) for Foundation Grant FDN167281 (Precision Health for Neuromuscular Diseases), Transnational Team Grant ERT 174211 (ProDGNE) and Network Grant OR2 189333 (NMD4C), from the Canada Foundation for Innovation (CFI JELF 38412), the Canada Research Chairs program (Canada Research Chair in Neuromuscular Genomics and Health, 950 232279), the European Commission (101080249) and the Canada Research Coordinating Committee New Frontiers in Research Fund (NFRFG 2022 00033) for SIMPATHIC, and from the Government of Canada First Research Excellence Fund (CFREF) for the Brain-Heart Interconnectome (CFREF 2022 00007). KP is a recipient of a Canadian Institutes of Health Research (CIHR) postdoctoral fellowship award under award no: MFE 491707. JP was supported by the Else Kroener Fresenius Stiftung Clinician Scientist program precise.net and by the intramural TUFF program (3049 0 0). AP has received funding from the Secretariat for Universities and Research of the Ministry of Business and Knowledge of the Government of Catalonia (2021SGR00899), and the Instituto de Salud Carlos III (ISCIII) (FIS PI23/00835) Fondo Europeo de Desarrollo Regional (FEDER), Union Europea, una manera de hacer Europa. ASC was supported by the grants FPU20/06692 and EST23/00463 from Ministerio de Universidades, Spain. AH was supported by a ZonMW (The Netherlands Organization for Health Research and Development) Vici grant (No. 09150182310053). KL receives support from the German Research Foundation (DFG, LO1555/10 1). DNdB was supported by Instituto de Salud Carlos III (Grant CP22/00141). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The ethics committee of University Hospital of Tuebingen gave ethical approval for this work (ClinicalTrials.gov ID: [NCT03491280][1], https://clinicaltrials.gov/study/[NCT03491280][1]). Informed consent for data sharing, including indirect identifiers within Europe for research, was obtained from all recruited individuals. All data submitters confirmed the code of conduct of RD-connect GPAP. This study adheres to the principles set out in the Declaration of Helsinki. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Raw data will be available at the European Genome-Phenome Archive (https://ega-archive.org/datasets/) under the Solve-RD study EGAS00001003851, and can be accessed following approval from the Solve-RD Data Access Committee. [1]: /lookup/external-ref?link_type=CLINTRIALGOV&access_num=NCT03491280&atom=%2Fmedrxiv%2Fearly%2F2026%2F02%2F14%2F2026.02.10.26345954.atom
Abstract Small interfering RNAs (siRNAs) are a clinically validated therapeutic modality, yet designing potent chemically modified siRNAs remains a costly and iterative process, limited by scarce public data. Computational prediction of siRNA efficacy is therefore essential for rational design and accelerated preclinical development. However, despite the critical role of chemical modifications in therapeutic performance, current state-of-the-art machine learning methods either are not designed to model the chemical diversity of therapeutic siRNAs, or exhibit poor generalization performance. Here, we present FENNEC (Fine-Tuned Ensemble of Neural Networks for siRNA Efficiency Characterization), a machine-learning framework for predicting siRNA activity across chemically diverse design spaces. To support this effort, we curated the largest patent-derived dataset to date of chemically modified siRNAs from 42 patents using OCR-based table extraction and stringent filtering. FENNEC combines temporal convolutional networks with thermodynamic descriptors, experimental covariates, and embeddings from RNA foundation models to capture both local chemical determinants and broader target-context information. Importantly, we show that language-model-derived embeddings provide meaningful higher-order representations of target transcripts, particularly in data-scarce settings. FENNEC achieved robust predictive performance across both gene-level and scaffold-level validation settings, with additional experimental validation on a novel AHSA1-targeting dataset further supporting its generalizability across chemically modified siRNAs. In benchmarking, FENNEC outperformed classical machine-learning and state-of-the-art deep learning models, demonstrating generalization to unseen chemistry. Model interpretation recovered established design principles, including position-specific effects of glycol nucleic acid, 2’-fluoro modifications, and phosphorothioate backbones. Furthermore, in silico perturbation analyses suggest that FENNEC can serve not only as a predictive model, but also as an oracle for the design and optimization of chemically modified siRNAs. Together, our work addresses a key gap in the field by enabling chemically aware deep learning for siRNA design, supported by a large and diverse collection of chemically modified siRNA measurements.
Despite advances in genomic diagnostics, the majority of individuals with rare diseases remain without a confirmed genetic diagnosis. The rapid emergence of advanced omics technologies, such as long-read genome sequencing, optical genome mapping and multiomic profiling, has improved diagnostic yield but also substantially increased analytical and interpretational complexity. Addressing this complexity requires systematic multidisciplinary collaboration, as recently demonstrated by targeted diagnostic workshops. Here, we highlight the experience of the Solve-RD consortium, a pan-European initiative, in implementing four structured workshops, termed 'Solvathons', as a regular and effective component of its operational workflow. We provide actionable insights, best practices and lessons learned for successful data integration, expert training and scalable collaborative diagnostics within large research consortia.
Exome and genome sequencing leave >50% of dystonia-affected individuals without a molecular diagnosis. Where DNA-oriented approaches remain insufficient, integrating multiomics methods and bioinformatics is essential to advance genome interpretation. Herein, we incorporated RNA sequencing (RNA-seq) from a collection of 167 fibroblast samples from individuals affected with dystonic diseases. We leveraged an RNA-seq analysis pipeline, focused on the identification of expression and splicing aberrations, on RNA-seq from skin biopsies. We evaluated a “variant-positive” group of patient samples with preexisting information on variants (36/167, 21.6%), and a “variant-negative” group in which genomic sequencing alone had been unsuccessful in yielding a diagnostic candidate (78.4%). We found that at least 80% of dystonia-associated genes from databases were sufficiently detected by RNA-seq in fibroblasts, highlighting broad applicability. Expression and splicing aberration analyses then produced a manageable number of statistically significant RNA defects affecting dystonia-associated genes for effective case-by-case review. Our approach successfully detected RNA underexpression and mis-splicing for different types of pre-identified dystonia-related variants, providing both benchmarks and insights into mutational mechanisms. Applied to 131 samples from patients without candidate variants from exome and genome sequencing, RNA-seq aided the identification of previously unprioritized causative intronic alterations on reanalysis, providing an added diagnostic yield of 6.9% (9/131). For observed events, we also report the integration of new machine-learning scores predicting corresponding aberrant gene expression in the brain. Fibroblast-based RNA-seq in our selected cohort improved variant interpretation and enabled diagnoses missed by genomic analysis alone, suggesting this framework could be generalized to other dystonias.
Understanding how regulatory sequences shape gene expression across individual cells is a fundamental challenge in genomics. Joint RNA sequencing and epigenomic profiling provides opportunities to build models capturing sequence determinants across steps of gene expression. However, current models, developed primarily for bulk omics data, fail to capture the cellular heterogeneity and dynamic processes revealed by single-cell multimodal technologies. Here, we introduce scooby, a framework to model genomic profiles of single-cell RNA-sequencing coverage and single-cell assay for transposase-accessible chromatin using sequencing insertions from sequence at single-cell resolution. For this, we leverage the pretrained multiomics profile predictor Borzoi and equip it with a cell-specific decoder. Scooby recapitulates cell-specific expression levels of held-out genes and identifies regulators and their putative target genes. Moreover, scooby allows resolving single-cell effects of bulk expression quantitative trait loci and delineating their impact on chromatin accessibility and gene expression. We anticipate scooby to aid unraveling the complexities of gene regulation at the resolution of individual cells.
Despite the frequent implication of aberrant gene expression in diseases, algorithms predicting aberrantly expressed genes of an individual are lacking. To address this need, we compile an aberrant expression prediction benchmark covering 8.2 million rare variants from 633 individuals across 49 tissues. While not geared toward aberrant expression, the deleteriousness score CADD and the loss-of-function predictor LOFTEE show mild predictive ability (1-1.6% average precision). Leveraging these and further variant annotations, we next train AbExp, a model that yields 12% average precision by combining in a tissue-specific fashion expression variability with variant effects on isoforms and on aberrant splicing. Integrating expression measurements from clinically accessible tissues leads to another two-fold improvement. Furthermore, we show on UK Biobank blood traits that performing rare variant association testing using the continuous and tissue-specific AbExp variant scores instead of LOFTEE variant burden increases gene discovery sensitivity and enables improved phenotype predictions.
Post-translational modifications (PTMs) play a central role in cellular regulation and are implicated in numerous diseases. Database searching remains the standard for identifying modified peptides from tandem mass spectra but is hindered by the combinatorial expansion of modification types and sites. De novo peptide sequencing offers an attractive alternative, yet existing methods remain limited to unmodified peptides or a narrow set of PTMs. Here, we curated a large dataset of spectra from endogenous and synthetic peptides from ProteomeTools spanning 19 biologically relevant amino acid-PTM combinations, covering phosphorylation, acetylation, and ubiquitination. We used this dataset to develop Modanovo, an extension of the Casanovo transformer architecture for de novo peptide sequencing. Modanovo achieved robust performance across these amino acid-PTM combinations (median area under the precision-coverage curve 0.92), while maintaining performance on unmodified peptides (0.93), nearly identical to Casanovo (0.94). The model outperformed π-PrimeNovo-PTM and InstaNovo-P and showed increased precision and complementarity to the database search tool MSFragger. Robustness was confirmed across independent datasets, particularly at peptide lengths frequently represented in the curated dataset. Applied to a phosphoproteomics dataset from monkeypox virus-infected cells, Modanovo recovered numerous confident peptides not reported by database search, including new viral phosphosites supported by spectral evidence, thereby demonstrating its complementarity to database-driven identification approaches. These results establish Modanovo as a broadly applicable model for comprehensive de novo sequencing of both modified and unmodified peptides.
Developmental disorders constitute a major class of genetic diseases, yet tools to identify splicing-disruptive variants during development are lacking. To address this need we extended the AbSplice framework to incorporate splicing dynamics across developmental stages. Moreover, we introduced several improvements including a refined ground truth from the aberrant splicing caller FRASER2, a continuous representation of splice site usage, and integration of the rich set of predictions from the sequence-based model Pangolin. These advances double the precision and recall of the original model at predicting aberrant splicing events and enable predictions across embryonic, childhood, and adult tissues. Genome-wide scores for all single-nucleotide variants and a web interface to score indels are available to facilitate the exploration of the predictions. Our genome-wide predictions reveal a class of variants with splicing-disruptive effects confined to early development, particularly abundant in the brain (>18,000 variants). Analysis of genomes of individuals with a suspected Mendelian disorder from Genomics England’s National Genomics Research Library identified 26 unique variants in disease-linked genes, with stronger predicted effects during development than in adulthood, including a candidate new diagnosis in the gene FGFR1 . Altogether, these results improve the accuracy of splice-disruptive variant prediction and provide developmental context to aid interpretation. ### Competing Interest Statement The authors have declared no competing interest. Deutsche Forschungsgemeinschaft, 461264291, 441914366 Bundesministerium fM-CM-<r Bildung und Forschung (BMBF), 01KU2016B European Research Council, https://ror.org/0472cxd90, 101118521 European Union, 101156595 Wellcome Trust, https://ror.org/029chgv08, 220134/Z/20/Z
Aging is a major risk factor for neurodegeneration and is characterized by diverse cellular and molecular hallmarks. To understand the origin of these hallmarks, we studied the effects of aging on the transcriptome, translatome, and proteome in the brain of short-lived killifish. We identified a cascade of events in which aberrant translation pausing led to altered abundance of proteins independently of transcriptional regulation. In particular, aging caused increased ribosome stalling and widespread depletion of proteins enriched in basic amino acids. These findings uncover a potential vulnerable point in the aging brain’s biology—the biogenesis of basic DNA and RNA binding proteins. This vulnerability may represent a unifying principle that connects various aging hallmarks, encompassing genome integrity, proteostasis, and the biosynthesis of macromolecules.