Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disease with limited therapeutic options. Riluzole remains the only widely available disease-modifying treatment for ALS, yet its survival benefit is modest and likely to vary substantially between patients. Cytochrome P450 2D6 (CYP2D6), is a highly polymorphic enzyme that contributes to interindividual variability in the metabolism of many drugs. CYP2D6 is also expressed in the brain, and experimental and translational studies indicate that brain CYP2D activity can influence local metabolism of neuroactive compounds. Accordingly, CYP2D6 poor function variants have been examined as susceptibility modifiers in the development of other neurodegenerative diseases, including Parkinson's disease and Alzheimer's disease, with heterogenous evidence; however, the role of CYP2D6 in ALS has not been established.
Summary Genomics is adopting autonomous AI agents that interpret genomes from natural-language instructions faster than it is building the means to trust them. We report the first large-scale controlled evaluation of where, in an agentic genomic pipeline, correctness must reside for the system to be trustworthy at clinical scale. Using pharmacogenomics, a domain where errors are measurable and sometimes lethal, we benchmarked nine frontier large language models across 44,550 scored evaluations on 110 pharmacogenomic cases, and tested model interpretation of real star-allele diplotypes from more than 7,000 individuals in three ancestrally diverse populations. Trustworthiness proved to be a property of pipeline architecture, not of the model. Letting the model reason was stochastic and unsafe, and grounding it in the correct guidelines by retrieval paradoxically increased lethal-class errors. Encoding the validated decision logic as a versioned skill and executing it as code made the pharmacogenomic mapping exact, auditable and identical across models, confining all residual error to a single input-interpretation step. On individual genomes, unguarded model interpretation degraded along an ancestry gradient; execution removes this gradient from the clinical mapping, relocating it to the auditable completeness of the input caller. This establishes a generalisable, auditable architecture for trustworthy agentic genome interpretation at scale. Highlights Correctness must be executed, not reasoned or retrieved, to be trustworthy Retrieval raises phenotype accuracy yet increases lethal-class errors; skills do not Execution makes the clinical mapping exact and model-invariant; error stays at input A deterministic input caller is the predicted route to all-correct emitted answers In brief Corpas and colleagues show that trustworthy agentic genome interpretation comes not from making language models reason correctly about biology, but from confining them to interpreting input while versioned, validated skills do the reasoning as executed code. Across nine large language models and 110 pharmacogenomics cases, executing the skill makes the clinical mapping deterministic, auditable and model-invariant. Significance Genomics is adopting autonomous, language-model-mediated agents faster than it is building the standards needed to trust them. On a pharmacogenomic benchmark with lethal-class consequences, we show that an agent’s trustworthiness is not a property of the model but of how the agent is constrained: correctness must be moved out of the stochastic model into a versioned skill executed as code, with the model confined to interpreting heterogeneous input. This gives the field a transferable architecture for trustworthy agentic genome interpretation, a predicted route to deploying it so that every emitted answer is correct (execute the validated skill, call the input deterministically, and abstain on the irreducible residual), and a way to develop genomic skills as validated, executable, versioned units rather than prompts. Following a validation framework described elsewhere, we use clinical-grade to mean determinism, auditability, traceability to versioned components and population-invariant performance, all achieved under skill-constrained execution. We distinguish two senses of population performance: the executed clinical mapping is population-invariant by construction, verified across European, Latin American and East African origin individuals, whereas the model’s interpretation of real, ancestrally diverse diplotypes is not, degrading along an ancestry gradient, which is precisely why the mapping must be executed rather than reasoned. We do not claim full clinical validation, which would additionally require non-canonical inputs, real-world genomic and clinical data, human comparators and multi-site concordance.
Abstract Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease characterised by progressive motor neuron degeneration. Mutations in the SOD1 gene represent the second most common genetic cause of ALS (ALS), and distinct SOD1 missense variants present with markedly different clinical profiles. A4V leads to an aggressive form of the disease (median survival ∼1y), H46R confers a mild, slowly progressive course and I113T exhibits an intermediate phenotype. The molecular basis by which these mutations produce divergent clinical outcomes remains poorly understood. We performed extensive classical molecular dynamics simulations of wild-type SOD1 and the three ALS-associated variants in the apo monomeric state to attempt to investigate the mechanisms behind such phenotypic differences. Structural stability, global compactness, and conformational flexibility, as well as analysis of collective motions between residues and estimation of free energy, were assessed. The H46R, A4V, and I113T variants exhibited distinct dynamic behaviours, highlighting differences in structural stability, local flexibility, and intramolecular interactions. These findings suggest that specific structural regions may contribute differently to protein dysfunction and could represent key elements for understanding the relationship between molecular dynamic properties and the differing clinical severity associated with these variants. Most strikingly, H46R exhibited exceptional structural stability across every analytical level, the lowest global deviation, most attenuated local flexibility, strongest internal dynamic coordination, and the deepest, most confined free energy basins of any system examined. This convergent multi-layered evidence of structural restraint provides a compelling mechanistic basis for the mild and slowly progressive clinical course of H46R ALS, suggesting that enhanced conformational rigidity, rather than bulk destabilisation, is the defining biophysical feature of this variant, and that its pathogenic mechanism operates through a route fundamentally decoupled from the aggregation-driven toxicity that characterises the more aggressive SOD1-ALS mutations. mutations.
Background: Motor neuron disease (MND) is a fatal neurodegenerative condition with significant clinical heterogeneity that is incompletely captured by existing phenotype classifications based on onset site. Electronic health records (EHRs) contain detailed symptom documentation in clinical narratives that may enable data-driven discovery of clinically meaningful patient subgroups. Methods: We developed a natural language processing (NLP) pipeline using MedCAT to extract symptoms from clinical notes of 2,361 people with a confirmed diagnosis of MND at a tertiary neurology center. MND cohort confirmation used three complementary methods: clinic attendance records, text-based diagnosis detection, and NLP extraction with negation detection. Extracted symptoms were filtered to Unified Medical Language System semantic type T184 (Sign or Symptom) with removal of negated concepts. Patients were clustered using latent class analysis on binary symptom profiles. Survival differences were assessed using Kaplan-Meier analysis, log-rank tests, and Cox proportional hazards regression. Results: From the first clinical notes, we identified four clusters of symptoms among 872 patients and 76 symptoms: Motor-Bulbar (n=373), Motor-Tremor (n=154), Sensory-Pain (n=222), and Motor-Respiratory (n=123). When extended to all clinical notes (n=2,065; 184 symptoms), these reorganized into three clusters: Autonomic-Respiratory (n=472), Nocturnal-Respiratory (n=338), and Classic Motor (n=1,255). Survival differences were significant across all clusters in both the first notes and all notes analyses (log-rank p < 0.001). Conclusions: NLP-based symptom extraction from EHRs identifies clinically meaningful MND subgroups that extend beyond traditional onset-site classifications. Autonomic-respiratory symptom burden is associated with poorer survival while a newly identified Sensory-Pain subtype with a better prognosis. These data-driven phenotypes may improve prognostication and inform targeted supportive care. ### Competing Interest Statement The authors have declared no competing interest. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The KERRI committee of King's College Hospital gave ethical approval for this work. The datasets used in your study had been de-identified prior to use in this study. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Clinical data cannot be shared due to patient confidentiality. EPSRC, EP/Y035216/1
Abstract The clinical and molecular heterogeneity observed in amyotrophic lateral sclerosis (ALS) presents a challenge for diagnosis, prognosis, and treatment. RNA sequencing of post-mortem brain samples from ALS patients has identified several subtypes with distinct molecular signatures. We sought to evaluate these subtypes across diverse tissues and datasets and assess the feasibility of supervised machine learning models for sample classification. Unsupervised clustering and pathway analysis were performed to confirm the presence of ALS subtypes in motor cortex samples. Three machine learning strategies were then used to create models based on post-mortem motor cortex expression data of 112 people with ALS from the London Neurodegenerative Diseases Brain Bank. These models were subsequently improved through feature selection and evaluated in independent cohorts from motor cortex (n = 257, NYGC ALS Consortium) and blood (n = 96, Macquarie University Neurodegenerative Disease Biobank) samples. Multi-class linear discriminant analysis (LDA) models were then used for subtype classification. Clustering of ALS post-mortem motor cortex samples confirmed the presence of three subtypes: neuroinflammation (ALS-Neu), extracellular matrix organisation and muscle contraction (ALS-OxA), and synaptic and neuropeptide signalling (ALS-SNs). Among all machine learning strategies, random forests produced the most accurate and stable models for binary classification (∼93% accuracy across the three subtypes). After feature selection, random forest models were able to classify samples from an independent post-mortem motor cortex cohort in their respective subtypes (AUC of ∼0.98 across the three subtypes). When these models were evaluated in blood using LDA, we found consistent clustering patterns, with samples aligning in the same subtype regions of the post-mortem motor cortex samples, with ALS-SNs being the subtype in which samples were classified with the highest confidence (LDA class probability ∼86%). Moreover, classification for this subtype improved when blood samples were collected closer to death. Our findings support the presence of three gene expression-based ALS subtypes in motor cortex samples and the utility of machine learning strategies for subtype classification. We also observed that the subtypes identified in the brain partially match those in the blood, with samples from the late stages of the disease more likely to be correctly predicted into the ALS-SNs cluster. This suggests a longitudinal effect in subtype identification that requires further investigation.
Abstract Amyotrophic Lateral Sclerosis (ALS) is a neurodegenerative disease characterised by progressive motor neuron loss and corticospinal tract degeneration. The genetic landscape of ALS is complex, with increasing recognition of shared genetic and phenotypic features with other neurodegenerative conditions, particularly those involving repeat expansions. Given that repeat expansions in disorders like spinocerebellar ataxia type 27B (SCA27B), caused by an intronic GAA repeat expansion in Fibroblast Growth Factor 14 ( FGF14 ), are recognised to extend beyond cerebellar ataxia with frequent pyramidal signs, we hypothesised that FGF14 repeat expansions might also contribute to ALS and degeneration of corticospinal pathways, and sought to investigate whether repeat length is associated with clinical phenotype. We screened 62 individuals with ALS using PacBio HiFi long-read whole-genome sequencing and compared repeat-size distributions with 256 healthy controls from the Human Pangenome Reference Consortium. Repeat expansions were confirmed using flanking PCR and repeat-primed PCR. We identified pathogenic-range FGF14 GAA ≥250 expansions, the established threshold for SCA27B, in 3/62 ALS cases (4.8%) and none in controls. Further analysis revealed that GAA expansions ≥200 repeats were enriched in ALS compared to controls (8.1% vs 0.4%; p = 0.0013), suggesting a broader pathogenic spectrum for FGF14 GAA repeats in ALS. In contrast, GAAGGA expansions were not significantly associated. Expanded pure GAA alleles were predicted to form triplex (H-DNA) structures, with the repeat-containing isoform (1B) being the predominant FGF14 transcript in motor neurons. These findings demonstrate that FGF14 GAA repeat expansions extend into the motor neuron disease spectrum.
Human endogenous retrovirus-K (HERV-K) reactivation is increasingly implicated in amyotrophic lateral sclerosis (ALS), with ongoing clinical trials investigating antiretroviral therapies. However, there is limited understanding of how HERV-K is trafficked in peripheral biofluids, and the role of exosomes, nano-sized extracellular vesicles, in this process remains largely unexplored. Exosomes offer a stable and cell-specific cargo reservoir that may reflect central pathogenic processes and serve as a minimally invasive biomarker source. In this study, we isolated plasma-derived exosomes from ALS patients (n = 21) and healthy controls (n = 16), and quantified exosomal HERV-K gag, env, and pol transcript levels using SYBR Green qPCR with RNase treatment and normalization to both traditional and exosome-enriched reference genes. HERV-K pol expression was significantly elevated in ALS, with fold-changes ranging from 1.59 to 1.85 (P = 0.037–0.051). env and gag also showed increased expression, though with greater variability. Normalization to the exosome-specific gene SOD2 provided the most consistent signal. These findings suggest that exosomal HERV-K transcripts, particularly pol, could serve as accessible biomarkers for patient stratification and treatment monitoring in HERV-K–targeted ALS trials. This work establishes proof-of-concept for using exosomal cargo to track endogenous retroviral activity in neurodegeneration and supports further investigation of liquid biopsy approaches in ALS precision medicine.
A variety of common and rare genetic factors have been implicated in the development of amyotrophic lateral sclerosis (ALS), and the evidence is that a genetic component is present in most affected individuals. However, our current understanding of ALS genetics causally explains only a small proportion of sporadic ALS, which accounts for over 90
Abstract Background Differential expression analysis is a central tool for studying the biological processes altered in human diseases via transcriptomic signatures. However, transcriptomic datasets are systematically confounded by latent variables from two distinct sources: unmeasured technical and biological heterogeneity within the expression data, and expression differences driven by population stratification. Correction using expression-based surrogate variables (SVs) and genotype-based principal components (PCs) addresses these sources independently, yet no study has directly evaluated their combined use against either method alone within a differential expression framework. In this study we hypothesised that simultaneously including both correction layers would produce more biologically valid and reproducible results than either approach alone, and tested this in two independent RNA-seq datasets of amyotrophic lateral sclerosis (ALS) cases and controls with matching genotype data. Results Four nested differential expression models (corrected for PC-only, SV-only, both SV and PC, and neither PCs nor SVs) were evaluated across the KCLBB (96 cases and 52 controls) and ALS Consortium (272 cases and 35 controls) datasets. Models were evaluated on: cross-dataset effect size concordance, cross-dataset replicability quantified by the Jaccard Similarity Index, and biological recall against a curated reference set of 66 known ALS genes. The combined SV+PC framework consistently outperformed simpler models across all metrics. Replicability improved nearly ten-fold compared to the non-corrected model, (Jaccard index: 2.28% to 19.5%), and the combined framework exhibited a statistically significant 2.1% gain over the SV-only model. The biological recall ALS genes recovered doubled comparing to the SV correction alone. Crucially, effect size stability was preserved, with the combined model expanding the shared transcriptomic signal without sacrificing consistency. These findings remained generally robust to PC number in sensitivity analyses. Conclusions This study found that SVs and genotype PCs address non-redundant sources of confounding, and we recommend their combined use as standard practice in differential expression analysis where matched genotype data are available. Notably PCs capturing population structure can also be derived directly from RNA-seq data, extending the applicability of this framework to studies lacking matched genotype data. Although this analysis was restricted to ALS datasets, we expect these findings to generalise to other traits.
Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2× to 3× , with strong statistical separation (p < 0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.
Amyotrophic lateral sclerosis (ALS) is a heritable disorder where rare variants with low-to-moderate penetrance are thought to dominate genetic risk. To identify such rare variants, we harmonized and analyzed exome data from 22 cohorts, totaling 17,919 individuals with ALS and 200,703 controls across discovery and replication phases. Rare variant analyses identified several new risk genes, with replication confirming association of YKT6 and supporting HTR3C, GBGT1 and KNTC1. We also provide strong, independent validation for genes with limited previous evidence: ARPP21, DNAJC7 and CFAP410. Notably, in ARPP21, we identified a new high-effect variant (p.P747L) and confirmed that p.P563L is an ALS-associated variant leading to an aggressive disease course. Beyond new discoveries, our analyses largely recapitulated the known genetic architecture of ALS, identifying risk variants in over 20% of cases and supporting a cumulative oligogenic risk model. These findings highlight new translational targets and show that rare variant analyses capture substantially more genetic risk than common variant genome-wide association studies.
Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease characterised by considerable heterogeneity in both its underlying biological mechanisms and clinical presentation. High-dimensional transcriptomic datasets offer an opportunity to characterise this variation at the molecular level; however, traditional statistical methods struggle with their scale and complexity. Machine learning approaches can reduce dimensionality and uncover latent patterns, enabling the identification of molecular subtypes that may refine prognosis and support patient stratification. Recent transcriptomic studies employing unsupervised machine learning have identified ALS subtypes with distinct molecular and clinical characteristics. Redefining ALS into more homogeneous molecular and clinical subtypes could transform all areas of ALS research by supporting novel experimental designs and precision medicine approaches. In this review, we summarise and critically assess these studies, discussing their findings, strengths, and limitations, and highlighting research gaps and challenges that must be addressed to enable their translation into biomedical and clinical practice.
Abstract Drug repurposing offers a practical strategy to identify new therapeutic uses for approved drugs, potentially reducing the time and cost associated with conventional drug development. We present a novel three-stage drug repurposing pipeline that integrates knowledge graph-based gene prediction, network-based drug-disease association analysis, and systematic classification of candidate drugs by therapeutic class. The pipeline integrates DGLinker to predict novel disease-associated genes, SAveRUNNER to identify drug repurposing candidates, and ATC Category Enrichment Analysis (ATCEA) to prioritise candidates by pharmacological class. We benchmarked the pipeline across twelve diseases using DrugBank and MEDI2-HPS as validation resources. Utilising DGLinker-expanded disease-gene sets as input increased the number of predicted repurposed drugs, while overall discriminative performance remained stable across diseases (AUROC 0.71-0.77). Application of ATCEA consistently improved precision, F1-score, and specificity, while reducing recall, reflecting a conservative prioritisation strategy that contracts the candidate space while retaining pharmacologically coherent drug-disease candidates. We further applied the pipeline to amyotrophic lateral sclerosis (ALS), a neurodegenerative disease with limited therapeutic options and performed a deeper literature-based validation of the results. Incorporation of DGLinker-predicted genes substantially increased the number of significant candidate drugs and uncovered enriched ATC categories not identified using known ALS genes alone, including antidepressants and antipsychotics. Moreover, several drugs with supporting evidence available in literature were identified only when DGLinker-predicted genes were used. Overall, 77 candidate drugs were prioritised within significantly enriched ATC categories, several of which are supported by previously published studies. To provide exploratory real-world support for these findings, we further evaluated candidate drugs in a longitudinal electronic health record (EHR) dataset of 2361 patients with ALS from King’s College Hospital. Although the number of evaluable drugs was limited due to sample size, the EHR analysis provided additional clinically relevant context for selected prioritised drugs and pharmacological classes. Our pipeline demonstrates potential to accelerate drug repurposing by integrating complementary computational approaches to each step of the process, providing an end-to-end framework that showed robust performance across benchmarking experiments and use cases. Graphical abstract
ATXN2 expansions of ≥33 CAG-repeats are associated with spinocerebellar ataxia type 2, while intermediate length expansions have been associated with amyotrophic lateral sclerosis (ALS). Yet, no consensus is established regarding the lengths that define a true association with ALS risk, with recent studies debating between a lower limit of ≥29 or ≥31 repeats. Here, we assessed the risk of ALS imparted by various ATXN2 repeat lengths to establish an accepted lower limit of repeats that impart risk of disease in the largest meta-analysis to-date. We identified 19 studies with carrier counts of expansions ranging from 24 to ≥34 repeats in cohorts of individuals with ALS and controls that we meta-analysed with the ATXN2 repeat lengths of the large-scale Project MinE ALS Consortium dataset (total individuals with ALS = 19202; total controls = 22177) and determined a lower limit of 30 repeats defining significant ALS risk. These findings were validated with a secondary assessment of the individuals with ALS captured within the meta-analysis using the gnomAD short tandem repeat dataset as a proxy control cohort. We also applied our defined ATXN2 repeat risk threshold to explore relationships with ALS clinical outcomes. While we did not observe a significant relationship between ATXN2 repeat lengths and age of ALS onset, we did identify a significant inverse correlation between ATXN2 repeat lengths as a continuous metric and duration of disease and found that individuals with ALS carrying the risk variant allele of ≥30 repeats had significantly shorter times to diagnosis than those without the repeat expansion. Our comprehensive analyses propose a lower-limit threshold of ≥30 ATXN2 trinucleotide repeats in length defining true ALS risk. These findings are imperative for allowing improved accuracy in risk interpretation and guidance for patients and their families, particularly as clinical genetic testing efforts continue to expand and there becomes increased need for guiding targeted clinical trial inclusion criteria. Our results may also aid in future analyses assessing ATXN2 pathogenic mechanisms and therapeutic strategies. ### Competing Interest Statement AA-C declares contracts with the MRC, NIHR and Darby Rimmer Foundation; consulting fees from Amylyx, Apellis, Biogen, Brainstorm, Clene Therapeutics, Cytokinetics, GenieUs, GSK, Lilly, Mitsubishi Tanabe Pharma, Novartis, OrionPharma, Quralis, Sano, and Sanofi). AAK declares contracts with the MRC ((MR/Z505705/1), the Motor Neurone Disease Association (MNDA), National Institute for Health and Care Research (NIHR) Maudsley Biomedical Research Centre, Amyotrophic Lateral Sclerosis (ALS) Association Milton Safenowitz Research Fellowship, Darby Rimmer MND Foundation, LifeArc, and the Dementia Consortium; equipment by NIHR Maudsley Biomedical Research Centre; and consulting fees from the UK National Endowment for Science, Technology and the Arts (NESTA). ### Funding Statement This study did not receive any funding. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: All Project MinE Sequencing Consortium participants provided written informed consent with study ethnical approval from the institutional review board of the University Medical Center Utrecht I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The data sets used that support the findings in this study are available from The Project MinE consortium public repository. To gain access to the data, an account request must be made to info{at}projectmine.com. Data access will require the completion of a data access request. Further information about data access can be found at https://www.projectmine.com/research/data-sharing/.
Background Despite several studies suggesting a potential oligogenic risk model in amyotrophic lateral sclerosis (ALS), case–control statistical evidence implicating oligogenicity with disease risk or clinical outcomes is limited. Considering its direct clinical and therapeutic implications, we aim to perform a large-scale robust investigation of oligogenicity in ALS risk and in the disease clinical course. Methods We leveraged Project MinE genome sequencing datasets (6711 cases and 2391 controls) to identify associations between oligogenicity in known ALS genes and disease risk, as well as clinical outcomes. Results In both the discovery and replication cohorts, we observed that the risk imparted from carrying multiple ALS rare variants was significantly greater than the risk associated with carrying only a single rare variant, both in the presence and absence of variants in the most well-established ALS genes. However, in contrast to risk, the relationships between oligogenicity and ALS clinical outcomes, such as age of onset and survival, did not follow the same pattern. Conclusions Our findings represent the first large-scale, case–control assessment of oligogenicity in ALS and show that oligogenic events involving known ALS risk genes are relevant for disease risk in ~6% of ALS but not necessarily for disease onset and survival. This must be considered in genetic counselling and testing by ensuring to use comprehensive gene panels even when a pathogenic variant has already been identified. Moreover, in the age of stratified medication and gene therapy, it supports the need for a complete genetic profile for the correct choice of therapy in all ALS patients.
IntroductionMotor neuron disease (MND), also known as amyotrophic lateral sclerosis (ALS), is a progressive neurodegenerative disorder characterized by motor neuron degeneration, leading to muscle weakness, paralysis, and eventual respiratory failure. Despite advances in understanding its pathology, effective therapies remain limited, underscoring the need for reliable biomarkers to aid early diagnosis, monitor disease progression, and optimize clinical trials. This systematic review explores the role of biomarkers in ALS, focusing on their application in clinical trials to accelerate therapeutic development and enhance patient care.MethodsA comprehensive search of PubMed, EMBASE, MedLine, and Google Scholar identified 93 studies investigating various biomarkers, including neurofilament light chain (NFL), inflammatory markers, genetic markers like SOD1 and C9orf72, and imaging modalities.ResultsNFL emerged as a robust biomarker, strongly correlating with disease progression and therapeutic response, and was frequently used in trials like RESCUE-ALS and CENTAUR. Genetic biomarkers, such as C9orf72 and SOD1 mutations, provided insights into ALS mechanisms and informed targeted therapeutic approaches. Emerging biomarkers, such as retroviral elements, show potential but require further validation. Included studies span key trials such as Lighthouse-II, MIROCALS, and MND-SMART.DiscussionThis systematic review evaluates which biomarkers are currently validated for monitoring disease progression and therapeutic response in ALS clinical trials, including protein, genetic, inflammatory, metabolic, and imaging markers. It also highlights the critical role of biomarkers in advancing MND clinical trials by enabling adaptive trial designs, patient stratification, and the use of surrogate endpoints, thereby reducing trial duration and improving efficiency. The review also highlights the translational gap between biomarker discovery and clinical application, emphasizing their potential to optimize trial design and patient stratification. While biomarkers like NFL have transformed trial methodologies, challenges such as disease specificity and inter-patient heterogeneity persist. Future efforts should focus on multimodal biomarker approaches to achieve comprehensive disease assessment and advance personalized therapeutic strategies, ultimately improving outcomes for patients with MND.
A variety of common and rare genetic factors have been implicated in the development of amyotrophic lateral sclerosis (ALS), and the evidence is that a genetic component is present in most affected individuals. However, our current understanding of ALS genetics causally explains only a small proportion of sporadic cases which represent over 90% of all people with ALS. This limits the utility of genetic testing in screening, diagnosis and management to the 15-20% of people with ALS who carry a known pathogenic variant. Capsule Networks (CapsNets) constitute a deep learning method that has demonstrated strong performance in using genotyping data to predict individuals at risk for ALS. However, their use is constrained by a lack of generalised, flexible, and validated implementations across comprehensive datasets that account for the technical, biological, and clinical heterogeneity found in real-world disease scenarios. In this study, we build upon this method to address existing limitations, to develop a new model that is validated across diverse ALS populations, can handle discrepancies between genotyping technologies, and is applicable to individual external samples. Using large-scale datasets from over 47,000 individuals from 13 countries, genotyped with nine different genotyping platforms, our model achieved high precision and sensitivity in distinguishing between individuals with ALS and non-affected controls. Moreover, in simulations of population screening for ALS, its performance was comparable to that of conventional genetic screening for known ALS gene mutations, such as FUS and C9orf72. Our results demonstrate that this flexible and validated method could support the development of a genetic screening test for identifying individuals at risk and expediting ALS diagnosis. This would be applicable to all individuals, regardless of their family history or presence of known ALS mutations. ### Competing Interest Statement VS received compensation for consulting services and/or speaking activities from AveXis, Cytokinetics, Italfarmaco, Liquidweb S.r.l., Amylyx, Novartis Pharma AG, Zambon Biotech SA, and Biogen. VS is in the Editorial Board of Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, European Neurology, American Journal of Neurodegenerative Diseases, Frontiers in Neurology, and Exploration of Neuroprotective Therapy. AAC reports receiving nonfinancial support from the National Institute for Health and Care Research (NIHR); consultant fees from Amylyx, Clene Therapeutics, GenieUs, GSK, Eli Lilly, Mitsubishi Tanabe Pharma, Novartis, OrionPharma, Quralis, SanoGenetics, Sanofi, Voyager Therapeutics, and Wave Pharmaceuticals; and having a patent for use of CSF-neurofilament determinations and CSF-neurofilament thresholds of prognostic and stratification value with regards to response to therapy in neuromuscular and neurodegenerative diseases pending. ### Funding Statement This is an EU Joint Programme-Neurodegenerative Disease Research (JPND) project. The project is supported through the following funding organisations under the aegis of JPND http://www.neurodegenerationresearchneurodegenerati onresearch.eu/ (UK, Medical Research Council (MR/L501529/1 and MR/R024804/1) and Economic and Social Research Council (ES/L008238/1). AA-C is an NIHR Senior Investigator. AA-C receives salary support from the National Institute for Health and Care Research (NIHR) Dementia Biomedical Research Unit at South London and Maudsley NHS Foundation Trust and King's College London. The work leading up to this publication was funded by the European Community's Health Seventh Framework Program (FP7/2007-2013; grant agreement number 259867) and Horizon 2020 Program (H2020-PHC-2014-two-stage; grant agreement number 633413). This project has received funding from the European Research Council (ERC) under the European Union's Horizon 2020 Research and Innovation Programme (grant agreement no. 772376-EScORIAL. This study represents independent research part funded by the NIHR Maudsley Biomedical Research Centre at South London and Maudsley NHS Foundation Trust and King's College London. AI is funded by South London and Maudsley NHS Foundation Trust, MND Scotland, Motor Neurone Disease Association, National Institute for Health and Care Research, Spastic Paraplegia Foundation, Rosetrees Trust, Darby Rimmer MND Foundation, the Medical Research Council (UKRI), LifeArc, and Alzheimer's Research UK. Project MinE Belgium was supported by a grant from IWT (n 140935), the ALS Liga Belgie, the National Lottery of Belgium and the KU Leuven Opening the Future Fund. AAK is funded by The Motor Neurone Disease Association (MNDA), NIHR Maudsley Biomedical Research Centre and ALS Association Milton Safenowitz Research Fellowship, the Darby Rimmer MND Foundation, LifeArc, and the Dementia Consortium. AAK is supported by the UK Dementia Research Institute through UK DRI Ltd, principally funded by the Medical Research Council. VS Receives or has received research supports from the Italian Ministry of Health, AriSLA, E-Rare Joint Transnational Call, and the ERN Euro-NMD. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The datasets used in your study were individual-level data all individual-level data had been de-identified. All data used in this study are publicly available. The ALS genome data and GWAS data used in this study are from Project MinE and can be accessed via online application (www.projectmine.com). Other data utilized in this study include the following: the Wellcome Trust Case Control Consortium (https://www.wtccc.org.uk/) and dbGaP datasets (phs000101.v3.p1, phs000101.v3.p1, phs000101.v3.p1, phs000101.v3.p1, phs000126.v1.p1, phs000196.v1.p1, phs000344.v1.p1, phs000344.v1.p1, phs000344.v1.p1). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data used in this study are publicly available. The ALS genome data and GWAS data used in this study are from Project MinE and can be accessed via online application (www.projectmine.com). Other data utilized in this study include the following: the Wellcome Trust Case Control Consortium (https://www.wtccc.org.uk/) and dbGaP datasets (phs000101.v3.p1, phs000101.v3.p1, phs000101.v3.p1, phs000101.v3.p1, phs000126.v1.p1, phs000196.v1.p1, phs000344.v1.p1, phs000344.v1.p1, phs000344.v1.p1).
Oxford Nanopore Technologies (ONT) long-read sequencing (LRS) has emerged as a promising genomic analysis tool, yet comprehensive benchmarks with established platforms across diverse datasets remain limited. This study aimed to benchmark LRS performance against Illumina short-read sequencing (SRS) and microarrays for variant detection across different genomic contexts and to evaluate the impact of experimental factors. We sequenced 14 human genomes using the three platforms and evaluated single nucleotide variants (SNVs), insertions/deletions (indels), and structural variants (SVs) detection, stratifying by high-complexity, low-complexity, and dark genome regions while assessing effects of multiplexing, depth, and read length. LRS SNV accuracy was slightly lower than that of SRS in high-complexity regions (F-measure: 0.954 vs. 0.967) but showed comparable sensitivity in low-complexity regions. LRS showed robust performance for small (1–5 bp) indels in high-complexity regions (F-measure: 0.869), but SRS agreement decreased significantly in low-complexity regions and for larger indel sizes. Within dark regions, LRS identified more indels than SRS, but showed lower base-level accuracy. LRS identified 2.86 times more SVs than SRS, excelling at detecting large variants (>6 kb), with SV detection improving with sequencing depth. Sequencing depth strongly influenced variant calling performance, whereas multiplexing effects were minimal. Our findings provide valuable insights for optimising LRS applications in genomic research and diagnostics.