Abstract Background Gene fusions are critical drivers of oncogenesis and diagnostic biomarkers in various cancers. However, their detection from RNA or DNA sequencing, when performed using traditional analytical methods, encounters challenges related to sample quality, computational complexity, and noise. Although deep learning is more robust, it usually requires large labeled datasets and substantial training resources. Genomic foundation models (GFMs), which are pre-trained on pangenome-scale data, offer a promising solution to these issues. Methods This study presents the first comprehensive benchmark of four transformer-based GFMs, Nucleotide Transformer (NT), Evo2, HyenaDNA, and DNABERT2, for the classification of gene fusion breakpoints. Using the curated FusionAI dataset of ~ 52,000 sequences, we extracted embeddings from 10-kilobase-pair (kbp) DNA sequences surrounding fusion breakpoints. We evaluated the quality of these representations qualitatively using t-SNE visualization and quantitatively by training lightweight classifiers (Support Vector Machines and simple Neural Networks) on the fixed embeddings. Results NT achieved the best performance with an accuracy of 0.967 and an F1 score of 0.967. This result outperformed the dedicated deep learning baseline (FusionAI, with an accuracy of 0.894). Evo2 was the second-best performer (accuracy: 0.920), demonstrating robustness derived from evolutionary pretraining. Conversely, DNABERT2 failed to compete (accuracy 0.677–0.723). Furthermore, sample efficiency analysis revealed that NT required only ~ 2,600 samples to reach 95% of its peak performance, whereas the baseline required over 14,000 samples. Conclusions These findings demonstrate that advanced GFMs, particularly the NT and Evo2 models, generate highly discriminative ‘out-of-the-box’ embeddings. These embeddings significantly outperform dedicated deep learning baselines while requiring a fraction of the training data and computational time. This suggests that GFMs could be a scalable, data-efficient way of developing precise genomic diagnostic tools, particularly for rare diseases.
Background:Aspiration pneumonia is a leading cause of death in Parkinson disease (PD). Expiratory muscle strength training (EMST) is a promising intervention for respiratory and swallowing dysfunction. However, long-term EMST adherence is frequently poor in PD. Objective:This study aims to determine whether mobile health (mHealth)-assisted EMST with the SpiroGym app (Czech Technical University) improves long-term adherence and physiological outcomes versus conventional EMST among participants at risk for nonadherence. Methods:In this single-center, parallel, phase 2 randomized controlled trial, 75 individuals with PD were randomized 1:1 to conventional EMST (control; n=38) or the same protocol enhanced with the SpiroGym app (experimental; n=37), using a simple computer-generated randomization sequence. The SpiroGym is an mHealth app that provides real-time performance monitoring, direct visual feedback, and longitudinal progress tracking. All participants completed 8 weeks of semisupervised intensive EMST with biweekly in-person reassessments, followed by 16 weeks of unsupervised maintenance training. The primary outcome was adherence during weeks 8 to 24 among participants at risk for nonadherence, defined a priori at week 8 as Self-Efficacy for Home Exercise Program Scale (SEHEPS) less than 59. Because risk status was determined at week 8 and all participants subsequently entered the unsupervised phase, individuals not classified as at-risk were not excluded. Their data from week 8 onward were reported alongside the at-risk group. Secondary outcomes were changes in maximum expiratory pressure and SEHEPS. Results:No study-related adverse events occurred. Groups were well matched at baseline (control vs experimental: mean disease duration 7.0 (SD 5.7) vs 7.3 SD 4.7) y; mean Hoehn-Yahr 1.97 (SD 0.6) vs 2.0 (SD 0.5)). The mixed-effects model showed no significant 3-way interaction (group×interval×SEHEPS risk; P=.14). At week 24, the at-risk category for the nonadherence cohort comprised 34 participants (control, n=17; experimental, n=17). In this at-risk cohort, the experimental group demonstrated a smaller decline in adherence during weeks 8 to 24 than controls (β=496.9, 95% CI 130.7-863.3; P=.008), completing 1073 (95% CI 643-1502) expiratory maneuvers versus 525 (95% CI 358-692). Maximum expiratory pressure increased in both groups from weeks 0 to 24, with larger gains in the experimental group (+43.1, 95% CI 32.4-53.8 cmH₂O) than in controls (+22.8, 95% CI 13.8-31.8 cmH₂O; P=.006; Cohen d=0.74). SEHEPS improved after intensive training in both groups, but only the experimental group exceeded the 12-point minimal detectable change at the 95% confidence limit. Conclusions:This is the first randomized controlled trial to integrate mHealth with EMST. Unlike prior studies in the EMST field, we focused on sustaining long-term exercise adherence. SpiroGym-assisted EMST resulted in higher long-term adherence and greater gains in expiratory muscle strength than conventional EMST. In real-world PD care, assessing self-efficacy after the supervised EMST phase may help identify individuals who would benefit from digital support, making mHealth-assisted EMST a practical approach for maintaining exercise adherence.
Tremor is the most prevalent human movement disorder, characterized by rhythmic oscillations of a body part. Accurate tremor assessment is essential for diagnosis, monitoring, and treatment evaluation. Traditional methods rely on accelerometry-based measurements, requiring direct sensor attachment, which may be impractical in some settings. We developed a novel algorithm for detecting tremors from video recordings based on the motion of the center of mass and implemented it in the open-source software TremAn3. Motion data were extracted from 2D video recordings of both hands and the head, and spectral analysis was then performed to quantify the tremor by calculating peak tremor power and peak power frequency. A total of 30 videos were recorded from 30 participants with essential or dystonic tremors. Simultaneously, acceleration signals were collected using inertial measurement units (IMUs) placed on the backs of the hands and forehead as a gold-standard reference. Agreement between video- and IMU-derived metrics was assessed using intraclass correlation coefficients (ICCs) and mean absolute error (MAE). For PP, video-based estimates showed moderate-to-good agreement (ICC: 0.70 left hand, 0.77 right hand, 0.80 head) with MAE of 8.12–10.80 dB. For PPF, agreement was moderate for the hands (ICC: 0.60 left, 0.67 right; MAE: 0.54–0.76 Hz) but poor for head PPF (ICC: 0.08; MAE: 2.06 Hz). Our results indicate that video analysis can serve as a viable alternative to traditional accelerometry for tremor quantification. This contactless method holds significant potential for telemedicine and research applications.
The rapid expansion of large-scale genomic variant datasets necessitates efficient database architectures for querying and programmatic access. This study systematically compares six database systems—four NoSQL document stores (MongoDB, Elasticsearch, RavenDB, CouchDB) and two PostgreSQL relational schemas—evaluating performance on typical variant queries using whole genome sequencing data. Through twelve representative scenarios including point lookups, range queries, and complex filters, MongoDB achieved the fastest median response times ( 2.9 ms), significantly outperforming relational approaches by two orders of magnitude. Elasticsearch showed strength in queries searching nested annotation arrays representing variant identifiers (rsIDs), leveraging its inverted indexing for efficient lookup while CouchDB exhibited performance bottlenecks for nested array operations. This work highlights critical considerations in schema design, query optimization, and API implementation for scalable variant data management, offering practical insights for genomic data infrastructure development and precision medicine applications.
Assessing muscle activity is essential for diagnosis and treatment of movement disorders such as dystonia and spasticity. While task-based muscle functional magnetic resonance imaging (m-fMRI) enables non-invasive imaging of muscle activation, conventional methods rely on comparisons between rest and activity, which are unsuitable for patients with sustained muscle contractions. This pilot study introduces a resting-state muscle fMRI (rs-m-fMRI) approach based on regional homogeneity (ReHo) to evaluate muscle activity from spontaneous BOLD fluctuations during sustained isometric contraction without block-design contrasts. Eight healthy male participants performed separate isometric plantar and dorsal foot flexion tasks during 3 T MRI scanning. rs-m-fMRI data were analyzed using ReHo to assess local synchronization of BOLD signal. Calf muscle activation was quantified as the percentage of suprathreshold z-transformed ReHo voxels within each segmented muscle and activation thresholds were derived via ROC analysis. ROC analysis demonstrated moderate discrimination between expected active and inactive muscle regions (AUC = 0.63), with sensitivity of 0.57 and specificity of 0.61 at the selected threshold. Consistent condition-related differences were observed between active and inactive muscles during both conditions, with a higher percentage of suprathreshold zReHo voxels in voluntarily contracted muscles. This pilot study demonstrates the feasibility of detecting contraction-related ReHo differences using rs-m-fMRI from a single continuous acquisition. The activation threshold was internally calibrated using expected agonist and antagonist muscle groups and therefore does not represent an externally validated classifier. Further studies incorporating independent physiological validation, reproducibility assessment and larger patient cohorts are required before clinical translation.
Speech impairments affect up to 90
Large language models (LLMs) are increasingly explored as tools for healthcare research and data analysis. However, their applicability to structured public health datasets, especially in non-English contexts, remains underexamined. We systematically evaluated 11 state-of-the-art LLMs on their ability to generate executable Python code for analytical queries over Czech public health datasets, focusing on incidence and prevalence data provided by the National Health Information Portal (known as NZIP). A set of representative analytical queries were designed, covering filtering, aggregation, weighted averages, and identification of primary diagnoses. Each model was prompted in Czech and assessed on code executability, correctness of results, and ability to adapt to local terminology. In the majority of cases, the models generated syntactically valid code within one minute, but performance varied. For the main objective of replicating “ground truth” queries as per dataset documentation, ChatGPT-4o achieved the highest accuracy, followed closely by GPT-4.1 mini. Claude and Gemini models frequently failed to apply critical filtering instructions, while Deepseek-R1, though accurate, defaulted to English output. Some models produced code that executed successfully but returned incorrect results, underscoring the need for systematic validation. Overall, LLMs show strong potential as coding assistants in public health analytics, even in Czech-language settings. Their integration into hybrid human–AI workflows, combined with validation mechanisms and retrieval-augmented generation, may accelerate the creation of reliable analytical pipelines. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This research was supported by the project CZ.02.01.01/00/23_025/0008743, funded by the European Union under the Operational Programme Johannes Amos Comenius (OP JAK). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study used ONLY openly available aggregated data. The datasets used are available at the following URL: I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Scripts and data produced within this study are available online at . If applicable, any other data produced in the present study are available upon reasonable request to the authors.
The substantia nigra (SN) has historically been regarded as a pivotal element of the brain’s motor circuits, notably within the context of the nigrostriatal pathway and Parkinson’s disease. However, recent advancements in neuroimaging techniques, particularly tractography, have facilitated the delineation of its anatomical projections. These techniques have revealed the involvement of the SN in a more extensive array of functional networks encompassing cognitive, emotional, and motivational domains. This paper reviews the current knowledge on the structural connectivity of the SN in humans based on diffusion tensor imaging and tractography. It summarizes the main projection pathways, including classical and newly described connections, such as the direct SN pars compacta connections to the thalamus, cortico–neural inputs, and connections to limbic regions and the hippocampus. Furthermore, the text delves into the distinctions between the SN pars compacta and SN pars reticulata subregions, exploring their parcellation based on connectivity. The paper demonstrates that the SN is a functionally diversified nucleus, the implications of which are significant for the understanding of both motor and neuropsychiatric disorders. The present study addresses the paucity of comprehensive treatment in this area and provides a framework for further research on dopaminergic circuits.
Head tremor is a common symptom in both essential tremor (ET) and cervical dystonia (CD). Distinguishing between these two conditions can be challenging in clinical practice, particularly when head tremor is the dominant feature. Our goal was to explore the potential of speech assessment in recognizing the mechanisms of head tremor in patients with ET and CD. Objective acoustic vocal assessments of oral diadochokinesis, phonatory stability, vocal tremor, and speech timing were performed. Of the 93 patients assessed, 39 had cervical dystonia (CD) with head tremor, 38 had ET with head tremor (ET-HT), and 16 had ET with no head tremor (ET-nHT). Compared to both CD and ET-nHT, ET-HT showed irregular sequential motion rate, excessive pitch fluctuations, increased noise, and higher extent of vocal vibrato. Compared to CD, ET-HT also demonstrated slower sequential motion rate, prolonged pauses, and a slower articulation rate. Additionally, ET-HT had more pronounced vocal tremolo compared to ET-nHT. Speech assessment provided discrimination between the CD and ET-HT groups with an area under curve of 0.80. This study underscores the promising potential of speech analysis in recognizing mechanisms of head tremor in patients with ET or CD, revealing more severe and distinct speech impairments in ET-HT patients compared to those with CD.
Parkinson’s disease (PD) and isolated REM sleep behaviour disorder (iRBD) are neurodegenerative conditions associated with alterations in the striatum, a subcortical structure essential for motor and cognitive functions. In this study, we apply diffusion tensor imaging (DTI) and probabilistic tractography to segment the striatum into two functionally distinct compartments: striosomes and matrix. A total of 152 subjects were included: 64 PD patients, 47 iRBD patients, and 41 healthy controls. Preprocessing, tractography, and segmentation were performed using the FSL toolbox; voxel-based morphometry (VBM) was used for group comparisons. Probabilistic group atlases were constructed in MNI space and compared with each other and with the standard MNI_152 atlas using RMSE and statistical testing. Results show significant differences in the matrix region, particularly in the right striatum of iRBD subjects compared to both PD and control groups. Our findings suggest that connectivity-based striatal segmentation may reflect early structural changes in neurodegeneration and could support future biomarker development for PD and related disorders. This tractography-based approach offers a promising tool for the personalized analysis of subcortical brain structures in clinical neuroimaging.
Assessing muscle activity is essential for diagnosis and treatment of neuromuscular disorders such as spasticity and dystonia. While muscle functional magnetic resonance imaging (MRI) enables non-invasive imaging of muscle activation, conventional methods rely on comparisons between rest and activity, which are unsuitable for patients with sustained muscle contractions. This pilot study introduces a novel resting-state fMRI (rs-fMRI) approach using regional homogeneity (ReHo) to evaluate muscle activity without requiring activation paradigms. Eight healthy male participants performed separate isometric plantar and dorsal foot flexion tasks during 3T MRI scanning. rs-fMRI data were analyzed using ReHo to assess local synchronization of BOLD signal. Calf muscle activation was quantified using z-transformed ReHo values and activation thresholds were derived via ROC analysis. Statistical differences in activation between tasks were assessed using the Wilcoxon signed-rank test. Significant ReHo differences were observed between active and inactive muscles during both tasks. This study demonstrates the potential of rs-fMRI combined with ReHo analysis as a non-invasive method for detecting muscle activity using one series of volumes. Further research in larger cohorts is warranted to validate and expand this approach. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study was supported by the National Institute for Neurological Research (Programme EXCELES, ID project no. LX22NPO5107) - Funded by the European Union - Next Generation EU. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethics committee of General University Hospital in Prague gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors.
Isolated REM sleep behavior disorder (iRBD) is associated with impaired colour discrimination, cognitive deficits and morphological changes. This study evaluates whether colour discrimination deficits in iRBD are mediated by cognitive functions or related to dopaminergic denervation and brain morphology. A sample of 73 patients with iRBD and 77 controls underwent neuropsychological assessment, and colour discrimination assessment using the Farnsworth Munsell 100 Hue Test, DAT-SPECT, and MRI. The data were analyzed using multiple regression, mediation analysis, and voxel-based morphometry. Significant between-group differences were found in total colour discrimination as well as in the red-yellow spectrum. The association between iRBD and performance in the yellow-green spectrum was mediated by cognitive functions, as measured by the Montreal Cognitive Assessment. In controls, a positive correlation between the yellow-green spectrum and the left inferior frontal gyrus was observed compared to patients, however, this association was largely driven by a single data point. The performance in the green–blue spectrum was associated with the activity of dopamine transporters in the caudate nucleus. No interactions were found for total colour discrimination in any analysis. The present findings demonstrate a colour vision deficit in iRBD, which is not directly linked to any of the proposed potential explanatory mechanisms.
Self-supervised pre-trained speech models such as wav2vec 2.0 provide rich frame-level embeddings that are increasingly used for clinical voice screening, including Parkinson's disease (PD). Optimally aggregating the frame-level embeddings is a significant task, with the underexplored question of how to best aggregate the frame-level embeddings into fixed-length utterance descriptors for the downstream task of binary classification (healthy controls (HC) vs. PD). To address this, our study compared three wav2vec 2.0 variants, the base model and two fine-tuned variants (one adapted for dysarthric corpora), across multiple layer depths, ten statistical aggregation functions (e.g., mean, median, quantiles), and two proposed advanced schemes (attention-based and multi-scale average pooling). We used the MDVR-KCL as a read-speech corpus (16 PD, 21 HC). Contrary to the expectation that sophisticated pooling would help, statistical aggregations such as mean pooling consistently provided better performance and robustness. Early representations (pre-Transformer and 1st Transformer block) were often the most informative, and mean aggregation produced relatively high, low-variance scores across models and depths. Attention and multi-scale pooling did not yield consistent gains. Moreover, wav2vec-based embeddings outperformed traditional acoustic baselines. Applying the supervised feature selection (ANOVA F-value) further improved performance, i.e. conservative selection (12 features) achieved a mean balanced accuracy of 0.87 and precision of 0.92, with top configurations exceeding a 0.93 balanced accuracy and had a precision of 1.0. The findings empirically support the continued use of mean pooling as a viable strategy for temporal aggregation of latent features for PD wav2vec-based detection.
Background: Freezing of gait (FoG) is a walking disturbance in the Parkinson's disease (PD). The freezing ratio (FoG-ratio) is a parameter used to quantify overall freezing severity rather than to assess single freezing episodes. Originally the FoG-ratio was designed to be computed from lower limb acceleration. However, some available measurement systems get their data from a single sensor located elsewhere, e.g. on the lower back. Purpose: The objective of our paper is to analyse whether acceleration signals measured on different body locations result in a consistent FoG-ratio. Methods: Eighty-four people with PD and 65 people without neurological disorders completed an instrumented Timed Up&Go Test (iTUG) twice. The FoG-ratios from inertial units placed on the chest, lower back, left and right lower limbs were calculated. Findings: There were significant differences between the tested FoG-ratios in the control group as well as in the PD group for both segments. Four significant, but not consistent, correlations were revealed for the turn segment in the PD group. Eight correlations were revealed in the control group. The inter-trial reliability of all the tested cases for gait was good (rho>0.75) but only in one case for turning. Conclusion: In conclusion, the placement of sensors affected the FoG-ratio parameter output. The different FoG-ratios reflect different amounts of power in the locomotion band of body segments. This could result in inconclusive validity and incomparability of freezing severity presented in studies when the sensor is placed somewhere other than on the lower limbs. (c) 2025 AGBM. Published by Elsevier Masson SAS. This is an open access article under the CC BY-NC-ND license
Diffusion tensor image analysis along the perivascular space (DTI-ALPS) is a non-invasive marker of glymphatic function that typically relies on manual region of interest (ROI) placement. This study compared glymphatic function in treatment-naive, de novo diagnosed patients with Parkinson's disease (PD), patients with isolated REM behavior disorder (iRBD), and healthy controls using both manual and automatic DTI-ALPS methods. ALPS scores were analyzed bilaterally and correlated with clinical severity (MDS-UPDRS) and nigrostriatal denervation (DAT-SPECT). The study included 79 PD patients (60±12 years), 57 iRBD patients (67±7 years), and 48 controls (62±10 years). ANCOVA revealed significant inter-group differences using both manual (p=0.018) and automatic (p=0.002) methods. Automatic analysis showed significantly lower ALPS scores in PD compared to controls (p=0.001) and iRBD (p=0.009). ALPS scores correlated with symptom severity and nigrostriatal degeneration. These findings highlight early glymphatic dysfunction in PD and demonstrate the reliability of the automatic DTI-ALPS method. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement The study was supported by: National Institute for Neurological Research (Programme EXCELES, ID Project No. LX22NPO5107) - Funded by the European Union - Next Generation EU; General University Hospital in Prague project MH CZ-DRO-VFN64165 and Na Homolce Hospital project CZ-DRO-NHH00023884 and Czech Health Research Council grant NU21-04-00535. Computational resources were provided by the e-INFRA CZ project (ID:90254), supported by the Ministry of Education, Youth and Sports of the Czech Republic. Access to CESNET storage facilities provided by the project "e-INFRA CZ" under the programme "Projects of Large Research, Development, and Innovations Infrastructures" LM2018140), is greatly appreciated. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethics committee of the General University Hospital in Prague gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The datasets used and analyzed during the current study are available from the corresponding author on request.
Atherosclerosis, a major cause of ischemic stroke worldwide, is characterized by plaque formation, particularly in the carotid bifurcation, leading to arterial stenosis. Traditional histology and light microscopy have been used to study atherosclerotic plaques, but the advent of digital pathology and artificial intelligence has provided new opportunities. In this work, we proposed an automatic segmentation method using convolutional neural networks (U-Net and DeepLabV3+) to delineate atherosclerotic carotid plaque tissue. The study included 835 images of histological slices stained with hematoxylin and eosin and Van Gieson's method from 114 patients. The results showed that DeepLabV3+ outperforms U-Net, achieving high accuracy for tissue types such as lumen, fibrous tissue, atheroma, calcification, and hemorrhage. Staining influenced segmentation results, with Van Gieson's stain excelling in fibrous tissue segmentation, while hematoxylin and eosin showed better results for calcification and hemorrhage. Moreover, the segmentation models facilitated clinical plaque classification, demonstrating good discrimination performance. Our study highlights the potential of deep neural networks in segmenting atherosclerotic plaques while emphasizing the need for careful consideration of staining effects in computerized analysis.
Introduction: Essential tremor (ET) is characterized by isolated action tremor of the upper limbs, with or without tremor of the head, voice, or lower limbs. The newly defined ET plus (ET+) syndrome is characterized by the presence of additional mild neurological symptoms that are not indicative of other disease. Aim: This study investigated the incidence of ET+ in our patients to better understand the pathophysiological basis of both entities, with a special focus on cerebellar involvement. Methods: In a cohort of 51 patients with ET (27 women, mean age 67.5 +/- 13.1 years), we verified anamnestic data and neurological findings. Tremor severity and the degree of disability in activities of daily living were assessed using The Essential Tremor Rating Assessment Scale (TETRAS), while symptoms of cerebellar involvement used the Scale for the Assessment and Rating of Ataxia (SARA). Results: ET+ was identified in 22 patients (mostly due to resting hand tremor), while 29 patients had "pure" ET. The two groups did not differ in age of onset or symptom duration. Patients with ET+ had a higher incidence of head tremor compared to those with ET (81 vs. 43%; P < 0.05) and higher scores on both the TETRAS and SARA scales (all P < 0.01). Conclusion: ET+ patients exhibited more severe overall tremor disability, a higher incidence of head tremor, and more pronounced signs of ataxia compared to patients with "pure" ET. Cerebellar involvement appears to play a significant role in additional manifestations of disability that characterize ET+ syndrome.
Background Gene fusions are critical drivers of oncogenesis and diagnostic biomarkers in various cancers. However, their detection from RNA or DNA sequencing, when performed using traditional analytical methods, encounters challenges related to sample quality, computational complexity, and noise. Although deep learning is more robust, it usually requires large labeled datasets and substantial training resources. Genomic foundation models (GFMs), which are pre-trained on pangenome-scale data, offer a promising solution to these issues. Methods This study presents the first comprehensive benchmark of four transformer-based GFMs, Nucleotide Transformer, Evo2, HyenaDNA, and DNABERT2, for gene fusion detection. Using the curated FusionAI dataset of ~ 52,000 sequences, we extracted embeddings from 10-kilobase-pair (kbp) DNA sequences surrounding fusion breakpoints. We evaluated the quality of these representations qualitatively using t-SNE visualization and quantitatively by training lightweight classifiers (Support Vector Machines and simple Neural Networks) on the fixed embeddings. Results The Nucleotide Transformer achieved the best performance with an accuracy of 0.967 and an F1 score of 0.967. This result outperformed the dedicated deep learning baseline (FusionAI, with an accuracy of 0.894). Evo2 was the second-best performer (accuracy: 0.920), demonstrating robustness derived from evolutionary pretraining. Conversely, DNABERT2 failed to compete (accuracy 0.677–0.723). Furthermore, sample efficiency analysis revealed that the Nucleotide Transformer required only ~ 2,600 samples to reach 95% of its peak performance, whereas the baseline required over 14,000 samples. Conclusions These findings demonstrate that advanced GFMs, particularly the NT and Evo2 models, generate highly discriminative 'out-of-the-box' embeddings. These embeddings significantly outperform dedicated deep learning baselines while requiring a fraction of the training data and computational time. This suggests that GFMs could be a scalable, data-efficient way of developing precise genomic diagnostic tools, particularly for rare diseases.
Speech and language technologies are effective tools for identifying the distinct speech changes associated with Parkinson's disease (PD), enabling earlier and more accurate diagnosis. Models leveraging recent advancements in self-supervised speech pretraining, such as Wav2Vec, have demonstrated superior performance over traditional feature extraction methods. While Wav2Vec 2.0 has been successfully utilized for PD detection, a rigorous quantitative comparison with Wav2Vec 1.0 is needed to comprehensively evaluate its advantages, limitations, and applicability across different speech modes in PD. This study presents a systematic comparison of Wav2Vec 1.0 and Wav2Vec 2.0 embeddings across three multilingual datasets using various classification approaches to classify normal (healthy controls; HC) and PD-affected speech. Additionally, both Wav2Vec 1.0 and 2.0 were benchmarked against traditional baseline features across diverse linguistic contexts, including spontaneous speech, non-spontaneous speech, and isolated vowels. A multicriteria TOPSIS approach was employed to rank feature extraction methods, revealing that Wav2Vec 2.0 excelled across speech modes, with its first transformer layer demonstrating the best performance for classifying read text and monologue, and its feature extractor performing best in vowel-based classification. In contrast, Wav2Vec 1.0, while generally outperformed by Wav2Vec 2.0, still provided a more efficient alternative with competitive performance. Finally, we combined selected layers from both architectures and have demonstrated improved diagnostic accuracy in vowel-based classification. This comparative analysis underscores the strengths of both Wav2Vec architectures and informs their optimal use in PD detection.