As a technique that can compactly represent complex patterns, machine learning has significant potential for predictive inference. K-fold cross-validation (CV) is the most common approach to ascertaining the likelihood that a machine learning outcome is generated by chance, and it frequently outperforms conventional hypothesis testing. This improvement uses measures directly obtained from machine learning classifications, such as accuracy, that do not have a parametric description. To approach a frequentist analysis within machine learning pipelines, a permutation test or simple statistics from data partitions (i.e., folds) can be added to estimate confidence intervals. Unfortunately, neither parametric nor non-parametric tests solve the inherent problems of partitioning small sample-size datasets and learning from heterogeneous data sources. The fact that machine learning strongly depends on the learning parameters and the distribution of data across folds recapitulates familiar difficulties around excess false positives and replication. A novel statistical test based on K-fold CV and the Upper Bound of the actual risk (K-fold CUBV) is proposed, where uncertain predictions of machine learning with CV are bounded by the worst case through the evaluation of concentration inequalities. Probably Approximately Correct-Bayesian upper bounds for linear classifiers in combination with K-fold CV are derived and used to estimate the actual risk. The performance with simulated and neuroimaging datasets suggests that K-fold CUBV is a robust criterion for detecting effects and validating accuracy values obtained from machine learning and classical CV schemes, while avoiding excess false positives.
Depression, a prevalent mental illness, often manifests itself in a set of symptoms, including subtle and sometimes subclinical abnormalities in vocal and facial expressions. However, unimodal analysis typically falls short in capturing the complexity of such signals. Multimodal learning has been the recent trend in affective computing, but many existing approaches follow independent modality-wise processing and late fusion, which limits their ability to model the intricate cross-modal interactions in a fine-grained manner. In this paper, we propose a Cross-Attention Multimodal Fusion framework for audio-visual depression detection to dynamically fuse the complementary information from speech and facial dynamics. Our model uses a two-stream encoder to first extract the audio and visual features and then aligns them using a bidirectional cross-attention mechanism. This allows the network to attend to the visual signal while processing the acoustic modality and vice versa to model the context-aware correlations related to depressive cues. We evaluate our approach on benchmark audio-visual depression datasets and show that it outperforms early- and late-fusion baselines, highlighting the potential of cross-attention based methods in multimodal affective analysis.
Autism Spectrum Condition (ASC) is a neurodevelopmental condition characterized by impairments in communication, social interaction and restricted or repetitive behaviors. Extensive research has been conducted to identify distinctions between individuals with ASC and neurotypical individuals. However, limited attention has been given to comprehensively evaluating how variations in image acquisition protocols across different centers influence these observed differences. This analysis focuses on structural magnetic resonance imaging (sMRI) data from the Autism Brain Imaging Data Exchange I (ABIDE I) database, evaluating subjects' condition and individual centers to identify disparities between ASC and control groups. Statistical analysis, employing permutation tests, utilizes two distinct statistical mapping methods: Statistical Agnostic Mapping (SAM) and Statistical Parametric Mapping (SPM). Results reveal the absence of statistically significant differences in any brain region, attributed to factors such as limited sample sizes within certain centers, noise effects and the problem of multicentrism in a heterogeneous condition such as autism. This study indicates limitations in using the ABIDE I database to detect structural differences in the brain between neurotypical individuals and those diagnosed with ASC. Furthermore, results from the SAM mapping method show greater consistency with existing literature.
Alzheimer's disease (AD) is a progressive neurodegenerative disorder characterized by cognitive decline and substantial brain atrophy. Early and accurate prediction of disease progression and staging is crucial for timely intervention and effective treatment planning. Previous studies, including those based on artificial intelligence techniques, have employed neuroimaging, biomarkers and clinical data to model AD progression; however, many of these approaches rely on strong parametric assumptions or lack robust statistical guarantees regarding model validity. To bridge this gap, this study proposes a novel framework for validating predictive and staging models of disease using a statistically agnostic methodology. The objective is to take the advantages of an unconventional method for robust validation of ML models related to AD. Validation is performed using the Statistical Agnostic Regression (SAR) methodology applied to the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset. The method tests for a linear relationship by resampling and estimating an upper bound on the expected risk (R) via a Bayesian bound under the worst-case scenario. The SAR power assesses the likelihood of detecting a true linear relationship using the test statistic R, via Monte Carlo simulations under the null distribution. Three predictive models related to structural neuroimaging are assessed: one for the Mini Mental State Examination (MMSE) score, another for the concentration of amyloid beta 1-42 protein in the cerebrospinal fluid, and a third for age. In addition, a model for staging based on Alzheimer's-related clinical groups is explored through the joint analysis of segmented gray matter and white matter images. The findings indicate that the SAR methodology not only facilitates robust validation of predictive ML models related to neuroimaging and AD but also enables an effective staging of the AD continuum. This SAR-proposed framework opens new perspectives for the validation of ML models for early diagnosis and provides a solid foundation for future research in computational neuroscience.
The abstract should briefly summarize the contents of the paper in The growing presence of multimedia data in contemporary digital spaces has posed a high demand of artificial intelligence systems that can comprehend information that is presented in a multi-modal form of vision, audio and language. The human perception offers a viable and naturalistically based paradigm to solve this problem because the brain has a natural habit of integrating heterogeneous information of sensory signals based on hierarchical processing, selective attention and contextual learning. Our proposed bio-inspired multimodal deep learning framework is intended to be used in cognitive multimedia perception in this paper. In particular, we propose a cross-modal attention mechanism based on biology which dynamically models the reliability of the modality and removes the noisy or missing senses. The given method uses modality-specific deep encoders and attention-based fusion strategy based on the human multisensory integration, which allows to adaptively weight visual, auditory and textual information depending on its relevance to the current situation. Extensive simulations on representative multimodal benchmarks show that the proposed framework always beats unimodal frameworks and traditional multimodal fusion schemes with respect to accuracy, robustness and interpretability. According to its experimental findings, the maximum absolute accuracy increase compared to its traditional fusion approaches is 4.4
BACKGROUND AND OBJECTIVES:Behavioral and neuropsychiatric symptoms are common in frontotemporal dementia (FTD) and primary progressive aphasia (PPA). However, little is known about their patterns, time course, and association with brain atrophy. We, therefore, aimed to describe behavioral and neuropsychiatric phenotypes in patients with FTD and PPA, leveraging a hypothesis-free/data-driven approach. METHODS:We included participants diagnosed with behavioral variant FTD (bvFTD) or PPA according to Rascovsky and Gorno-Tempini criteria from the German Center for Neurodegenerative Diseases Clinical Registry Study of Neurodegenerative Diseases-FTD prospective multicenter observational cohort study. Symptoms were assessed using the Neuropsychiatric Inventory-Questionnaire. Principal component analysis (PCA) was used to delineate symptom groups. Subsequently, frequency and severity across diagnostic groups were examined. We applied linear mixed-effects models to describe the longitudinal evolution of symptoms. Associations with MRI-assessed atrophy were investigated using linear regression models. RESULTS:A total of 314 patients (42.4% female, mean age 65.52 [SD 9.0] years) with bvFTD or PPA were included. MRI was available for 134 of 314 individuals. PCA revealed 4 natural symptom groups, labeled active behavioral, passive behavioral, affective, and psychotic phenotypes. Symptom groups were observed at comparable frequencies across diagnostic groups. Time from symptom onset (0.130 [0.044-0.217], p < 0.003), sex (1.376 [0.666-2.087], p < 0.001), and the interaction between the nonfluent variant of PPA and sex (-1.940 [-3.242 to -0.638], p = 0.004) showed a significant effect on the active behavioral phenotype, with symptom severity increasing over time and being most pronounced in men with bvFTD. Patients with bvFTD exhibited more severe passive behavioral symptoms compared with any other diagnostic group. For the affective phenotype, a significant interaction between time and sex (0.063 [0.010-0.117], p = 0.021) indicated a progressive increase in symptom severity in men over time. Furthermore, we found robust neuroanatomical correlations of passive behavioral symptoms with subcortical and bilateral frontal and cingulate cortical atrophy. DISCUSSION:Our findings demonstrate that behavioral and neuropsychiatric symptoms are prevalent in both bvFTD and PPA. Their severity depends on the disease duration, phenotypic group, and sex. This detailed understanding of symptomatology is crucial for optimizing patient care, diagnostic evaluations, and the design of clinical trials. Limitations comprise the lack of neuropathologic validation and the limited availability of MRI data.
Major Depressive Disorder (MDD) is a major global health challenge and is commonly assessed through semi-structured interviews and questionnaires, which are clinically validated but partly subjective. Affective Computing has therefore explored digital biomarkers from multimodal behavioural data, including acoustic, visual, and linguistic cues. This paper evaluates a hierarchical deep learning framework for depression detection that combines a frozen DistilBERT encoder for text, Bi-LSTMs for temporal modelling of audio and video feature streams, and a Transformer-based fusion module to capture inter-modal dependencies. Using the Extended DAIC-WOZ corpus ( N=275 ), we conduct a systematic ablation study across unimodal, bimodal, and multimodal variants. While the multimodal configuration attains strong performance (F1=0.992), text-enhanced variants approach near-ceiling results, suggesting potential shortcut learning and/or leakage effects in interview-based settings. Overall, the results highlight both the promise of multimodal fusion and the need for stricter leakage controls and more clinically grounded linguistic inputs to improve generalisability and interpretability for computer-aided screening.
Medical imaging fusion combines complementary information from multiple modalities to enhance diagnostic accuracy. However, evaluating the quality of fused images remains challenging, with many studies relying solely on classification performance, which may lead to incorrect conclusions. We introduce a novel framework for improving image fusion, focusing on preserving fine-grained details. Our model uses a siamese autoencoder to process T1-MRI and FDG-PET images in the context of Alzheimer’s disease (AD). The framework optimizes fusion by minimizing reconstruction error between generated and input images, while maximizing differences between modalities through cosine distance. Additionally, we propose a supervised variant, incorporating binary cross-entropy loss between diagnostic labels and probabilities. Fusion quality is rigorously assessed through three tests: 1) classification of AD patients and controls using fused images; 2) an atlas-based occlusion test for identifying regions relevant to cognitive decline; and 3) analysis of structural-functional relationships via Euclidean distance. Results show an AUC of 0.92 for AD detection, reveal the involvement of brain regions linked to preclinical AD stages, and demonstrate preserved structural-functional brain networks, indicating that subtle differences are successfully captured through our fusion approach.
OBJECTIVES:Determining the involvement of specific peripheral nerves (PNs) in the upper limb associated with signs of muscle denervation can be challenging. This study aims to develop, compare, and validate various large language models (LLMs) to automatically identify and establish potential relationships between denervated muscles and their corresponding PNs. MATERIALS AND METHODS:We collected 300 retrospective MRI reports in Spanish from upper limb examinations conducted between 2018 and 2024 that showed signs of muscle denervation. An expert radiologist manually annotated these reports based on the affected peripheral nerves (median, ulnar, radial, axillary, and suprascapular). BERT, DistilBERT, mBART, RoBERTa, and Medical-ELECTRA models were fine-tuned and evaluated on the reports. Additionally, an automatic voting system was implemented to consolidate predictions through majority voting. RESULTS:The voting system achieved the highest F1 scores for the median, ulnar, and radial nerves, with scores of 0.88, 1.00, and 0.90, respectively. Medical-ELECTRA also performed well, achieving F1 scores above 0.82 for the axillary and suprascapular nerves. In contrast, mBART demonstrated lower performance, particularly with an F1 score of 0.38 for the median nerve. CONCLUSIONS:Our voting system generally outperforms the individually tested LLMs in determining the specific PN likely associated with muscle denervation patterns detected in upper limb MRI reports. This system can thereby assist radiologists by suggesting the implicated PN when generating their radiology reports.
Peripheral Nerves (PNs) are traditionally evaluated using US or MRI, allowing radiologists to identify and classify them as normal or pathological based on imaging findings, symptoms, and electrophysiological tests. However, the anatomical complexity of PNs, coupled with their proximity to surrounding structures like vessels and muscles, presents significant challenges. Advanced imaging techniques, including MR-neurography and Diffusion-Weighted Imaging (DWI) neurography, have shown promise but are hindered by steep learning curves, operator dependency, and limited accessibility. Discrepancies between imaging findings and patient symptoms further complicate the evaluation of PNs, particularly in cases where imaging appears normal despite clinical indications of pathology. Additionally, demographic and clinical factors such as age, sex, comorbidities, and physical activity influence PN health but remain unquantifiable with current imaging methods. Artificial Intelligence (AI) solutions have emerged as a transformative tool in PN evaluation. AI-based algorithms offer the potential to transition from qualitative to quantitative assessments, enabling precise segmentation, characterization, and threshold determination to distinguish healthy from pathological nerves. These advances could improve diagnostic accuracy and treatment monitoring. This review highlights the latest advances in AI applications for PN imaging, discussing their potential to overcome the current limitations and opportunities to improve their integration into routine radiological practice.
Brain tumor resection is a complex procedure with significant implications for patient survival and quality of life. Predictions of patient outcomes provide clinicians and patients the opportunity to select the most suitable onco-functional balance. In this study, global features derived from structural magnetic resonance imaging in a clinical dataset of 49 pre- and post-surgery patients identified potential biomarkers associated with survival outcomes. We propose a framework that integrates Explainable AI (XAI) with neuroimaging-based feature engineering for survival assessment, offering guidance for surgical decision-making. In this study, we introduce a global explanation optimizer that refines survival-related feature attribution in deep learning models, enhancing interpretability and reliability. Our findings suggest that survival is influenced by alterations in regions associated with cognitive and sensory functions, indicating the importance of preserving areas involved in decision-making and emotional regulation during surgery to improve outcomes. The global explanation optimizer improves both fidelity and comprehensibility of explanations compared to state-of-the-art XAI methods. It effectively identifies survival-related variability, underscoring its relevance in precision medicine for brain tumor treatment.
Transfer learning (TL) is a strategic solution to handle vast data volume requirements in deep learning (DL). It transfers knowledge learned from a large base dataset, as a pretrained model (PTM), to a new domain. In this study, we introduce an ensemble of classifiers trained on features extracted from some intermediate layers of a PTM for Tuberculosis (TB) detection task. We use different EfficientNet variants: EfficientNet-B0-EfficientNet-B3, as the PTM. Moreover, we introduce a rejection mechanism and implement post-hoc calibration methods to enhance the reliability and trustworthiness of the developed models. Additionally, we conduct analyses on domain-shift distribution, a topic rarely discussed in the context of TB detection. Through a fivefold cross-validation on two prominent chest X-ray datasets, the Montgomery County (MC) and Shenzhen (SZ), our ensemble approach achieved competitive results with accuracies of 94.89% (MC) and 92.75% (SZ). The incorporation of the devised rejection mechanism resulted in enhanced model accuracy, albeit with a coverage tradeoff. In domain-shift experiments, the proposed approach achieved an accuracy of 83.57% (63% coverage) when applying the MC-trained model on SZ, and an accuracy of 88.50% (82% coverage) when applying the SZ-trained model on MC.
INTRODUCTION:Regression analysis is a central topic in statistical modeling, aimed at estimating the relationships between a dependent variable, commonly referred to as the response variable, and one or more independent variables, i.e., explanatory variables. Linear regression is by far the most popular method for performing this task in various fields of research, such as data integration and predictive modeling when combining information from multiple sources. OBJECTIVES:Classical methods for solving linear regression problems, such as Ordinary Least Squares (OLS), Ridge, or Lasso regressions, often form the foundation for more advanced machine learning (ML) techniques, which have been successfully applied, though without a formal definition of statistical significance. At most, permutation or analyses based on empirical measures (e.g., residuals or accuracy) have been conducted, leveraging the greater sensitivity of ML estimations for detection. METHODS:In this paper, we introduce Statistical Agnostic Regression (SAR) for evaluating the statistical significance of ML-based linear regression models. This is achieved by analyzing concentration inequalities of the actual risk (expected loss) and considering the worst-case scenario. To this end, we define a threshold that ensures there is sufficient evidence, with a probability of at least 1-η, to conclude the existence of a linear relationship in the population between the explanatory (feature) and the response (label) variables. CONCLUSIONS:Simulations demonstrate that the proposed agnostic (non-parametric) test can perform an analysis of variance comparable to the classical multivariate F-test for the slope parameter, without relying on the underlying assumptions of classical methods. A power analysis on a putative regression task revealed an overinflated false positive rate in standard ML methods, whereas the SAR test exhibited excellent control. Moreover, the residuals computed using this method represent a trade-off between those obtained from ML approaches and classical OLS.
Image super-resolution (SR) is a classic visual problem that aims to generate high-quality, high-resolution images from low-resolution inputs. However, most deep learning methods are designed for visible images and often overlook infrared images, which play a crucial role in numerous research fields such as aerospace and remote sensing. Due to hardware limitations, infrared images possess a lower resolution and exhibit characteristics distinct from visible images, including low contrast and indistinct gradients. These unique patterns are challenging to extract and represent. To address this issue, we propose a Infrared Feature Fusion Network (InfraFFN) for infrared image SR in this paper. Specifically, we design a Residual Feature Fusion Block (RFFB) for deep feature extraction. Each Feature Fusion Block (FFB) within RFFB effectively combines the advantages of convolution and self-attention, and utilizes bi-directional information interactions across branches to better model in both channel and spatial dimensions. Furthermore, considering the low contrast in infrared images, we designed a dual-path convolution structure to extract features under different sizes of receptive fields and fuse features at various scales. Extensive experiments demonstrate that our InfraFFN achieves superior visual improvement on multiple infrared image datasets compared to state-of-the-art methods. The source codes are available at https://github.com/szw811/InfraFFN.
In the realm of neuroscience, brain activity is often characterized by rhythmic oscillations at different frequency bands. These oscillations underlie various cognitive processes and constitutes the basis of communication between populations of neurons. Cross-frequency coupling (CFC) refers to techniques directed to study the interactions between oscillations at different frequencies, providing a more comprehensive view of neural dynamics than traditional measures of connectivity or based on the distribution of the power spectral density. In this paper, we propose a method to explore CFC local patterns in an explainable way, allowing to visualize them over time and to easily identify functional brain areas activated during a task development from the Phase-Amplitude Coupling (PAC) point of view.
Myocarditis is a significant public health concern because of its potential to cause heart failure and sudden death. The standard invasive diagnostic method, endomyocardial biopsy, is typically reserved for cases with severe complications, limiting its widespread use. Conversely, non-invasive cardiac magnetic resonance (CMR) imaging presents a promising alternative for detecting and monitoring myocarditis, because of its high signal contrast that reveals myocardial involvement. To assist medical professionals via artificial intelligence, the authors introduce generative adversarial networks - multi discriminator (GAN-MD), a deep learning model that uses binary classification to diagnose myocarditis from CMR images. Their approach employs a series of convolutional neural networks (CNNs) that extract and combine feature vectors for accurate diagnosis. The authors suggest a novel technique for improving the classification precision of CNNs. Using generative adversarial networks (GANs) to create synthetic images for data augmentation, the authors address challenges such as mode collapse and unstable training. Incorporating a reconstruction loss into the GAN loss function requires the generator to produce images reflecting the discriminator features, thus enhancing the generated images' quality to more accurately replicate authentic data patterns. Moreover, combining this loss function with other regularisation methods, such as gradient penalty, has proven to further improve the performance of diverse GAN models. A significant challenge in myocarditis diagnosis is the imbalance of classification, where one class dominates over the other. To mitigate this, the authors introduce a focal loss-based training method that effectively trains the model on the minority class samples. The GAN-MD approach, evaluated on the Z-Alizadeh Sani myocarditis dataset, achieves superior results (F-measure 86.2%; geometric mean 91.0%) compared with other deep learning models and traditional machine learning methods.
Parkinson’s disease (PD), a complex and debilitating neurological disorder, often leads to progressive cognitive decline, including mild cognitive impairment (MCI) and dementia. Over the years, various methods have been developed to diagnose PD, with neuroimaging modalities, particularly electroencephalogram (EEG) recording, gaining significant traction among specialist doctors. This article presents a novel PD detection method employing deep learning (DL) techniques to analyze EEG signals. The proposed method utilizes the UC San Diego (UCSD) resting-state EEG dataset and involves a meticulous preprocessing phase encompassing filtering, channel selection, and EEG signal windowing. Subsequently, a novel 1D CNN-LSTM architecture is introduced for extracting salient features from EEG signals. In the classification stage, three algorithms, namely Softmax, support vector machine (SVM), and decision tree (DT), are employed and their performances compared. To assess the robustness of the classification models, k-fold cross-validation with k=10 is implemented. The results demonstrate that the SVM algorithm exhibits superior performance, achieving an impressive 99.51 % accuracy for binary classification and 99.75 % accuracy for multi-class classification tasks. To gain insights into the model’s decision-making process and enhance interpretability, t-distributed Stochastic Neighbor Embedding (t-SNE) is utilized as an explainable artificial intelligence (XAI) method in the post-processing stage.
As the world's population ages, Alzheimer's disease is currently the seventh most common cause of death globally; the burden is anticipated to increase, especially among middle-class and elderly persons. Artificial intelligence-based algorithms that work well in hospital environments can be used to identify Alzheimer's disease. A number of databases were searched for English- language articles published up until March 1, 2024, that examined the relationships between artificial intelligence techniques, eye movements, and Alzheimer's disease. A novel non-invasive method called eye movement analysis may be able to reflect cognitive processes and identify anomalies in Alzheimer's disease. Artificial intelligence, particularly deep learning, and machine learning, is required to enhance Alzheimer's disease detection using eye movement data. One sort of deep learning technique that shows promise is convolutional neural networks, which need further data for precise classification. Nonetheless, machine learning models showed a high degree of accuracy in this context. Artificial intelligence-driven eye movement analysis holds promise for enhancing clinical evaluations, enabling tailored treatment, and fostering the development of early and precise Alzheimer's disease diagnosis. A combination of artificial intelligence-based systems and eye movement analysis can provide a window for early and non-invasive diagnosis of Alzheimer's disease. Despite ongoing difficulties with early Alzheimer's disease detection, this presents a novel strategy that may have consequences for clinical evaluations and customized medication to improve early and accurate diagnosis.
The analysis of neuroimaging data by means of computer systems has become a general practice in the diagnosis and monitoring of Alzheimer’s disease. In recent years, different systems based on neural networks have been proposed to aid in the diagnosis of this disorder. These systems usually contain millions of parameters that must be adjusted during the training process. This requires large datasets, often not available in most studies. The use of pre-trained systems (also known as transfer learning) would help to alleviate this problem, however these systems are often pre-trained with data of a very different nature than neuroimaging data, so their use could be counterproductive. In this work we evaluate the use of transfer learning for the development of a computer-aided diagnosis system for Alzheimer’s disease. To this end, we compared the performance obtained by different systems with and without transfer learning. The results show that transfer learning improves the results in some cases, reaching higher performances than other systems, however, the improvement depends largely on the optimizer used. In addition, we propose a novel method to convert neuroimaging data from 3D volumes to the 2D images used as input for most pre-trained systems.
Automatic facial expression recognition is a big challenge in human–computer interaction. Analyzing the changes in the face during a facial expression can be used for this purpose. In this paper, these changes are extracted as a number of motion vectors. These motion vectors are extracted using an optical flow algorithm. Then, they are used to analyze facial expressions by some of the data mining algorithms. This analysis has not only determined what changes occur in the face during facial expression but has also been used to recognize facial expressions. Cohen-Kanade facial expression dataset was used in this research. Based on our findings, the vertical lengths of motion vectors created in the lower part of the face have the greatest impact on the classification of facial expressions. Among the investigated classification algorithms, deep learning, support vector machine, and C5.0 had better performance, yielding an accuracy of 95.3%, 92.8%, and 90.2% respectively.