The human brain undergoes dynamic, potentially pathology-driven, structural changes throughout a lifespan. Longitudinal Magnetic Resonance Imaging (MRI) and other neuroimaging data are valuable for characterizing trajectories of change associated with typical and atypical aging. However, the analysis of such data is highly challenging given their discrete nature with different spatial and temporal image sampling patterns within individuals and across populations. This leads to computational problems for most traditional deep learning methods that cannot represent the underlying continuous biological process. To address these limitations, we present a new, fully data-driven method for representing aging trajectories across the entire brain by modelling subject-specific longitudinal T1-weighted MRI data as continuous functions using Implicit Neural Representations (INRs). Therefore, we introduce a novel INR architecture capable of partially disentangling spatial and temporal trajectory parameters and design an efficient framework that directly operates on the INRs' parameter space to classify brain aging trajectories. To evaluate our method in a controlled data environment, we develop a biologically grounded trajectory simulation and generate T1-weighted 3D MRI data for 450 healthy and dementia-like subjects at regularly and irregularly sampled timepoints. In the more realistic irregular sampling experiment, our INR-based method achieves 81.3
Most deep learning models for ocular condition classification from color fundus photographs (CFPs) operate on images that are rescaled to a uniform spatial resolution, which is usually lower than the diverse acquisition resolutions of modern imaging devices. While required for streamlining data processing in standard deep learning pipelines, such downsampling may potentially lead to a loss of clinically relevant details needed for an accurate diagnosis. To circumvent the uniformity requirement and to make better use of the full image information originally acquired, resolution-independent deep learning approaches to retinal image analysis are highly desirable. We propose to address this limitation at an architectural level via a novel framework based on Implicit Neural Representations (INRs). Specifically, our framework models individual discrete CFPs as continuous functions represented by neural networks that map spatial coordinates to RGB values. This design enables resolutionin-dependent image representations with the potential to process CFPs without any downsampling. We then present a multilayer perceptron (MLP)-based classifier that directly operates on these INR representations for a challenging multi-class classification of five ocular conditions and a healthy class. We train and test the model on INRs fit to a diverse set of 19,565 CFP images and benchmark it against three standard deep learning classifiers operating on the images themselves. While our model shows a slightly lower overall performance than the baselines, it achieves a strong performance on conditions such as cataract and myopia, with class-specific accuracies of 98.10% and 95.38%, respectively. We believe that these results demonstrate the general feasibility of INR-based classification for multi-class CFP analysis and that our novel method is an important step towards developing resolution-independent deep learning approaches in ophthalmology.
Cervical cancer remains a major global health burden, where early detection through cytological screening is critical for improving outcomes. However, Pap smear analysis is often manual, time-consuming, and prone to inter-observer variability. Although deep learning offers promise for automating cytology analysis, challenges such as limited data, class imbalance, staining variability, and overlapping cells hinder reliable model development. In this work, we address these challenges in the RIVA Cytology Challenge, which focuses on nucleus localization in Pap smear images. We propose a cascaded detection framework using multi-dataset pretraining followed by domain-specific finetuning. A YOLO11x model is pretrained on a harmonized multi-dataset collection and then finetuned on the RIVA dataset. During inference, a class-aware pipeline with test-time augmentations and nucleus-centering refinement improves localization accuracy. Our approach achieved 4th place in the challenge evaluation phase. The code is available at this repository: https://github.com/vibujithan/RIVA-Challenge-2026-TrackB
Despite the growing interest in biological brain age prediction models that rely on T1-weighted magnetic resonance imaging (MRI) data, evaluating model properties and validity remains a challenge. Obstacles include the scarcity of individual-level longitudinal neuroimaging data and the difficulty of establishing references for voxelwise brain aging patterns. In this work, we introduce a novel evaluation framework that utilizes diverse, realistic synthetic longitudinal MRI data to objectively assess global and voxelwise brain age prediction models. Our new framework specifically enables the investigation of inherent longitudinal prediction biases, sensitivities to localized brain aging patterns, and the impact of focal pathologies on performance, evaluations that are scarcely feasible with real-world data. We demonstrate the utility of the framework by applying it to established brain age prediction models for global and voxelwise brain age prediction and reveal new insights into their capabilities and limitations.
Mild Behavioural Impairment (MBI) is defined by later-life onset of persistent behavioural changes and is recognized as a risk marker for cognitive decline and dementia. Apathy, a core MBI domain characterized by diminished interest, initiative, and emotional reactivity, can emerge before dementia and is hypothesized to be associated with structural brain changes. While previous studies have explored Alzheimer disease (AD)-related neuroanatomical substrates of apathy in the dementia clinical stage, few have investigated these associations in cognitively normal (CN) or mild cognitive impairment (MCI) individuals with persistent apathy consistent with MBI. Thus, this study explores structural brain differences between individuals with MBI-apathy and those without neuropsychiatric symptoms (no-NPS). Participants (n = 446; mean age = 69.6 years; 79.8% CN; 62.8% female) were drawn from the National Alzheimer's Coordinating Center and categorized into MBI-apathy (n = 59) and no-NPS (n = 387) groups. Linear regressions were used to model associations between NPS group and regional brain measures, with adjustments for age, sex, years of education, apolipoprotein E4 carrier status, intracranial volume, and Mini-Mental State Examination score, with false discovery rate (FDR) correction for multiple comparisons. Primary outcomes included two predefined AD meta-regions-of-interest (ROIs): 1) thickness: a composite measure of mean cortical thickness across the entorhinal cortex, inferior temporal gyrus, middle temporal gyrus, inferior parietal lobule, fusiform gyrus, and precuneus; and 2) volume: a composite measure of mean cortical and subcortical grey matter volume across the hippocampus, entorhinal cortex, amygdala, middle temporal gyrus, inferior parietal lobule, and precuneus. Primary outcomes also included cortical thickness and grey matter volume among individual ROIs including the ventral striatum (VS), anterior cingulate cortex (ACC), orbitofrontal cortex (OFC), ventrolateral prefrontal cortex (vlPFC), and dorsolateral prefrontal cortex (dlPFC). MBI-apathy status was associated with significantly lower AD-meta-ROI cortical thickness (Z-score difference [95% CI]; FDR-corrected p-value, -0.43 [-0.73 - [-0.12]]; 0.025) and lower AD meta-ROI grey matter volume (-0.50 [-0.71 - [-0.30]]; <0.001). MBI-apathy was also associated with significantly lower dlPFC thickness (-0.40, [-0.70 - [-0.09]]; 0.02) and volume (-0.28 [-0.50 - [-0.06]]; 0.026) and lower OFC volume (-0.32, [-0.57 - [-0.07]]; 0.026) compared to the no-NPS group. Within a non-dementia sample, MBI-apathy was more strongly associated with established AD-vulnerable regions than with regions that have been traditionally implicated in apathy in dementia. Results suggests that during CN and MCI stages, MBI-apathy may reflect early AD-related neurodegeneration, with conventional apathy-related structural changes becoming more prominent as disease progresses.
Perinatal stroke affects millions of individuals worldwide, often leading to lifelong complications for those who survive and who require targeted rehabilitation to limit disability. Although biological brain age prediction may be a valuable biomarker to analyze neurodevelopment following perinatal stroke and guide rehabilitation, there have been no prior studies exploring its utility. Therefore, in this work, we analyzed neurodevelopment in children with two forms of perinatal stroke, namely arterial ischemic stroke (AIS) and periventricular venous infarction (PVI), by using T1-weighted neuroimaging data and machine learning-based biological brain age prediction at a global and voxel level. Specifically, we analyzed trends in the brain age gap (BAG) at both global and voxel-wise levels in the contralesional hemisphere, alongside correlation analyses with motor scores in stroke cohorts. Global and voxel-wise biological brain age prediction machine learning models were developed and trained using 5969 T1-weighted MRI scans of typically-developing children (mean age: 12.11 ± 2.72 years). These trained models were then applied to T1-weighted MRI data from N = 105 subjects with perinatal stroke (mean age: 11.41 ± 3.27 years) and N = 105 age- and sex-matched controls. The Wilcoxon signed-rank test and Mann-Whitney U-test were used to identify differences in BAGs in the stroke vs. controls subgroups, and AIS vs. PVI subgroups, respectively. Spearman's correlation coefficient was used to identify relationships between BAGs and motor scores. Lastly, sex-specific analyses were performed to identify sexually dimorphic characteristics in the data. Children with perinatal stroke exhibited, on average, more positive global and voxel-wise BAGs, particularly those with AIS. Motor scores were negatively correlated with global and voxel-wise BAGs in the contralesional hemisphere. The correlations between BAGs and motor scores differed between the sexes. This is the first study to propose and explore voxel-wise brain age prediction for a pediatric cohort and the first to utilize brain age prediction to study perinatal stroke. Further development of these methods may reveal biomarkers that are valuable to study other pediatric diseases and promote personalized rehabilitation for affected individuals.
IntroductionBehaviors of concern, such as aggression and self-injury, are common in children with neurodevelopmental disorders and contribute substantially to caregiver stress and health system burden. However, most evidence on behaviors is drawn from diagnostically homogeneous samples, limiting relevance to real-world presentations encountered in tertiary care. Our aim was to characterize the diagnostic and behavioral complexity of children referred to tertiary clinics and use machine learning to identify transdiagnostic features associated with behaviors across diverse neurodevelopmental presentations.MethodThis retrospective chart review and analysis examined all children aged 2 to 17 years undergoing first-time assessment at a tertiary developmental pediatrics clinic between May 2022-2023. Unstructured physician notes and supporting documentation were transformed into structured data and supervised machine learning models identified factors associated with behaviors of concern.ResultsAmong 600 children, 83% exhibited at least one behavior of concern (mean, 3.15 [SD 2.98]). Most had multiple neurodevelopmental diagnoses (73%) and frequent co-occurrence of physical (80%) and mental health conditions (21%). Emotional dysregulation and non-cooperation were the most common behaviors, while aggression and self-injury were frequently moderate to severe. Machine learning models identified sleep difficulties, speech and language delay, restricted and repetitive behaviors (B symptoms) of autism, gastrointestinal issues, and service use as top-ranked associated factors across multiple behavioral types. Factors associated with any behaviors included ADHD, sensory impairments, and genetic contributors to autism.DiscussionTransdiagnostic factors highlight key domains often under-assessed in routine care. These findings support integrated assessment models that address modifiable clinical correlates for behaviors of concern across diagnostic boundaries to improve outcomes for high-need neurodevelopmental populations.
Early and accurate computer-aided detection of abnormal cervical cells in Papanicolaou smear images is essential for preventing cervical cancer, in particular for centers with limited access to expert cytotechnologists. The RIVA Cytology Challenge provides a large-scale, expert-annotated collection of high-resolution Papanicolaou smear images from Hospital Rivadavia, offering the unique opportunity to develop a highly accurate classification and detection model for automated cervical cancer screening. In this work, we describe our solution for the challenge in Track A, which involves both the classification and detection tasks. Our proposed solution achieved the mean Average Precision (mAP@0.50:0.95) score of 0.14183, and demonstrate many strategies, which were found to be successful in progressively improving the test scores, as well as those which failed to do so.
Vision-influencing ocular conditions such as diabetic retinopathy, age-related macular degeneration, and degenerative myopia affect millions globally. Color fundus photography is a cost-effective, non-invasive, and widely-used imaging modality that captures key ocular structures and helps to diagnose patients. However, manual interpretation of a fundus image by a clinician is time-consuming. This has spurred the development of discriminative deep learning models for the detection of ocular conditions. While they demonstrate impressive classification performance, they fall short in providing transparent explanations to the clinicians. On the other hand, so-called generative classifiers (generative models repurposed as a classifier) not only perform competitively with discriminative deep learning models but also provide meaningful and interpretable explanations. Moreover, generative classifiers have been shown to be less prone to shortcut learning in computer vision tasks and neuroimaging compared to discriminative deep learning models. However, it is unclear if this holds true in ocular condition detection. In this paper, we investigate the presence of potentially shortcut-induced performance disparities of a recently presented generative classifier across subgroups in a multi-class ocular condition detection setting and compare the results to those achieved by a discriminative classifier. We evaluate the balanced accuracy score across the three attributes: sex, age, and camera. The generative classifier model achieves a higher balanced accuracy (84.86%) compared to the discriminative model (83.60%). Furthermore, the results reveal that the generative classifier suffers less from performance disparities across sex and age groups compared to the discriminative model. We believe that our results reinforce the idea that generative classifiers are an important step towards trustworthy AI in retinal image analysis.
Predicting functional outcomes after acute ischemic stroke is critical for prognosis and treatment planning, as it could help identify patients most likely to benefit from interventions. These outcomes are typically measured by the 90-day modified Rankin Scale (mRS90), a seven-level disability score. However, most deep learning models for functional outcome prediction reduce this to a binary variable, losing information about intermediate recovery levels and limiting clinical interpretability. So far, only a few studies have attempted to model the full mRS90 range using multiclass classification, ignoring the ordered nature of the scale. To address this, we propose a multimodal deep learning model that predicts mRS90 from pre-treatment 4D computed tomography perfusion (CTP) imaging and clinical metadata using CORAL-based ordinal regression. Using data from 111 thrombectomy-treated patients, our results suggest that the proposed ordinal approach outperforms standard multiclass and linear regression models by better leveraging the ordered nature of stroke outcomes.
Vision impairment remains a major global health challenge, and early diagnosis of retinal diseases is essential to prevent avoidable vision loss. Color fundus photography is widely used for retinal assessment, but manual interpretation is time-consuming and prone to inter-clinician variability. Although deep learning has improved automated retinal disease detection, most classification models lack transparency. We introduce RetCond, a self-explanatory generative classifier designed to provide both accurate predictions and counterfactual visual explanations. RetCond is a diffusion model repurposed as a classifier whose generative capabilities enable counterfactual synthesis for built-in explanations. These counterfactual images depict the retinal changes associated with alternative diagnoses, offering insight into the model's decisions. We curated a diverse dataset of 19,565 images containing five retinal disease conditions to train and evaluate the model. Performance is assessed using standard classification metrics and by qualitatively and quantitatively evaluating the self-explanatory component. Comparisons are made against strong discriminative baselines. RetCond achieves a classification accuracy of 96.98%, comparable to RETFound (96.99%). In contrast to standard models, RetCond produces counterfactual images that highlight condition-specific retinal features, confirming its learned concepts and demonstrating its self-explanatory properties. These results show that RetCond matches state-of-the-art discriminative performance while addressing limitations of post-hoc interpretability techniques in standard deep learning models. RetCond, therefore, constitutes a contribution toward more trustworthy AI solutions in ophthalmology.
Purpose:Autism is a common neurodevelopmental condition (NDC) that is characterized by restricted, repetitive behaviors and social communication differences that can impact the daily functioning of individuals. The clinical diagnosis of autism can be challenging, mainly due to its behavioral variability and frequent co-occurrence with other NDCs. This study investigates the ability of machine learning-based classification models trained using multimodal neuroimaging data combined with feature-importance analyses to identify development-specific brain characteristics associated with autism. Approach:A total of 144 participants aged 5 to 18 years with structural MRI (sMRI), diffusion MRI (dMRI), and resting-state functional MRI (rs-fMRI) data available were obtained from the Autism Brain Imaging Data Exchange (ABIDE) database. Radiomic features were extracted from each MRI data modality and used to train support vector machine (SVM) classifiers to identify neuroimaging patterns associated with autism. Single MRI modality classifiers, as well as one combining all three modalities, were trained for comparison purposes. To investigate age-specific effects, the same approach was followed for three age sub-groups: younger children (5-11 years), adolescents (12-18 years), and the entire 5-18 years age cohort. Model performance was evaluated using leave-one-out cross-validation across 30 diagnosis-balanced data splits. Feature-importance analyses were conducted to identify the most important neuroimaging features for classification. Results:The classification accuracies of the unimodal models ranged from 68.3% to 75.3% for sMRI, from 69.3% to 77.6% for dMRI, and from 66.3% to 69.9% for rs-fMRI data across age groups. Among all single imaging modalities and age groups, dMRI showed the highest performance with a 77.6% accuracy in younger children (5-11 years). The multimodal approach improved classification performance when compared to the unimodal models in all age groups, achieving accuracies of 78.9%, 76.7%, and 70.5% in the younger, adolescent, and entire age cohorts, respectively. Our findings indicate that multimodal classifiers integrating complementary structural, microstructural, and functional imaging features result in a more comprehensive representation of brain features that strengthens model performance. The most informative brain regions for classification differed between children and adolescents while several diffusion-derived features significantly correlated with social responsiveness scores, emphasizing the clinical importance of studying white and gray matter microstructure in autism. Conclusions:This study demonstrates the potential of multimodal neuroimaging-based machine learning models to identify development-specific biomarkers associated with autism. The results highlight the value of integrating age-stratified analyses of multimodal neuroimaging to better capture autism-associated developmental brain characteristics. The framework adopted in this study could be extended to explore other NDCs in the future.
Despite an abundance of high-quality mouse neuroimaging data, anatomical divergence from humans limits how findings from this data can be translated into clinical insights or applications. In this work, we present a proof-of-concept for a deep learning-driven image-to-image translation framework that translates parcellations of 23 common regions between mouse and human brain MRI datasets. Our model utilizes species-specific pre-trained autoencoders with principal component analysis dimensionality reduction coupled with a multi-layer perceptron that learns mappings from mouse to human latent encodings. Moreover, a novel objective function enforces the latent mapping to maintain realism and volume correlation of common regions between the input mouse image and the translated human image. Our results demonstrate that the proposed method produces realistic and varied human brain parcellations (FID: 31.9 real-translated vs. 25.3 real-real) while preserving common features (strong correlation between translated regional volume ratios). Overall, this work presents the first framework with considerable potential for an extension to full cross-species neuroimage translation, turning abundant mouse imaging data into clinically relevant synthetic human data and enabling data augmentation when human data is scarce.
The rapid evolution of artificial intelligence (AI) and in particular machine learning (ML) for healthcare applications has opened many new exciting opportunities to automatically analyze increasingly complex clinical data, especially related to medical imaging, one of the biggest data contributors in health care. However, while machine learning models have demonstrated considerable potential in the research setting to improve and accelerate diagnostic accuracy, reduce clinician workload, and enable integration of multi-modal data, their real-world deployment in clinical settings remains limited in many cases. A major obstacle is that ML models trained on medical images consistently fail to generalize well and perform poorly when applied to data from institutions, scanners, and patient populations that were not or poorly represented in the training set. These performance gaps often stem from biological and non-biological variations and biases that are only spuriously but not causally correlated with the medical task of interest. However, it still remains an open question how such biases propagate through deep learning model architectures and shape learned representations. This invited paper summarizes our recent efforts to address these challenges through controlled bias experiments and distributed learning methods. More precisely, the Simulated Bias in Artificial Medical Images (SimBA) framework is introduced, which enables the generation of realistic brain MRI datasets with known and fully controllable morphological and intensity-based biases, thereby facilitating counterfactual analyses of how biases are encoded and used by deep learning models. Using SimBA, we demonstrate that standard convolutional neural networks encode biases across all layers and that shortcut learning depends on various factors, including spatial proximity, bias salience, and class prevalence. We further discuss distributed training strategies as a scalable solution to overcome structural barriers to data sharing and enhance deep learning model generalizability. Together, this work provides foundational insights toward developing robust, interpretable, and equitable ML solutions for the analysis of medical imaging data and downstream computer-aided diagnosis tasks.
Neurofibromatosis type 1 (NF1) is a multisystem disorder with wide-ranging clinical presentations. Patients with NF1 may manifest with macrocephaly, strokes, and cognitive deficits, abnormal neural development, and other neurologic symptoms. This study used an atlas-based approach to quantitatively examine structural and physiologic changes of the brain in children with NF 1. Children evaluated for NF1 over a 9-year period at a children’s hospital were retrospectively reviewed (n = 34). Children with intracranial tumors or prior strokes were excluded. Patients received diffusion-weighted imaging (DWI) and arterial spin labeling (ASL) perfusion imaging on a 3T MRI scanner. Using an atlas-based approach, quantitative assessment of regional brain volumes, median apparent diffusion coefficient (ADC), and cerebral blood flow (CBF) was performed for the cerebral cortex, thalamus, caudate, putamen, globus pallidus, hippocampus, amygdala, nucleus accumbens, brainstem, and cerebral white matter. Differences were tested between NF1 patients and 100 typically developing controls. Compared to controls, children with NF1 demonstrated significantly increased volume measurements in all brain regions (p < 0.001), significantly higher median ADC values in all structures except for the putamen and nucleus accumbens (p < 0.001), and significantly lower median CBF most notable in the cerebral white matter (p < 0.001), globus pallidus (p < 0.001), hippocampus (p < 0.001), amygdala (p < 0.001), and brainstem (p = 0.001). This study measured microstructural and physiologic brain changes in children with NF1 compared to typically developing children. Further studies are needed to elucidate the cellular and molecular basis for these differences. With further refinement, atlas-based quantitative MRI brain signatures may serve as useful biomarkers of neural development, cognitive dysfunction, and risks for vasculopathy-related strokes in children with NF 1.
BackgroundMild behavioral impairment (MBI), characterized by later-life emergence of persistent neuropsychiatric symptoms (NPS), is an early clinical indicator of dementia risk. Global MBI has been associated with Alzheimer's disease (AD) pathology; studies have also explored MBI domains. Prior work has linked MBI-apathy to AD cerebrospinal fluid (CSF) biomarkers, but whether associations are detectable using plasma-based biomarkers such as phosphorylated tau (p-tau) is unknown. Establishing such relationships is critical, as plasma biomarkers are more accessible than CSF.ObjectiveTo explore cross-sectional and longitudinal associations between MBI-apathy and plasma p-tau181 levels using Alzheimer's Disease Neuroimaging Initiative data.MethodsOlder adults with normal cognition or mild cognitive impairment were categorized as MBI-apathy (n = 69), non-MBI NPS (n = 112), and no-NPS (n = 215) based on Neuropsychiatric Inventory scores and symptom persistence over one year. Linear regression modelled cross-sectional associations between NPS group and plasma p-tau181, adjusting for age, sex, education, apolipoprotein E4 status, and Mini-Mental State Examination score. Hierarchical linear mixed-effects modelling assessed associations over two and three years, including time-by-NPS group interactions.ResultsMBI-apathy was associated with significantly higher plasma p-tau181 levels at baseline (24.05% [6.06-45.08%]; adjusted p = 0.014), and over two (26.46% [7.24-49.12%]; adjusted p = 0.012) and three years (29.28% [10.17-51.72%]; adjusted p = 0.004) compared to no-NPS. No significant associations were observed for non-MBI NPS. In sensitivity analyses, non-MBI apathy was not associated with plasma p-tau181 at baseline (-9.96% [-32.69-20.44%]; unadjusted p = 0.478).ConclusionsMBI-apathy is associated with elevated plasma p-tau181 cross-sectionally and longitudinally, supporting MBI-apathy as a potential proxy marker of tau pathology for early AD detection.
Purpose:Assessing treatment response in bone metastases from non-small cell lung cancer (NSCLC) remains a major clinical challenge, particularly for patients receiving immune checkpoint inhibitors (ICIs). The existing response criteria are not optimized for osseous disease, leading to inconsistent evaluation. We aimed to develop and validate a radiomics-based machine learning (ML) framework to non-invasively distinguish immunotherapy response categories-progression, stable disease, and partial response-in NSCLC patients with bone metastases. Approach:Chest computed tomography (CT) scans from 99 NSCLC patients were analyzed before and during ICI therapy. Bone structures were automatically segmented using TotalSegmentator, and 1051 radiomic features were extracted per time point. Clinical variables were incorporated as optional features. Three ML classifiers-random forest, XGBoost, and support vector machine-were trained using fivefold cross-validation. A multistep feature selection pipeline (correlation filtering, mutual information, recursive feature elimination, and ReliefF ranking) was applied. Model performance was evaluated using area under the curve (AUC), F 1 -score, accuracy, sensitivity, and specificity, with additional statistical testing using Kruskal-Wallis, Mann-Whitney U , bootstrapping, and permutation analysis. Results:Inter-rater agreement for radiological response categories was high (Cohen's kappa = 0.91). Post-treatment radiomic features yielded the best performance. The random forest model achieved an AUC of 0.94, an F 1 -score of 0.79, an accuracy of 0.79, a sensitivity of 0.80, and a specificity of 0.83. Clinical features did not meaningfully improve performance. Models based on the largest lesion showed lower accuracy than those using the overall response. Conclusions:Post-treatment CT radiomics captured therapy-induced skeletal changes and enabled differentiation of immunotherapy response categories in NSCLC bone metastases. These findings highlight radiomics as a non-invasive tool for response assessment and guiding personalized treatment strategies.