Voxel-based morphometry (VBM), a popular approach in neuroimaging research, uses magnetic resonance imaging data to assess variations in the local density of brain tissue and to examine its associations with biological and psychometric variables. Here we present deepmriprep, a preprocessing pipeline designed to leverage neural networks to perform all the necessary preprocessing steps for the VBM analysis of T1-weighted magnetic resonance imaging. Utilizing the graphics processing unit, deepmriprep is 37 times faster than CAT12, the leading VBM preprocessing toolbox. The proposed method matches CAT12 in accuracy for tissue segmentation and image registration across more than 100 datasets and shows strong correlations in the VBM results. Tissue segmentation maps from deepmriprep have more than 95% agreement with ground-truth maps, and its nonlinear registration predicts smooth deformation fields comparable to CAT12. The high computational speed of deepmriprep enables rapid preprocessing of large datasets and opens the door to real-time applications.
Penile measurement is clinically relevant across male reproductive and urogenital health, including conditions such as micropenis, congenital and endocrine disorders, and sexual or urinary dysfunction. However, quantitative assessment of penile size has relied mainly on external length or circumference measurements, which are difficult to standardize, sensitive to measurement conditions, and unable to capture the internal portion of the penis. MRI enables volumetric assessment of the whole penis in vivo, but automated segmentation has not previously been established at population scale. Automated whole-organ volumetry would enable high-throughput phenotyping for multi-omics and clinical studies of male reproductive disease. Here, we present a deep learning framework for whole-penis segmentation in multi-channel DIXON MRI. Using a newly curated expert-annotated training dataset (n = 145 subjects; 13,050 annotated slices) and a double-annotated independent test benchmark (n = 24 subjects; 2,160 double-annotated slices), we optimized a 3D nnU-Net architecture. The model achieved a 5-fold cross-validation Dice score of 0.90 and performed at observer-level accuracy on the independent test set (Dice: 0.92; Hausdorff distance: 3.58). We deployed the model in 34,412 UK Biobank participants, enabling automated quantification of total penile tissue, including both external and internal components. Longitudinal evaluation in 2,282 men demonstrated high inter-session reproducibility (r = 0.87). This framework establishes a reproducible and population-scalable method for MRI-based assessment of penile anatomy and provides an open technical resource for future studies in urological imaging and male reproductive health. The trained model weights will be publicly released.
BACKGROUND:The brain age biomarker estimates biological age from brain structure and is discussed as a potential screening tool for clinically relevant brain aging patterns in individuals. For brain age estimates to be of clinical utility, they must be meaningful for individual patients and free from systematic bias. Here, we investigate how biases from training data age skewness, termed distribution bias, impact the reliability and biological interpretability of this promising biomarker. METHODS:Using Monte Carlo simulations with data from 9305 individuals and external validation in neuropsychiatric cohorts (1345 individuals), we trained 100 brain age models for each of the 4 differently age-skewed training distributions, respectively. For each model, we evaluated predictive performance, conducted standard group-level analyses for different neurodegenerative and psychiatric diseases, and evaluated the clinical utility of the prediction as an individual risk marker. RESULTS:Training data age distribution significantly influenced model predictions, causing substantial fluctuations in predicted brain ages across the aging continuum. Statistical analyses revealed that these fluctuations impacted effect sizes and statistical significance across all diseases. Moreover, we found limited effectiveness of the brain age gap (BAG) as an individual risk marker and different levels of disease-associated brain age across the aging continuum. CONCLUSIONS:Skewed training data age distributions significantly impact brain age model predictions and may compromise scientific results. Based on our findings, we want to raise awareness about distribution bias and propose agewise interpretation of BAGs as a practical solution for robust research and meaningful clinical application.
Sampling error yields exclusively reactive, non-lesional brain parenchyma in a significant proportion of intracranial biopsies, leaving the underlying disease undiagnosed. We benchmark four pathology foundation models (UNI2-h, Virchow2, Prov-GigaPath, H-optimus-0) as frozen patch encoders within a shared attention-based multiple-instance learning framework using 245 whole-slide images from 186 patients with confirmed downstream diagnoses. We first show that coarse disease-category prediction can be reproduced largely from slide size alone. After restricting classification to three finer diagnostic distinctions within common tissue categories, this confound no longer explains performance, yet disease labels remain predictable above chance under permutation testing (p ≤ 10^-4 throughout). Surprisingly, performance is statistically indistinguishable across all foundation-model encoders, suggesting that recovering these weak morphological signatures is not limited by current patch representations. Signed instance-contribution maps and expert review further test whether predictive evidence localizes to reactive parenchyma rather than sampling-induced bias like blood introduced during tissue sampling. These results position acquisition-shortcut auditing via a provenance-only baseline as a necessary control in computational-pathology benchmarks, and show, once that confound is removed, that weakly supervised models still recover disease signal from tissue conventionally regarded as non-diagnostic.
Despite their promise, current neuroimaging biomarkers often fail to capture the full spectrum of inter-individual variability in brain structure and aging effects. This limits their ability to detect subtle norm deviations and impacts their utility for personalized care. We introduce Nearest Neighbor Normativity ( N 3 ) , a novel framework designed to resolve the confound between natural diversity and subtle pathological patterns. It evaluates individual brain structures from several meaningful viewpoints, accommodates a variety of co-existing normative prototypes and accounts for individually varying progression rates of brain structural decline. Using MRI data of 36,896 individuals, we provide empirical evidence that the N 3 biomarker effectively disentangles natural inter-individual variability from pathological alterations, significantly outperforming brain age models and traditional normative modeling approaches in the detection of neurodegenerative diseases. The N 3 framework is easily adaptable to various medical domains, fostering individualized and context-rich biomarkers and paving the way for more targeted and personalized therapeutic strategies.
Machine learning approaches pave a promising avenue to advance individual predictions about psychiatric illnesses, possibly using biomarkers. Here, we investigate longitudinal individual-level predictions of depressive relapse and the level of psychosocial functioning. Clinical variables (containing detailed symptom profile, previous disease course, as well as environmental and psychological protective and risk factors) and resting-state functional connectivity (rsFC) measures were used to predict relapse and the level of psychosocial functioning after a two-year follow-up interval in 346 patients (240 female) with Major Depressive Disorder (MDD). Random Forest machine learning models were computed to test the incremental predictive capability of clinical and rsFC data compared to a reference model containing confounding variables, and of a multimodal model (combining rsFC and clinical data) compared to the clinical model. Clinical information significantly predicted future psychosocial functioning beyond the reference model (21% versus 12% explained variance, p < 0.001). Depression relapse can be predicted by clinical information, however not significantly better than by the reference model alone (64.64% versus 57.70% balanced accuracy, p = 0.062). Resting-state data did not yield above-chance accuracies on its own (12% explained variance and 51.72% balanced accuracy) and did not hold incremental predictive value compared to clinical variables for either outcome. Baseline clinical information can be used to predict individual future psychosocial functioning, while rsFC patterns fall short of predicting clinical trajectories in MDD over a two-year interval. Sample size, model complexity and methodological considerations are discussed as potential sources of poor translation of MDD biomarkers.
We investigated whether the brain age gap (BAG)—the difference between chronological age and age estimated from structural MRI scans—is associated with long-term disease course in affective disorders, using a prospective nine-year follow-up design. T1-weighted MRI data were collected at two time points (mean interval = 8.98 ± 2.20 years) from patients with Major Depressive Disorder (MDD; N = 32), Bipolar Disorder (BD; N = 6), and healthy controls (HC; N = 37) across two sites. Using a brain age prediction model trained on a sample of over 10,000 subjects of the German National Cohort (GNC), we estimated individual BAG at baseline and follow-up using gray matter segments derived from MRI images. Employing linear-mixed-effects models, we tested main effects of diagnosis and hospitalizations as well as their interaction with time on BAG. In an exploratory analysis, we tested if BAG at baseline was predictive of hospitalizations during the nine-year follow-up using logistic regression and 10-fold nested cross-validation. MDD patients showed significantly higher BAG compared to HC (2.27 ± 5.68 vs. 1.00 ± 5.12 years, d = –0.23), while BAG in BD patients was descriptively elevated (4.71 ± 5.40 years). In the Münster subsample (N = 52), patients with at least one hospitalization had higher BAG than those without (4.16 ± 5.74 vs. 1.65 ± 5.41 years, d = –0.45). No group-by-time interaction was observed. Higher BAG at baseline predicted hospitalization during follow-up (p = 0.035), although cross-validated prediction accuracy (64.3%) did not reach significance (p = 0.071). BAG remained stable over time and was not influenced by future recurrence, supporting its role as a potential trait-like marker of vulnerability to illness recurrence. While exploratory, these findings suggest that BAG may capture individual risk for future hospitalization in affective disorders.
Major depressive disorder (MDD) is a complex psychiatric disorder that affects the lives of hundreds of millions of individuals around the globe. Even today, researchers debate if morphological alterations in the brain are linked to MDD, likely due to the heterogeneity of this disorder. The application of deep learning tools to neuroimaging data, capable of capturing complex non-linear patterns, has the potential to provide diagnostic and predictive biomarkers for MDD. However, previous attempts to demarcate MDD patients and healthy controls (HC) based on segmented cortical features via linear machine learning approaches have reported low accuracies. In this study, we used globally representative data from the ENIGMA-MDD working group containing 7012 participants from 31 sites (N = 2772 MDD and N = 4240 HC), which allows a comprehensive analysis with generalizable results. Based on the hypothesis that integration of vertex-wise cortical features can improve classification performance, we evaluated the classification of a DenseNet and a Support Vector Machine (SVM), with the expectation that the former would outperform the latter. As we analyzed a multi-site sample, we additionally applied the ComBat harmonization tool to remove potential nuisance effects of site. We found that both classifiers exhibited close to chance performance (balanced accuracy DenseNet: 51%; SVM: 53%), when estimated on unseen sites. Slightly higher classification performance (balanced accuracy DenseNet: 58%; SVM: 55%) was found when the cross-validation folds contained subjects from all sites, indicating site effect. In conclusion, the integration of vertex-wise morphometric features and the use of the non-linear classifier did not lead to the differentiability between MDD and HC. Our results support the notion that MDD classification on this combination of features and classifiers is unfeasible. Future studies are needed to determine whether more sophisticated integration of information from other MRI modalities such as fMRI and DWI will lead to a higher performance in this diagnostic task.
Macroscale brain modeling using neural mass models (NMMs) offers a framework for simulating human whole-brain dynamics. These models are pivotal for investigating the brain as a complex dynamic system, exploring phenomena like bifurcations, oscillatory patterns, and responses to stimuli. While connectome-based NMMs allow for the creation of personalized NMMs, their utility in capturing individual-specific neural characteristics remains underexplored, with current studies constrained by small sample sizes and computational inefficiencies. To address these limitations, we employed an algorithmically differentiable version of the reduced Wong Wang (RWW) model, enabling efficient optimization for large datasets. Applying this to resting-state fMRI data from 1444 samples, we optimized models with varying parameter complexities (n = 4, 658, and 23,875), which were derived from creating biologically plausible model variants. The optimized models achieved 4%, 19%, and 56% variance explanation in empirical functional connectivity (FC), respectively. Subject identification accuracy, based on simulated FC patterns, improved from < 1% (n = 4) to almost 100% (n = 23,875). Despite this precision, individual-level correlations between model parameters and attributes like age, gender, or intelligence quotient were small (effect sizes: η partial 2 ≤ 0.03 $$ {\eta}_{\mathrm{partial}}^2\le 0.03 $$ , standardized β ≤ 0.234 $$ \beta \le 0.234 $$ ). Machine learning analyses confirmed that these parameters lack the granularity to encode personal traits effectively. These findings suggest that, while current implementations of the RWW NMM can robustly replicate resting-state dynamics, the resulting parameters may lack the granularity required to map onto individual-specific behavioral metrics. This highlights a critical alignment problem: neural patterns and behavioral constructs such as intelligence may not correspond in a one-to-one fashion but instead represent higher-level abstractions. Bridging this gap will require the development of new tools capable of uncovering the underlying mapping manifolds, likely situated at the level of functional dynamics rather than isolated parameters. Future efforts should build on individual-level mechanistic modeling by exploring more expressive model classes and integrating richer sources of data, such as multimodal imaging or task-based paradigms, to better capture individual variability in both neural dynamics and behavioral traits. Such approaches may ultimately help to bridge the gap between model-based neural similarity and clinically meaningful personalization.
Objective Acute kidney injury (AKI) is a frequent complication in critically ill patients, affecting up to 50% of patients in the intensive care units. The lack of standardized and open-source tools for applying the Kidney Disease Improving Global Outcomes (KDIGO) criteria to time series, requires researchers to implement classification algorithms of their own which is resource intensive and might impact study quality by introducing different interpretations of edge cases. This project introduces pyAKI, an open-source pipeline addressing this gap by providing a comprehensive solution for consistent KDIGO criteria implementation. Materials and methods The pyAKI pipeline was developed and validated using a subset of the Medical Information Mart for Intensive Care (MIMIC)-IV database, a commonly used database in critical care research. We constructed a standardized data model in order to ensure reproducibility. PyAKI implements the Kidney Disease: Improving Global Outcomes (KDIGO) guideline on AKI diagnosis. After implementation of the diagnostic algorithm, using both serum creatinine and urinary output data, pyAKI was tested on a subset of patients and diagnostic accuracy was compared in a comparative analysis against annotations by physicians. Results Validation against expert annotations demonstrated pyAKI’s robust performance in implementing KDIGO criteria. Comparative analysis revealed its ability to surpass the quality of human labels with an accuracy of 1.0 in all categories. Discussion The pyAKI pipeline is the first open-source solution for implementing KDIGO criteria in time series data. It provides a standardized data model and a comprehensive solution for consistent AKI classification in research applications for clinicians and data scientists working with AKI data. The pipeline’s high accuracy make it a valuable tool for clinical research and decision support systems. Conclusion This work introduces pyAKI as an open-source solution for implementing the KDIGO criteria for AKI diagnosis using time series data with high accuracy and performance.
Importance:Soft drink consumption is linked to negative physical and mental health outcomes, but its association with major depressive disorder (MDD) and the underlying mechanisms remains unclear. Objective:To examine the association between soft drink consumption and MDD diagnosis and severity and whether this association is mediated by changes in the gut microbiota, particularly Eggerthella and Hungatella abundance. Design, Setting, and Participants:This multicenter cohort study was conducted in Germany using cross-sectional data from the Marburg-Münster Affective Cohort. Patients with MDD and healthy controls (aged 18-65 years) recruited from the general population and primary care between September 2014 and September 2018 were analyzed. Data analyses were conducted between May and December 2024. Main Outcomes and Measures:Primary analyses included multivariable regression and analysis of variance (ANOVA) models examining the association between soft drink consumption and MDD diagnosis and symptom severity, controlling for site and education, and Eggerthella and Hungatella abundance, controlling for site, education, and library size. Mediation analyses tested whether microbiota abundance mediated the soft drink-MDD link. Results:A total of 405 patients with MDD (275 female patients [67.9%]; mean [SD] age, 36.37 [13.33] years) and 527 healthy controls (345 female controls [65.5%]; mean [SD] age, 35.33 [13.13] years) were included. Soft drink consumption predicted MDD diagnosis (odds ratio [OR], 1.081; 95% CI, 1.008-1.159; P = .03) and symptom severity (P < .001; partial η2 [ηp2], 0.012; 95% CI, 0.004-0.035), with stronger effects in women (diagnosis: OR, 1.167; 95% CI, 1.054-1.292; P = .003; severity: P < .001; ηp2, 0.036; 95% CI, 0.011-0.062). In women, consumption was linked to increased Eggerthella (P = .007; ηp2, 0.017; 95% CI, 0.0002-0.068), but not Hungatella abundance. Mediation analyses confirmed that Eggerthella significantly mediated the soft drink-MDD association (diagnosis: P = .011; severity: P = .005), explaining 3.82% and 5.00% of the effect, respectively. Conclusions and Relevance:In this cohort study, it was found that soft drink consumption may contribute to MDD through gut microbiota alterations, notably involving Eggerthella. Public health strategies to reduce soft drink intake may help mitigate depression risk, especially among vulnerable populations; in addition, interventions for depression targeting the microbiome composition appear promising.
The biological aging process exhibits heterogeneous effects on different tissues, manifesting as tissue-specific variations in structural integrity and functional decline. Previously developed models are able to predict age from DNA methylation in the blood and the difference between estimated epigenetic age and chronological age is suggested to reflect accelerated or decelerated biological aging. While most prior studies have focused on the association between epigenetic age acceleration and global cortical thickness, it remains to be determined whether biological aging varies across specific cortical regions. This study aimed to assess associations between epigenetic age acceleration and regional cortical thickness, brain age gaps, as well as intra- and interindividual neuroanatomical heterogeneity in 756 participants of the BiDirect Study, including 430 participants from the general population and a cohort of 326 individuals with depression. Epigenetic age was estimated from whole blood DNA methylation data using the GrimAge algorithm. We observed an association of epigenetic age acceleration with cortical thinning across almost all cortical regions, suggesting a global association without regional confinement. This result was additionally underpinned by showing that accelerated epigenetic aging was also associated with increased interindividual neuroanatomical heterogeneity in contrast to a lack of association between epigenetic aging acceleration and intraindividual neuroanatomical heterogeneity. Accelerated epigenetic aging was furthermore paralleled by higher neuroimaging-based brain age gaps, suggesting at least partly shared aging processes. Together, these findings highlight that accelerated epigenetic aging reflects a global rather than region-specific neuroanatomical aging process, linking molecular and structural markers of brain aging and underscoring the potential of epigenetic clocks as biomarkers for brain health and neurodegenerative risk.
Disclosure: P. Beeken: None. J. Ernsting: None. L. Ogoniak: None. J. Kockwelp: None. T. Hahn: None. B. Risse: None. A.S. Busch: None. Background: The genetic factors influencing male fertility remain elusive. While testicular volume closely mirrors quantitative spermatogenesis, testicular size serves as a reliable proxy for male reproductive capacity. Previous approaches have faced limitations, as large-scale assessment has been hindered by traditional invasive measurement techniques, resulting in biased samples and restricted applicability. Despite its significance, the genetic basis of testicular volume remains largely unexplored. Methods: We applied machine learning to evaluate bi-testicular volume in 22,499 male participants from the UK Biobank using abdominal DIXON MRI scans. A U-Net model, trained on manual segmentations performed by medical experts, generated automated segmentations for the dataset. Following quality control, excluded incomplete segmentations and participants with conditions affecting segmentation, e.g. undescended testis, 18,998 participants were included. A genome-wide association study (GWAS) on bi-testicular volume was conducted using PLINK adjusting for age and the first ten principal components. Results: The U-Net model for MRI segmentation achieved Dice score of 0.87 (median), enabling precise measurement of testicular volume. The mean (SD) bi-testicular volume was 52 (16) mL. The GWAS identified 14 genome-wide significant loci (p<5×10⁻⁸), offering new insights into the genetic determinants of testicular volume. The most significant association was at rs12271187 within the FSHB locus on chromosome 11 (p = 1.33×10⁻⁴¹), with additional findings at the FSHR locus on chromosome 2, emphasizing the role of follicle-stimulating hormone and its receptor in male reproductive health. Conclusion: Circumventing previous challenges in assessing testicular volume, our study offers a large-scale approach, effectively addressing the limitations of traditional methods. By identifying genetic loci associated with testicular volume, we provide a foundation for further research into the genetic determinants of male fertility and reproductive health. Presentation: Monday, July 14, 2025
Testis size is known to be one of the main predictors of male fertility, usually assessed in clinical workup via palpation or imaging. Despite its potential, population-level evaluation of testicular volume using imaging remains underexplored. Previous studies, limited by small and biased datasets, have demonstrated the feasibility of machine learning for testis volume segmentation. This paper presents an evaluation of segmentation methods for testicular volume using Magnetic Resonance Imaging data from the UKBiobank. The best model achieves a median dice score of 0.89, compared to median dice score of 0.85 for human interrater reliability on the same dataset, enabling large-scale annotation on a population scale for the first time. Our overall aim is to provide a trained model, comparative baseline methods, and annotated training data to enhance accessibility and reproducibility in testis MRI segmentation research.
Background: We investigated associations of the brain age gap (BAG), the difference between actual and estimated age derived from MRI scans, with disease course over nine years in patients with affective disorders in a long-term prospective design. Methods: At two time-points, we acquired T1-weighted MRI images (mean [SD] follow-up period 8.98 [2.20] years) of patients with Major Depressive Disorder (MDD; N=32) and Bipolar Disorder BD; N=6) and healthy controls (HC; N = 37) at two sites (Dublin, Muenster). Using a brain age prediction model trained on a sample of over 10,000 subjects of the German National Cohort (GNC), we estimated individual BAG at two time-points (baseline and follow-up) using gray matter segments derived from MRI images. Employing linear-mixed-effects models, we tested main effects of diagnosis and hospitalizations during follow-up on BAG at baseline and follow-up, as well as their interaction with time respectively. In an exploratory analysis, we tested if BAG at baseline was predictive of hospitalizations during the nine-year follow-up using logistic regression and 10-fold nested cross-validation. Results: MDD patients showed a larger BAG compared to HC (MDD>HC: p=.039, MDD vs. BD: n.s.), while BD patients only showed a tendency for a larger BAG (p = .066). In the Muenster subsample (N=52), patients with hospitalizations showed a higher BAG compared to patients without hospitalizations (p=.001). No significant group-by-time interaction could be detected. However, higher BAG at baseline was associated with the number of hospitalizations during follow- up (p=.018), however, the cross-validation of our prediction with an accuracy of 64.3% was not significant (p=.071). Discussion: Our results show that BAG did not change over time as a function of patients' course of disease. The present study rather suggests that a higher estimation of biological aging (higher BAG) predicts future hospitalizations. Therefore, BAG may indicate a patient's vulnerability to future recurrence. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was funded by the German Research Foundation (DFG, grant FOR2107 DA1151/5-1 and DA1151/5-2 to UD; SFB-TRR58, Projects C09 and Z02 to UD) and the Interdisciplinary Center for Clinical Research (IZKF) of the medical faculty of Muenster (grant Dan3/012/17 to UD), and the Else Kroener-Fresenius-Stiftung (grant 2022_EKEA.102 to KF) as well as the Graduate Academy of the TU Dresden with funds of the Federal Ministry of Education and Research (BMBF) and the Freestate of Saxony under the Excellence Strategy of the Federal Government and the Laender (to KF). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The institutional review boards of the medical faculty of the University of Muenster and the Trinity College Dublin gave ethical approval for this work. All participants gave written informed consent according to the Declaration of Helsinki. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All raw MRI data in the present study are available upon reasonable request to the authors. All processed data and scripts are available online at https://osf.io/qadxz/.
Acute Kidney Injury (AKI) is a frequent complication in critically ill patients, affecting up to 50% of patients in the intensive care units. The lack of standardized and open-source tools for applying the Kidney Disease Improving Global Outcomes (KDIGO) criteria to time series data has a negative impact on workload and study quality. This project introduces pyAKI, an open-source pipeline addressing this gap by providing a comprehensive solution for consistent KDIGO criteria implementation. The pyAKI pipeline was developed and validated using a subset of the Medical Information Mart for Intensive Care (MIMIC)-IV database, a commonly used database in critical care research. We defined a standardized data model in order to ensure reproducibility. Validation against expert annotations demonstrated pyAKI's robust performance in implementing KDIGO criteria. Comparative analysis revealed its ability to surpass the quality of human labels. This work introduces pyAKI as an open-source solution for implementing the KDIGO criteria for AKI diagnosis using time series data with high accuracy and performance.