Major depressive disorder (MDD) involves complex interactions across multiple physiological systems, necessitating a comprehensive and integrative perspective for objective diagnosis and personalized treatment. Although multi-omics technologies provide broad molecular insights, their high cost and complexity limit clinical translation. We developed a cost-effective and practical machine learning framework using digital biomarkers from dried serum near-infrared spectroscopy (NIRS) for MDD diagnosis and antidepressant treatment outcome prediction. A total of 126 MDD patients and 86 healthy controls were enrolled. Six feature selection methods were systematically compared via rigorous nested cross-validation, and the final models were constructed using partial least squares discriminant analysis, evaluated with decision curve analysis (DCA), and deployed as a clinical diagnostic and decision-support web tool. Dried serum NIRS provided sufficient biochemical information to support the development of accurate diagnostic and treatment outcome prediction models. Competitive adaptive reweighted sampling (CARS) combined with SHapley Additive exPlanations (SHAP) emerged as the best-performing approach for identifying decisive wavelengths. The diagnostic model achieved an AUC of 0.91 (95
While longitudinal brain PET imaging is the gold standard for quantifying the spatiotemporal accumulation of Beta-amyloid, its widespread clinical utility is constrained by high operational costs and cumulative radiation risks. Recent deep generative models show promise in longitudinal image synthesis; however, they often fail to capture subtle pathological progression due to identity drift and a persistent bias toward trivially replicating baseline signal intensities rather than modeling temporal transition. To this end, we propose Delta-Diffusion, a novel progression-aware framework that redefines longitudinal PET synthesis as a conditional Poisson Diffusion Bridge (PDB) process. Unlike standard diffusion models that start from Gaussian noise, our PDB formulation is mathematically anchored to the subject's baseline PET, effectively transforming the generative task into a conditional distribution transition of the amyloid trajectory. To handle heteroscedastic nature of PET imaging, we introduce a physically-grounded Poisson perturbation within a Diffusion Transformer (DiT). This architecture uses adaptive scale-shift modulation to precisely calibrate the synthesis with the elapsed clinical interval and structural MRI context. A volume-of-interest balanced objective is designed to emphasize sparse, high-risk regions of amyloid accumulation. Validated on two cohorts with 542 subjects, Delta-Diffusion demonstrates superior performance in capturing longitudinal variations in amyloid deposition compared to state-of-the-art methods, offering a robust computational framework for tracking disease progression.
Accurate identification of late-life depression (LLD) using structural brain MRI is essential for monitoring disease progression and facilitating timely intervention. However, existing learning-based approaches for LLD detection are often constrained by limited sample sizes (e.g., tens), which pose significant challenges for reliable model training and generalization. Although incorporating auxiliary datasets can expand the training set, substantial domain heterogeneity, such as differences in imaging protocols, scanner hardware, and population demographics, often undermines cross-domain transferability. To address this issue, we propose a Collaborative Domain Adaptation (CDA) framework for LLD detection using T1-weighted MRIs. The CDA leverages a dual-branch architecture integrating a Vision Transformer (ViT) and a Convolutional Neural Network (CNN) to exploit their complementary strengths in capturing global anatomical topology and local texture details. The CDA framework consists of three stages: (a) supervised training on labeled source data, (b) self-supervised target feature adaptation to refine decision boundaries, and (c) collaborative training on unlabeled target data. The collaborative stage employs a reliability-aware JSD-based dual-constraint mechanism to filter high-quality pseudo-labels, encouraging robust prediction consistency without target supervision. Extensive experiments on two multi-site benchmarks demonstrate that CDA consistently outperforms state-of-the-art unsupervised domain adaptation methods, highlighting its superior generalization ability for real applications.
Ultrasound is a non-invasive, real-time, and cost-effective imaging technique widely used in clinical diagnosis. However, its diagnostic efficacy is often compromised by inherent speckle noise that degrades image quality and obscures underlying anatomical structures. Existing speckle reduction methods tend to over-smooth tissue boundaries and generalize poorly to heterogeneous noise levels. To address these limitations, we propose a Noise-Aware Boundary-Enhanced Generative Learning (NBGL) framework for ultrasound speckle reduction, which simultaneously preserves annotated anatomical boundaries and adapts to varying noise levels. The NBGL framework consists of a speckle reduction branch and a boundary enhancement branch. The former leverages generative learning to suppress speckle noise, while the latter learns boundary-sensitive representations to preserve target anatomical structures. Furthermore, a noise-aware interaction weight generation (NIWG) module estimates the speckle noise level via 3D Laplacian filtering and a median absolute deviation estimator, and translates it into an adaptive interaction weight. This weight is incorporated into a weighted feature-wise linear modulation (wFiLM) module to adaptively modulate cross-branch feature coupling, thereby improving robustness to varying noise levels. Extensive evaluations on 141 3D transvaginal ultrasound volumes demonstrate that NBGL consistently outperforms state-of-the-art methods in speckle reduction and structural preservation across six noise levels, while maintaining consistency with annotated anatomical boundaries.
Magnetic resonance imaging (MRI) and positron emission tomography (PET) are increasingly used in multimodal analysis of neurodegenerative disorders. While MRI is broadly utilized in clinical settings, PET is less accessible. Many studies have attempted to use deep generative models to synthesize PET from MRI scans. However, they often suffer from unstable training and inadequately preserve brain functional information conveyed by PET. To this end, we propose a functional imaging constrained diffusion (FICD) framework for 3D brain PET image synthesis with paired structural MRI as input condition, through a new constrained diffusion model (CDM). The FICD introduces noise to PET and then progressively removes it with CDM, ensuring high output fidelity throughout a stable training phase. The CDM learns to predict denoised PET with a functional imaging constraint introduced to ensure voxel-wise alignment between each denoised PET and its ground truth. Quantitative and qualitative analyses conducted on 293 subjects with paired T1-weighted MRI and 18F-fluorodeoxyglucose (FDG)-PET scans suggest that FICD achieves superior performance in generating FDG-PET data compared to state-of-the-art methods. We further validate the effectiveness of the proposed FICD on data from a total of 1,262 subjects through three downstream tasks, with experimental results suggesting its utility and generalizability.
Early identification of dementia at the early or late stages of mild cognitive impairment (MCI) is crucial for a timely diagnosis and early prevention of Alzheimer’s disease (AD). Recent studies have demonstrated that multimodal neuroimaging data provide complementary information of the brain, and the fusion of such data yields promising results for AD diagnosis. Most existing methods preserve data structural information by measuring pairwise similarity between samples based on conventional distances (e.g., Euclidean distance). However, these measures are unable to capture the dynamic structure of data due to the static characteristics of conventional distances. To this end, we propose a unified effective distance-based multi-modality feature selection method for AD analysis with multi-modality data. Specifically, we first adopt a sparse representation-based algorithm to compute the effective distance in each modality. Then, the effective distance-based Laplacian regularizer is introduced to preserve the structure information. Besides, a group sparsity regularization is introduced to learn the intrinsic relatedness among different modalities and jointly select the shared features across multiple modalities. Finally, a multi-kernel support vector machine is used to fuse the features selected from different modalities for final prediction. Extensive experiments on a benchmark database show the superiority of our method for disease diagnosis.
Aggregating multi-site brain MRI data can enhance deep learning model training, but also introduces non-biological heterogeneity caused by site-specific variations (e.g., differences in scanner vendors, acquisition parameters, and imaging protocols) that can undermine generalizability. Recent retrospective MRI harmonization seeks to reduce such site effects by standardizing image style (e.g., intensity, contrast, noise patterns) while preserving anatomical content. However, existing methods often rely on limited paired traveling-subject data or fail to effectively disentangle style from anatomy. Furthermore, most current approaches address only single-sequence harmonization, restricting their use in real-world settings where multi-sequence MRI is routinely acquired. To this end, we introduce MMH, a unified framework for multi-site multi-sequence brain MRI harmonization that leverages biomedical semantic priors for sequence-aware style alignment. MMH operates in two stages: (1) a diffusion-based global harmonizer that maps MR images to a sequence-specific unified domain using style-agnostic gradient conditioning, and (2) a target-specific fine-tuner that adapts globally aligned images to desired target domains. A tri-planar attention BiomedCLIP encoder aggregates multi-view embeddings to characterize volumetric style information, allowing explicit disentanglement of image styles from anatomy without requiring paired data. Evaluations on 4,163 T1- and T2-weighted MRIs demonstrate MMH's superior harmonization over state-of-the-art methods in image feature clustering, voxel-level comparison, tissue segmentation, and downstream age and site classification.
Multi-tracer positron emission tomography (PET) provides critical insights into diverse neuropathological processes such as tau accumulation, neuroinflammation, and β-amyloid deposition in the brain, making it indispensable for comprehensive neurological assessment. However, routine acquisition of multi-tracer PET is limited by high costs, radiation exposure, and restricted tracer availability. Recent efforts have explored deep learning approaches for synthesizing PET images from structural MRI. While some methods rely solely on T1-weighted MRI, others incorporate additional sequences such as T2-FLAIR to improve pathological sensitivity. However, existing methods often struggle to capture fine-grained anatomical and pathological details, resulting in artifacts and unrealistic outputs. To this end, we propose RelA-Diffusion, a Relativistic Adversarial Diffusion framework for multi-tracer PET synthesis from multi-sequence MRI. By leveraging both T1-weighted and T2-FLAIR scans as complementary inputs, RelA-Diffusion captures richer structural information to guide image generation. To improve synthesis fidelity, we introduce a gradient-penalized relativistic adversarial loss to the intermediate clean predictions of the diffusion model. This loss compares real and generated images in a relative manner, encouraging the synthesis of more realistic local structures. Both the relativistic formulation and the gradient penalty contribute to stabilizing the training, while adversarial feedback at each diffusion timestep enables consistent refinement throughout the generation process. Extensive experiments on two datasets demonstrate that RelA-Diffusion outperforms existing methods in both visual fidelity and quantitative metrics, highlighting its potential for accurate synthesis of multi-tracer PET.
Multimodal neuroimages, such as diffusion tensor imaging (DTI) and resting-state functional MRI (fMRI), offer complementary perspectives on brain activities by capturing structural or functional interactions among brain regions. While existing studies suggest that fusing these multimodal data helps detect abnormal brain activity caused by neurocognitive decline, they are generally implemented in Euclidean space and can't effectively capture the intrinsic hierarchical organization of structural/functional brain networks. This paper presents a hyperbolic kernel graph fusion (HKGF) framework for neurocognitive decline analysis with multimodal neuroimages. It consists of a multimodal graph construction module, a graph representation learning module that encodes brain graphs in hyperbolic space through a family of hyperbolic kernel graph neural networks (HKGNNs), a cross-modality coupling module that enables effective multimodal data fusion, and a hyperbolic neural network for downstream predictions. Notably, HKGNNs represent graphs in hyperbolic space to capture both local and global dependencies among brain regions while preserving the hierarchical structure of brain networks. Extensive experiments involving over 4,000 subjects with DTI and/or fMRI data demonstrate the superiority of HKGF over state-of-the-art methods in two neurocognitive decline prediction tasks. The proposed HKGF is a general framework for multimodal data analysis, facilitating objective quantification of brain structural or functional connectivity changes associated with neurocognitive decline.
Major depressive disorder (MDD) is a common mental disorder that typically affects a person's mood, cognition, behavior, and physical health. Resting-state functional magnetic resonance imaging (rs-fMRI) data are widely used for computer-aided diagnosis of MDD. While multi-site fMRI data can provide more data for training reliable diagnostic models, significant cross-site data heterogeneity would result in poor model generalizability. Many domain adaptation methods are designed to reduce the distributional differences between sites to some extent, but usually ignore overfitting problem of the model on the source domain. Intuitively, target data augmentation can alleviate the overfitting problem by forcing the model to learn more generalized features and reduce the dependence on source domain data. In this work, we propose a new augmentation-based unsupervised cross-domain fMRI adaptation (AUFA) framework for automatic diagnosis of MDD. The AUFA consists of 1) a graph representation learning module for extracting rs-fMRI features with spatial attention, 2) a domain adaptation module for feature alignment between source and target data, 3) an augmentation-based self-optimization module for alleviating model overfitting on the source domain, and 4) a classification module. Experimental results on 1,089 subjects suggest that AUFA outperforms several state-of-the-art methods in MDD identification. Our approach not only reduces data heterogeneity between different sites, but also localizes disease-related functional connectivity abnormalities and provides interpretability for the model.
Graph convolutional networks (GCNs) have been widely used and achieved remarkable results in brain disorder classification using functional connectivity (FC) graphs derived from resting-state functional magnetic resonance imaging (rs-fMRI). In GCN-based methods, the topological structure of FC graphs dominates feature aggregation, making it crucial for learning representative features. However, many existing approaches rely on static FC (sFC) with predefined topologies, which may limit the expressive power of GCNs. Furthermore, although dynamic FC (dFC) has been explored in some studies, the associated dynamic topological variations are often underutilized in disease recognition. To address these issues, we propose a novel topology-learnable static-dynamic graph convolution network (TSD-GCN) that adaptively learns topological structures from both static and dynamic FC graphs to capture complementary information for automated brain disorder identification. Specifically, TSD-GCN is designed as a dual-branch structure to comprehensively model the topological characteristics of each sample’s static and dynamic FC patterns. The static branch performs adaptive topology learning on sFC using a neural network layer to enhance the representational capacity of GCNs. The dynamic branch models topological variations of dFC by learning differential information across multiple consecutive time steps, thereby refining the dynamic topology and boosting feature expressiveness. Finally, a cross-branch collaborative block is employed to integrate holistic features from both branches for disease classification. Extensive experiments on two public datasets, ADNI and ABIDE, demonstrate that our method outperforms several state-of-the-art approaches, and the discovered discriminative FC patterns are biologically meaningful.
Brain tumor segmentation is crucial for accurate diagnosis and treatment planning, but the small sizes and irregular shapes of tumors pose significant challenges. Existing methods often fail to effectively incorporate medical domain knowledge such as tumor grade, which correlates with tumor aggressiveness and morphology, providing critical insights for more accurate detection of tumor subregions during segmentation. We propose an Automated and Editable Prompt Learning (AEPL) framework that integrates tumor grade into the segmentation process by combining multi-task learning and prompt learning with automatic and editable prompt generation. Specifically, AEPL employs an encoder to extract image features for both tumor-grade prediction and segmentation mask generation. The predicted tumor grades serve as auto-generated prompts, guiding the decoder to produce precise segmentation masks. This eliminates the need for manual prompts while allowing clinicians to manually edit the auto-generated prompts to fine-tune the segmentation, enhancing both flexibility and precision. The proposed AEPL achieves state-of-the-art performance on the BraTS 2018 dataset, demonstrating its effectiveness and clinical potential. The source code can be accessed online.
Radiology report generation is crucial in medical imaging, but the manual annotation process by physicians is time-consuming and labor-intensive, necessitating the development of automatic report generation methods. Existing research predominantly utilizes Transformers to generate radiology reports, which can be computationally intensive, limiting their use in real applications. In this work, we present R2Gen-Mamba, a novel automatic radiology report generation method that leverages the efficient sequence processing of the Mamba with the contextual benefits of Transformer architectures. Due to lower computational complexity of Mamba, R2Gen-Mamba not only enhances training and inference efficiency but also produces high-quality reports. Experimental results on two benchmark datasets with more than 210,000 radiograph-report pairs demonstrate the effectiveness of R2Gen-Mamba regarding report quality and computational efficiency compared with several state-of-the-art methods. The source code can be accessed online.
Learning to estimate and classify brain functional networks (BFNs) has become an increasingly important way of predicting neurological or mental disorders at their early stages. The traditional methods conduct BFN estimation and classification in two separate steps, thus preventing the interaction and joint optimization. In contrast, Transformer provides a natural architecture to learn BFNs with downstream tasks in an end-to-end manner. Despite their great potential, Transformer-based methods involve a large number of parameters that need to be learnt from big data and often lead to poor model interpretability. Considering the challenge in acquiring data and the high demand for model interpretability in medical scenarios, in this paper, we propose a minimalist Transformer architecture, referred to as Miniformer, by simplifying the projection matrices in the self-attention module into a single diagonal matrix, which greatly reduces the number of parameters, alleviates the risk of overfitting, and improves the interpretability. Additionally, the clear physical meaning of parameters in Miniformer makes the integration of domain knowledge or prior easier and more natural. Therefore, we further develop two variants of Miniformer by incorporating sparsity for removing potentially noisy time points from fMRI signals, and smoothness for capturing the temporal correlations in fMRI signals, respectively. To evaluate the effectiveness of the proposed methods, we perform brain disease diagnosis experiments on three public datasets. The results show that Miniformer and its variants tend to achieve higher classification performance than comparison methods with good interpretability.
Significant memory concern (SMC), a preclinical phase of Alzheimer’s disease (AD), is at increased risk of underlying AD pathology. Monitoring the progression of SMC is crucial for timely intervention of AD and related disorders. Learning-based neuroimage analysis provides a non-invasive and objective solution for SMC prognosis. However, existing studies usually focus on utilizing imaging data, without considering subjects’ demographic information which is essential for individual-level analysis. In addition, due to the characteristics of SMC progression requiring longitudinal analysis (e.g., 2 years), the data used for SMC analysis are usually very limited (e.g., tens), which poses a huge challenge to model training. To address these limitations, we propose a prompt-driven multi-view learning (PML) framework for predicting the clinical progression of SMC by integrating T1-weighted MRI with demographic information. Specifically, PML comprises four key components: (1) data-driven MRI feature extraction using a residual neural network to learn representative features from 3D MRI scans; (2) handcrafted MRI feature extraction to incorporate domain knowledge on brain tissues (e.g., cortical thickness); (3) demographic feature encoding using a prompt-based strategy through a contrastive language-image pretraining encoder; and (4) feature fusion and classification to jointly model multimodal information. To alleviate data scarcity challenges, we initialize our model with pretrained weights and employ transfer learning to enhance performance. Experimental results on a total of 469 subjects demonstrate the efficacy of PML in predicting SMC progression.
Domain Generalization (DG) has been widely used in image classification tasks to effectively handle distribution shifts between source and target domains without accessing target domain data. Traditional DG methods typically rely on static models trained on the source domain for inference on unseen target domains, limiting their ability to fully leverage target domain characteristics. Test-Time Adaptation (TTA)-based DG methods improve generalization performance by adapting the model during inference using target domain samples. However, this often requires parameter fine-tuning on unseen target domains during inference, which may lead to forgetting of source domain knowledge or reduce real-time performance. To address this limitation, we propose a Dynamic Decision Boundary-based DG (DDB-DG) method for image classification, which effectively leverages target domain characteristics during inference without requiring additional training. In the proposed DDB-DG, we first introduce a Prototype-guide Multi-lever Prediction (PMP) module, which guides the dynamic adjustment of the decision boundary learned from the source domain by leveraging the correlation between test samples and prototypes. To enhance the accuracy of prototype computation, we also propose a data augmentation method called Uncertainty Style Mixture (USM), which expands the diversity of training samples to improve model generalization performance and enhance the accuracy of pseudo-labeling for target domain samples in prototypes. We validate DDB-DG using different backbone networks on three publicly available benchmark datasets: PACS, Office-Home, and VLCS. Experimental results demonstrate that our method achieves superior performance on both ResNet-18 and ResNet-50, surpassing the state-of-the-art DG and TTA methods.
Multi-site structural MRI is increasingly used in neuroimaging studies to diversify subject cohorts. However, combining MR images acquired from various sites/centers may introduce site-related non-biological variations. Retrospective image harmonization helps address this issue, but current methods usually perform harmonization on pre-extracted hand-crafted radiomic features, limiting downstream applicability. Several image-level approaches focus on 2D slices, disregarding inherent volumetric information, leading to suboptimal outcomes. To this end, we propose a novel 3D MRI Harmonization framework through Conditional Latent Diffusion (HCLD) by explicitly considering image style and brain anatomy. It comprises a generalizable 3D autoencoder that encodes and decodes MRIs through a 4D latent space, and a conditional latent diffusion model that learns the latent distribution and generates harmonized MRIs with anatomical information from source MRIs while conditioned on target image style. This enables efficient volume-level MRI harmonization through latent style translation, without requiring paired images from target and source domains during training. The HCLD is trained and evaluated on 4158 T1-weighted brain MRIs from three datasets in four tasks, assessing its ability to remove site-related variations while retaining essential biological features. Qualitative and quantitative experiments suggest the effectiveness of HCLD over several state-of-the-arts.
Function-structure connectivity (FSC) coupling helps reveal alterations in the interplay between brain functional connectivity (FC) and structural connectivity (SC) caused by neurocognitive decline. Existing studies on FSC coupling typically focus on modeling interactions between static FC and SC features, ignoring temporal dynamics conveyed in functional MRI (fMRI) time series. Additionally, conventional strategies often compute global whole-brain FSC correlation or assess local region-specific FSC correspondences, without capturing complex inter-region dependencies between FC and SC patterns. To this end, we propose a dynamic function-structure connectivity coupling (DFSC) framework to predict progression trajectories in neurocognitive decline with fMRI and diffusion tensor imaging (DTI) data. In DFSC, we first construct static SC and dynamic FC graphs and use graph neural networks (GNNs) for feature learning, yielding new SC and FC embeddings. Based on these embeddings, we construct dynamic local-to-global FSC coupling graphs to capture both region-specific and inter-region dependencies between FC and SC, followed by GNNs to generate dynamic FSC coupling embeddings. These multi-view embeddings are finally fed into a squeeze-excitation readout module and a Transformer for feature fusion and prediction. Experimental results on two datasets with paired fMRI and DTI data from a total of 231 subjects demonstrate that our DFSC outperforms several state-of-the-art methods. With the DFSC, one can identify both discriminative brain regions and between-group FSC coupling difference, facilitating objective quantification of structural and functional brain changes associated with neurocognitive decline.
Domain Generalization-based Medical Image Segmentation (DGMIS) aims to enhance the robustness of segmentation models on unseen target domains by learning from fully annotated data across multiple source domains. Despite the progress made by traditional DGMIS methods, they still face several challenges. First, most DGMIS approaches rely on static models to perform inference on unseen target domains, lacking the ability to dynamically adapt to samples from different target domains. Second, current DGMIS methods often use Fourier transforms to simulate target domain styles from a global perspective, but relying solely on global transformations for data augmentation fails to fully capture the complexity and local details of the target domains. To address these issues, we propose a Dynamic Domain Generalization (DDG) method for medical image segmentation, which improves the generalization capability of models on unseen target domains by dynamically adjusting model parameters and effectively simulating target domain styles. Specifically, we design a Dynamic Position Transfer (DPT) module that decouples model parameters into static and dynamic components while incorporating positional encoding information to enable efficient feature representation and dynamic adaptation to target domain characteristics. Additionally, we introduce a Global-Local Fourier Random Transformation (GLFRT) module, which jointly considers both global and local style information of the samples. By using a random style selection strategy, this module enhances sample diversity while controlling training costs. Experimental results demonstrate that our method outperforms state-of-the-art approaches on several public medical image datasets, achieving average Dice score improvements of 0.58%, 0.76%, and 0.76% on the Fundus dataset (1,060 retinal images), Prostate dataset (1,744 T2-weighted MRI scans), and SCGM dataset (551 MRI image slices), respectively. The code is available online (https://github.com/ZMC-IIIM/DDG-Med).
Chest X-ray (CXR) images play a crucial role in diagnosing COVID-19 by facilitating the rapid identification of lung damage. Recently, deep hashing technology has enhanced our capabilities for retrieving large CXR image databases, providing healthcare professionals with more comprehensive tools for pandemic analysis. However, current methods typically depend on single-scale features to represent CXR images, without modeling the multi-scale contextual information present in the original images. Additionally, lung segmentation and image classification from CXR offer anatomical and semantic insights that could enhance the feature extraction of CXR images, but this remains largely unexplored in the field. To this end, we develop a segmentation-enhanced multi-scale deep hashing (SMDH) framework for automated CXR image retrieval, comprising a feature extraction module and an image retrieval module. The feature extraction module comprises a multi-scale neural network architecture, aiming to deeply mine and integrate key semantic information within CXR images through a multi-level feature fusion strategy. In particular, to capture rich anatomical and semantic information in CXR images, we utilize lung segmentation and image classification as auxiliary tasks to guide the feature extraction process. During retrieval, the trained feature extraction module is used to convert each input query CXR image and all CXR images in the retrieval database into binary hash codes, followed by Hamming distance-based ranking for fast image matching. Experiments on the COVID-QU-Ex dataset with 33,920 CXR images suggest that SMDH outperforms several state-of-the-art methods.