Brain age estimation from structural MRI is an effective approach for detecting abnormal neurodevelopment and neurodegeneration.However, most existing methods produce global biomarkers that lack tissue-level specificity and fail to leverage medical prior knowledge. To address these limitations, we propose T2AgeNet, a dual-path image-text framework for tissue-level brain age estimation that integrates anatomical features with clinical semantics. The framework first segments brain MRI to generate tissue-specific masks, forming the basis for localized age prediction. To further incorporate medical prior knowledge, the model first aligns visual features with personalized clinical descriptions to guide semantic understanding of tissue-level variation. In parallel, it transforms handcrafted aging-related features into textual representations through an auxiliary branch using a large language model, enabling enriched interpretation and representation. We evaluate T2AgeNet on five datasets spanning fetal development, preterm infants, Alzheimer's disease, and autism spectrum disorder. Results demonstrate accurate age estimation across diverse populations. On the OASIS-3 and ABIDE-I datasets, the model further identifies tissue-specific structural abnormalities consistent with known neurological patterns.
Functional magnetic resonance imaging (fMRI) has greatly advanced our understanding of neurodevelopment. However, head motion during fMRI acquisition remains a significant challenge, especially for pediatric subjects. Excessive head motion can introduce substantial artifacts into fMRI scans, degrading the accuracy of subsequent analyses. Although motion correction methods have been proposed to directly eliminate motion effects from fMRI signals, the resulting functional connectivity (FC), the key component in fMRI studies, still contains substantial motion-induced artifacts. Effective motion correction methods applicable to FC are therefore highly desirable but remain unexplored. To address this gap, given the complementary information provided by different brain atlas parcellation schemes, we propose a novel Multi-Atlas Representation Alignment Transformer (MARA-Former) for joint motion correction of FCs derived from multiple atlases by leveraging their intrinsic relationship. Specifically, (1) we develop an FC-specified conditional Transformer architecture. By conditioning our model on both the input high-motion degree and the target low-motion degree, it can adaptively extract motion-aware features and generate low-motion results, thereby enhancing its ability in handling different motion degrees. (2) We employ optimal transport to align the representations of different atlases and design an atlas fusion block to enable comprehensive information exchange and joint learning across atlases, thereby effectively integrating complementary information and intrinsic relationships across atlases to jointly correct multi-atlas high-motion FCs. Extensive experiments on 1,289 resting-state fMRI scans of infants demonstrate the superiority of our MARA-Former in generating low-motion FCs from varying high-motion inputs. Moreover, downstream experiments further validate its effectiveness in FC-based infant brain development analysis.
Cortical surface registration is essential for neuroimaging analysis, enabling anatomical correspondence across subjects. Previous advances in learning-based registration have predominantly focused on architectures with limited receptive fields, often overlooking the necessity of modeling the holistic structure of long-range cortical dependencies. In this work, we propose STU-Net, a hybrid convolutional neural network (CNN)-Transformer framework, to achieve robust and anatomically plausible registration. First, We design a Global Self-Attention (GSA) module into the spherical UNet bottleneck, enabling its effective application to spherical cortical surfaces. Second, we propose a novel coarse-tofine training strategy that leverages sulcal depth for robust global alignment and curvature for local refinement. Finally, we integrate a Rodrigues-based diffeomorphic transform to guarantee smooth, invertible, and topology-preserving deformations. Validated on the developing Human Connectome Project (dHCP) dataset, STU-Net demonstrates significant improvements in registration accuracy and consistency of cortical alignment across subjects, providing an efficient and anatomically plausible solution.
We propose an EEG-based framework for depression subtype assessment using emotion-modulated neural dynamics elicited by immersive virtual reality (VR). EEG was recorded from 70 participants (31 first depressive episode, FDE; 18 recurrent depressive episode, RDE; 21 control participants, HC) using a compact frontal montage (Fp1/Fpz/Fp2) during positive and negative VR conditions, focusing on the immediate post-stimulus regulation period. Subject-level features were analyzed with a linear mixed-effects model to probe group-by-condition interaction patterns, and a Bi-Emotional Siamese Network (BESN) was developed to model baseline-anchored within-subject deviations by fusing positive/negative reactivity streams with path-signature-based temporal encoding. Feature-level analyses revealed interaction signals suggestive of differential emotion-related modulation across groups. In subject-wise evaluation, BESN achieved 83.1% accuracy for HC versus FDE discrimination and 70.6% accuracy for three-class classification (HC versus FDE versus RDE), outperforming conventional machine-learning baselines. Robustness was further supported on an external public resting-state EEG dataset (MODMA), achieving 78.2% accuracy. These results suggest that baseline-anchored, emotion-modulated EEG dynamics combined with path-signature modeling provide an objective and generalizable computational approach for depression-related group discrimination.
White matter fiber bundle parcellation is crucial for understanding brain connectivity, yet faces challenges due to the enormous number of streamlines and the need for anatomically meaningful classification. Existing methods often struggle to utilize the sequential nature of streamlines efficiently and fail to balance local and global feature extraction, limiting their accuracy in complex white matter architectures. This paper presents the Streamline Signature Net, a novel deep learning framework that addresses these limitations. The key contributions include: (1) leveraging path signature transforms and a dual-branch network to encode the geometric and sequential properties of streamlines, (2) implementing multi-scale window slicing to extract both fine-grained local details and global trajectory patterns, and (3) introducing a dynamic attention mechanism to weight discriminative slices within streamlines. Comprehensive experiments were conducted on two public datasets: ORG-800, an 800-cluster white matter atlas, and 105HCP-72, a semi-automatically annotated dataset with 72 fiber bundle classes. Our proposed SSN achieved state-of-the-art performance(reaching 93.53 https://github.com/RenchZhao/Streamline_Signature_Net
Effective automated electrocardiogram (ECG) interpretation hinges on disentangling waveform morphology from rhythm dynamics, a challenge for existing multimodal models that often conflate these heterogeneous attributes and introduce semantic ambiguity. We introduce ECG-MTDA, a framework that explicitly decouples these components. It learns morphology-oriented representations via a PQRST-guided masked autoencoder, while separately modeling temporal dynamics using continuous wavelet transform. Crucially, we align the learned morphology with concise, label-conditioned textual descriptions generated by a large language model (LLM) using a contrastive objective, creating a semantically grounded embedding space. ECG-MTDA demonstrates superior performance on the PTB-XL and CPSC 2018 benchmarks (e.g., AUC 93.16 on PTB-XL Superclass), with statistically significant gains over a strong multimodal baseline. Furthermore, on a challenging in-house cohort (n=620) for short-term paroxysmal atrial fibrillation (pAF) progression prediction, the model achieves an AUC of 0.97 $\pm$ 0.02 with high sensitivity (0.80 $\pm$ 0.04) and specificity (0.98 $\pm$ 0.01). Ablation studies and qualitative analyses confirm the benefits of our decoupled design and morphology-text alignment. Our results demonstrate that this clinically-inspired decoupling strategy yields more precise and robust multimodal representations for complex ECG analysis, enhancing both diagnostic classification and near-term risk stratification.
Accurate assessment of postmenstrual age (PMA) is critical for evaluating neonatal brain development and predicting neurodevelopmental outcomes. This paper presents a novel multi-scale multi-modal MRI framework for neonatal brain age estimation that integrates body weight as a core feature alongside imaging data. Our approach leverages T2-weighted structural MRI (T2w) and fractional anisotropy (FA) of diffusion MRI (dMRI) employing a three-branch network to extract multi-scale anatomical and microstructural features from both modalities. A transformer module then performs weight-aware feature fusion, effectively combining imaging characteristics with clinical weight information for precise PMA prediction. Evaluated on the dHCP neonatal cohort, our method achieves a mean absolute error of 0.48 weeks in term-born infants. Furthermore, our analysis reveals significant brain maturation delays in preterm neonates, the degree of which correlates strongly with both gestational age and body weight. These findings demonstrate the framework's potential as an effective tool for assessing early brain development and identifying deviations from typical maturation trajectories.
Brain age estimation based on structural MRI can serve as a powerful tool for exploring the impact of abnormal neurodegeneration on brain structure. Existing methods primarily focus on biomarkers associated with global structural changes in the brain, making it difficult to localize abnormalities to specific brain tissues. Furthermore, these approaches generally allow models to autonomously learn the relationship between brain structure and brain age from the images, without incorporating effective medical prior knowledge. To address these limitations, we propose TissueAgeNet: a dual-path image-text architecture that integrates imaging data with medical prior information for brain-tissue-level age estimation. Specifically, the proposed method utilizes a segmentation network to extract brain tissue masks, which, along with MRI images, are input into a visual encoder. Concurrently, tissue-level prior attributes—such as morphological and signal intensity features—are transformed into textual representations by a Large Language Model and serve as inputs to the text-stream. Finally, the two modalities are fused via simple linear integration to achieve accurate brain-tissue-level age estimation. We validate our approach on three datasets, including those of fetuses, preterm infants and Alzheimer’s patients. The results demonstrate accurate age prediction across diverse populations, and on the OASIS-3 dataset (Alzheimer’s patients), we show that our model can identify structural neurodegeneration at the tissue level.
Current fMRI decoders face a performance-fidelity trade-off where efficient ID encoders outperform geometrically-aligned surface-based models. We argue this is an artifact of inefficient surface tokenization and the failure to use anatomy as a predictive signal. We present , a framework that improves surface-based decoding by reframing anatomical variation from a nuisance to a powerful inductive prior. NeurIPS unites two innovations: a for efficient geometric encoding, and a that explicitly models individual anatomy using cortical features. On the Natural Scenes Dataset, NeurIPS establishes a new state-of-the-art for surface decoders and achieves performance comparable to strong 1D baselines. This is achieved with unprecedented efficiency, as the model converges dramatically faster (). This efficiency enables rapid adaptation to new subjects using only of data and remains stable when scaling the training cohort (4 to 8 subjects). Ablations provide evidence that these gains are driven by the model's use of cortical features, not by memorizing subject IDs. By leveraging anatomical priors, NeurIPS provides a principled and scalable path toward robust, generalizable brain decoding.
Predicting the development of functional connectivity (FC) derived from resting-state functional MRI is pivotal for elucidating the intrinsic brain functional organization and modeling its dynamic development during infancy. Existing deep learning methods typically predict FC at a target timepoint from each available FC independently, yielding inconsistent predictions and overlooking longitudinal dependencies, which introduce ambiguity in practical applications. Furthermore, the scarcity and irregular distribution of longitudinal rsfMRI data pose significant challenges in accurately predicting and delineating the trajectories of early brain functional development. To address these issues, we propose a novel Triplet Cycle-Consistent Masked Autoencoder (TC-MAE) for the trajectory prediction of the development of infant FC. Our TC-MAE has the capability to traverse FC over an extended period, extract unique individual characteristics, and predict target FC at any given age in infancy with longitudinal consistency. Extensive experiments on 368 longitudinal infant rs-fMRI scans demonstrate the superior performance of the proposed method in longitudinal FC prediction compared with state-of-the-art approaches.
Precise segmentation of fetal brain tissues in MRI is essential for studying brain development and for the early diagnosis and treatment of neurological disorders. However, the complex and variable anatomy of the fetal brain, significant morphological changes at different gestational ages, and the low-quality MRI and inherent noise of fetal acquisition pose significant challenges. To address these, we propose a novel 3D Dilated Projection U-net segmentation framework, DP-Net, which incorporates large kernel convolutions, atrous convolution for receptive field expansion, and dual skip connections mechanism to enhance global semantic consistency. Specifically, we introduce a Dilated Projection Block (DPB) that leverages atrous convolution to capture global context across multiple anatomical regions without additional parameters. Furthermore, we propose a Dual Skip Connection (DSC) mechanism to maintain encoder-decoder global consistency by fusing low-level and projected high-level features, mitigating blind spots introduced by atrous convolution. Extensive experiments show that our method significantly outperforms state-of-the-art methods, demonstrating its robustness and effectiveness in addressing the challenges of fetal brain tissue segmentation.
Diffeomorphic transformation-based cortical surface reconstruction typically involves a series of deformation processes to extract the cerebral cortex from brain magnetic resonance images (MRI). While most methods are designed for adult brains using Neural Ordinary Differential Equations (NODE) with fixed step sizes, the neonatal brain, which exhibits dramatic changes in cortical folding patterns early in life, requires a more adaptive approach. To address this, we develop a dual-task framework to directly characterize the brain development trajectory through processes of cortical surface reconstruction. For white matter (inner surfaces), we employ an Age-Conditioned ODE with adaptive step sizes. It is initially trained on a limited set of longitudinal paired data to establish a coarse trajectory, which is then refined through sample training of single-point data and knowledge distillation. For the pial surfaces (outer surfaces), we position the midthickness surfaces as intermediates and employ a cycle-consistent semi-supervised training strategy to depict a coherent brain development trajectory between the inner and outer surfaces. Our approach achieves precise developmental prediction directly on triangular meshes. Furthermore, by enhancing interpretability at each stage of the deformation process, this approach improves the applicability of diffeomorphic transformation-based methods. The proposed method has demonstrated state-of-the-art performance in modeling developmental trajectories and cortical surface reconstruction within the developing Human Connectome Project dataset (dHCP). The Code will be available at https://github.com/SCUT-Xinlab/DDT-from-CSR.
How to harmonize site effects is a fundamental challenge in modern multi-site neuroimaging studies. Although many statistical models and deep learning methods have been proposed to mitigate site effects while preserving biological characteristics, harmonization schemes for multi-site resting-state functional magnetic resonance imaging (rs-fMRI), particularly for functional connectivity (FC), remain undeveloped. Moreover, statistical models, though effective for region-level data, are inherently unsuitable for capturing complex, nonlinear mappings required for FC harmonization. To address these issues, we develop a novel, flexible deep learning method, Mamba-based Residual Generative adversarial network (MR-GAN), to harmonize multi-site functional connectivities. Our method leverages the Mamba Block, which has been proven effective in traditional visual tasks, to define FC-specified sequential patterns and integrate them with a multi-task residual GAN to harmonize multi-site FC data. Experiments on 939 infant rs-fMRI scans from four sites demonstrate the superior performance of the proposed method in harmonization compared to other approaches.
Brain functional connectivity (FC) constructed from resting-state functional MRI (rs-fMRI) is the predominant method for studying brain functional organization of infants. Predicting the full dynamic developmental trajectory of infant FC from existing incomplete longitudinal data can enrich our understanding of brain function developmental patterns and mechanisms and help identify neurodevelopmental disorders. However, the scarcity of longitudinal infant functional MRI scans with frequent irregular missing data poses significant challenges in accurately predicting and delineating the dynamic trajectory of early normal and abnormal brain development. Moreover, existing deep learning methods typically predict FC at a single target timepoint from each available FC independently, overlooking longitudinal dependencies and yielding temporally inconsistent and inaccurate predictions during infancy. To this end, we propose a novel Triplet Longitudinal Masked Autoencoder (TL-MAE) for the prediction of the full dynamic developmental trajectory of infant FC. Specifically, we adopt the following novel strategies: 1) Creating a longitudinally consistent prediction strategy to ensure the temporal consistency and robustness in the FC generation process; 2) Introducing the FC-specified Masked Autoencoder to capture FC domain features and pre-training this model by leveraging large-scale high-quality data; 3) Developing a dual triplet network alongside an identity conditional module to disentangle entangled identity and age information, enabling individualized predictions at any given age. Experiments on 696 longitudinal infant fMRI scans from two datasets demonstrate that our method not only yields more accurate and temporally consistent predictions of FC developmental trajectories, but also excels at capturing individualized features compared to state-of-the-art techniques.
Deep learning models are vulnerable to adversarial attacks. Although various defense methods have been proposed, such as incorporating perturbations during training, removing them in preprocessing steps or using image-to-image mapping to counter these attacks, these methods often struggle to robustly defend against diverse adversarial attacks and may affect the model's predictions on normal samples. To address this issue, we propose an adversarial example defense method based on image transformation. First, we designed an image transformation combiner that integrates multiple image transformations for defending against adversarial examples, thereby enhancing the robustness of the method. Second, we divide the image into patches and apply different combinations of image transformations to each patch to ensure the retention of useful information and increase the flexibility of the transformations. We combined 12 geometric or color transformations using the image transformation combiner and tested it on adversarial examples generated from the MNIST, CIFAR- 10, and ImageNet datasets. Experimental results show that our method outperforms other advanced detection methods in terms of accuracy and effectively mitigates the impact of adversarial perturbations on the model.
Reconstructing high quality 3D Magnetic Resonance Imaging (MRI) fetal brain volumes from motion-corrupted 2D slices is critical for accurate prenatal diagnosis. However, existing methods face challenges such as incomplete brain extraction regions and severe motion artifacts. To this end, we propose D-SVG (Diffusion-Based Slice-to-Volume Generation), a novel generation framework that integrates diffusion models with Implicit Neural Representation (INR) for reliable fetal brain volume generation. Our approach consists of two processes: (1) diffusion process, where a pre-trained diffusion model generates consistent 3D volumes with prior anatomical knowledge, and (2) slice constrained fidelity process, where an INR-based reconstruction scheme iteratively refines the generated volume using the scanned MRI slices in three stages. We conduct different experiments on clinical and simulated fetal brain datasets, including quantitative experiments, ablation study, visual analysis, demonstrating its superior performance over existing methods. We also perform two downstream tasks to prove the effectiveness and quality of generated brain volumes. The code of this paper is available at https://github.com/SCUT-Xinlab/dsvg.
Research has shown that prenatal depression impacts newborn brain development, with previous studies primarily using statistical analysis. This paper introduces deep learning models to investigate this effect, using rsfNIRS data to predict maternal prenatal depression. Traditional fNIRS classification methods struggle with high-dimensional dependencies and global features, leading to unsatisfactory performance. To overcome this, we propose the SSS model, which leverages path Signature, temporal Sequence, and global Statistical features, achieving an accuracy of 73% and an F1-score of 84%. Our experiments reveal that using data from the entire brain yields the best classification results, suggesting that prenatal depression may affect the newborn's entire brain.
This paper proposes a novel multi-sensor fusion SLAM approach for mobile robots, which integrates data from an IMU, a 2D LiDAR, and an RGB-D camera within a unified fusion architecture. First, to address the challenges of inaccurate and inefficient visual point cloud acquisition, enhanced neural networks are employed to extract global-local visual feature point clouds. These are then matched with LiDAR point clouds to improve spatial consistency. Second, to overcome the limitations of conventional fuzzy adaptive unscented Kalman filters (UKF) - namely low positioning accuracy and poor adaptability - an improved fuzzy adaptive UKF algorithm is introduced, which fuses IMU data and point cloud matching results. The optimal robot pose and environmental map are subsequently obtained through joint optimization. The proposed approach is evaluated through both simulation and real-world experiments. Results demonstrate that the method achieves accurate pose estimation, precise localization and mapping performance, and strong robustness in complex environments. The proposed approach significantly outperforms existing LiDAR-Visual-IMU SLAM approaches.
The sparsity of the Fourier transform domain has been applied to magnetic resonance imaging (MRI) reconstruction in k -space. Although unsupervised adaptive patch optimization methods have shown promise compared to data-driven-based supervised methods, the following challenges exist in MRI reconstruction: 1) in previous k -space MRI reconstruction tasks, MRI with noise interference in the acquisition process is rarely considered. 2) Differences in transform domains should be resolved to achieve the high-quality reconstruction of low undersampled MRI data. 3) Robust patch dictionary learning problems are usually nonconvex and NP-hard, and alternate minimization methods are often computationally expensive. In this article, we propose a method for Fourier domain robust denoising decomposition and adaptive patch MRI reconstruction (DDAPR). DDAPR is a two-step optimization method for MRI reconstruction in the presence of noise and low undersampled data. It includes the low-rank and sparse denoising reconstruction model (LSDRM) and the robust dictionary learning reconstruction model (RDLRM). In the first step, we propose LSDRM for different domains. For the optimization solution, the proximal gradient method is used to optimize LSDRM by singular value decomposition and soft threshold algorithms. In the second step, we propose RDLRM, which is an effective adaptive patch method by introducing a low-rank and sparse penalty adaptive patch dictionary and using a sparse rank-one matrix to approximate the undersampled data. Then, the block coordinate descent (BCD) method is used to optimize the variables. The BCD optimization process involves valid closed-form solutions. Extensive numerical experiments show that the proposed method has a better performance than previous methods in image reconstruction based on compressed sensing or deep learning.