Precise brain segmentation is fundamental for quantitative neuroimaging analysis. However, most existing methods lack generalization across the human lifespan and diverse imaging modalities, limiting their utility for Comprehensive Brain Segmentation (CBS) (i.e., tissue segmentation, parcellation, and lesion labeling). To address this, we propose BrainSeg, a novel unified framework, for CBS by using large-scale datasets spanning the entire lifespan, with adaptability to diverse uni- and multimodal input scenarios without the need for retraining or finetuning. Comprehensive experiments are conducted on lifespan data ranging from 14 gestational weeks to 100 years of age, consisting of 45,998 multimodal scans from 26 datasets, which are further augmented by our proposed synthesis strategy. Systematic validation and in-depth analysis demonstrate that our BrainSeg can achieve state-of-the-art performance across all three core CBS tasks, with the averaged Dice ratios reaching up to 96.94% for tissue segmentation, 94.25% for brain parcellation, and 91.06% for lesion labeling in the internal validations. It maintains similarly high accuracy in external validations, with averaged Dice ratios achieving 94.01% for tissue segmentation, and 91.20% for brain parcellation, underscoring its robustness and generalizability across diverse conditions. In summary, BrainSeg serves as a versatile foundation tool, providing flexible and reliable analysis for large-scale neuroimaging studies.
Background : In vivo whole-cortex quantification of intracortical signal-defined layering on the routinely acquired structural MRI remains limited. Purpose : To develop and validate an automated framework to reconstruct three intracortical signal-defined layers from 5T three-dimensional (3D) T2-weighted fluid-attenuated inversion recovery (FLAIR) and to characterize whole-cortex morphometrics and regional organization across a prespecified cortical organizational framework. Materials and Methods : In this retrospective study, 5T 3D FLAIR images were acquired between February and July 2024. Brain Multi-Layer Surface Reconstruction (BrainMLSR) reconstructed three intracortical signal-defined layers, and derived intracortical layer thickness and surface area measures and ratios. Performance was evaluated against manual annotations and assessed for test-retest repeatability (n=13) and cross-site feasibility (n=2). Paired two-tailed t-tests and linear mixed-effects models were used. A proof-of-concept analysis compared Heschl's gyrus ratios between 19 patients with temporal lobe epilepsy (TLE) and 19 age-matched healthy controls (HC Results : A total of 270 healthy participants (mean age, 54.4±14.5 years; 146 men) were included. Agreement with manual hypointense-layer annotations was high (Dice, 0.960±0.003), and was similar in the cross-site dataset (Dice, 0.954±0.009). In the test-retest dataset, average symmetric surface distance was less than 0.1 mm. Across prespecified systems, thickness and surface area ratios varied by region; within an auditory-perisylvian hierarchy, banksSTS showed a localized turning point with an increased hyperintense layer thickness ratio and decreased hypointense layer thickness ratio, accompanied by inflections in surface area ratios (P < .001). In bilateral Heschl's gyrus, hypointense (left: 0.619±0.262 vs 0.881 ± 0.102; right: 0.607±0.310 vs 0.907±0.141 mm) and isointense (left: 0.406±0.225 vs 0.678± 0.128; right: 0.478±0.232 vs 0.808 ± 0.176 mm) layer thicknesses were lower in TLE than in HC (all P<.001). Conclusion : BrainMLSR enabled accurate and repeatable in vivo reconstruction of three intracortical signal-defined layers from a single 5T 3D T2-weighted FLAIR acquisition and provided whole-cortex boundary-based morphometry with interpretable regional organization. ### Competing Interest Statement The authors have declared no competing interest. National Natural Science Foundation of China, 82441023, U23A20295, 62131015, 82394432 China Ministry of Science and Technolog, S20240085, STI2030-Major Projects-2022ZD0209000, STI2030-Major Projects-2022ZD0213100 Shanghai Municipal Central Guided Local Science and Technology Development Fund, YDZX20233100001001 The Key R&D Program of Guangdong Province, China, 2023B0303040001
Automatic sleep staging methods designed on healthy datasets often suffer from performance decline when applied to patients with obstructive sleep apnea (OSA). We attribute We attribute this to two factors: First, OSA sleep exhibits atypical inter-stages dependencies that are not captured by short-context models; Second, many prior approaches rely on one or a few channels and thus ignore the rich multichannel relationships present in clinical Polysomnography (PSG) recordings. To address these issues, we propose TransGATNet, which fuses long-term temporal-frequency features using a Graph-Attentional Transformer. Specifically, each TransGAT layer applies a Transformer encoder to capture global channel context, followed by a top-k sparsified graph attention network to isolate the most informative inter-electrode relationships. On the Sleep-EDF-2018 dataset, our TransGATNet achieves 86.8
Attenuation correction is critical for PET imaging to correctly reflect physiological activity. Recent studies perform PET self-attenuation correction, enabling attenuation correction from PET data itself instead of using additional CT or MRI. These methods either predict attenuation-corrected PET (AC-PET) from non-attenuation-corrected PET (NAC-PET) directly or synthesize intermediate CT images to guide PET attenuation correction. However, cross-modality synthesis of CT from NAC-PET is challenging, especially for small yet clinically important lesions. To address these challenges, we propose the Lesion-aware Mutual Guidance Diffusion Model (LMGDM), a coarse-to-fine dual-branch diffusion model that jointly performs PET attenuation correction and CT synthesis, with particular focus on lesion regions. Specifically, we first generate coarse predictions of both AC-PET and CT using individual residual diffusion models. Subsequently, the coarse AC-PET and CT are jointly refined by a proposed dual-branch mutual guidance module to enable feature fusion of the AC-PET and CT branches. Moreover, a lesion-aware refinement module is embedded into the PET branch, encouraging the network to focus on regions with pathologically high uptake rather than physiologically high uptake. In addition, an attenuation prior learned from real CT images is also introduced to further enhance the fidelity of the synthesized AC-PET and CT. Extensive evaluation on eight data centers demonstrates strong superiority of our LMGDM over the state-of-the-art methods.
Mild cognitive impairment (MCI) is the prodromal stage of dementia involving complex interactions between the brain and peripheral organs. Emerging evidence indicates that heart dysfunction and gut microbiota dysbiosis contribute to MCI pathogenesis. Here, we present a framework integrating brain-heart-gut interactions using whole-body positron emission tomography (PET) to enhance brain-only diagnostic performance. Our brain-only model achieves diagnostic performance comparable to that of whole-body PET and shows promising generalizability across four datasets comprising 1,543 whole-body PET and 1,721 brain PET images. We identify key brain regions involving the limbic, parietal, frontal, and temporal cortices that engage the default mode, central autonomic, and sensorimotor networks. These regions, along with specific myocardium and distal colon, constitute an integrated brain-heart-gut metabolic network, underscoring multi-organ crosstalk mediated by neural, biochemical, and mechanical pathways. Overall, our generalizable framework not only shows great potential for clinical translation in MCI diagnosis but also provides broad applicability to other systemic diseases beyond MCI.
Super-resolution reconstruction (SRR) from motion-corrupted thick-slice stacks of fetal brain MR images is essential for comprehensive prenatal examination and precise quantification of fetal brain development. Conventional approaches for fetal brain SRR require human intervention to extract cerebral structures from 2D slices, and their performance is often limited by blurred and less informative acquisitions due to fast scanning or irregular fetal movement. In order to overcome these challenges, we propose an automatic fetal brain SRR framework by integrating both individual- and group-level priors of fetal brain MRI for improved SRR. Specifically, we develop a robust fetal brain extraction approach based on Segment Anything Model (SAM). The extracted fetal brain region is segmented into brain tissues to serve as anatomical prior for the follow-up SRR. The SRR is performed based on an iterative optimization scheme by alternatingly performing slice-to-volume registration and volumetric reconstruction. Specifically, we ingeniously integrate anatomical priors obtained by tissue segmentation into both slice-to-volume registration and volumetric reconstruction, which emphasizes boundary alignment during registration and mitigates misalignment stemming from indistinct boundaries of cerebral tissues. Furthermore, we harness the available longitudinal fetal brain atlases to serve as specific guidance for volumetric reconstruction, thereby enriching structural details of the reconstructed images and also circumventing reconstruction of outliers. Experimental results on 184 clinical fetal brain MR images show that our proposed framework largely outperforms state-of-the-art methods for fetal brain SRR quantitatively and qualitatively.
Aligning medical images with partial anatomical overlap presents a significant challenge for various clinical applications that involve comparison of images with varying Fields of View. However, most existing registration methods implicitly assume full anatomical overlap and rely on local similarity measures, which limits their effectiveness when non-overlapping conditions exist. In this paper, we present a keypoint-driven registration framework for partial-overlap medical images, which builds robust cross-image correspondences by embedding anatomical contextual information directly into keypoint descriptors. The framework first detects sparse but anatomically representative keypoints and encodes each with a learned descriptor. These descriptors are then iteratively enhanced through a Dual-scale Context Aggregation Module (DualCAM), which jointly models global structures and local details to enhance descriptor discriminability. This process produces descriptors enriched with anatomical context for more reliable correspondence matching. Furthermore, we introduce an overlap-aware guidance mechanism that encourages the model to focus on the most reliable and anatomically consistent overlapping regions, mitigating interference from irrelevant non-overlapping regions and improving overall alignment accuracy. Extensive experiments conducted on two public multi-organ abdominal CT datasets demonstrate that our method surpasses state-of-the-art approaches across a wide range of overlap ratios, highlighting its robustness and strong generalization capability.
Many diseases are now recognized as systemic, involving extensive crosstalk between brain and body. Although existing reviews have elucidated the biochemical basis of these brain-body connections, their dynamic interactions remain largely unexplored by imaging approaches. To fill this gap, this review investigates brain-body axes from an imaging perspective across three research paradigms: 1) the impact of brain diseases on the body, including Alzheimer's disease, Parkinson's disease, psychiatric disorders, and cerebral small vessel disease; 2) the influence of body diseases on the brain, particularly those originating in the eye, heart, liver, and gut; and 3) the new frontier encompassing multi-organ data acquisition and integrative modeling methods for brain-body interactions. Spanning from biological hypotheses and inter-organ crosstalk to recent technical advances, this review not only summarizes the systemic manifestations of diseases across organs but also offers perspectives toward improved diagnostics and therapies. The integration of whole-body networks from an imaging perspective holds promise for advancing precision medicine and provides a new paradigm for investigating and diagnosing complex systemic diseases.
Brain extraction is a fundamental step in neuroimaging analysis, yet existing methods are often constrained to specific modalities or age groups, limiting their applicability in largescale, lifespan studies. In this work, we present BrainExt, a unified foundation model for brain extraction across multiple imaging modalities and the whole lifespan. To address inter-modality and across-age variations, we first collect a large-scale multi-modal, lifespan brain dataset of 25,487 scans spanning from fetal to elderly subjects and introduce a Structure-Intensity Disentanglement Synthesis (SIDSyn) module to generate real-world distribution-aligned data for robust pre-training. The pre-trained model is subsequently fine-tuned on real clinical scans to better adapt to real-world data distributions. BrainExt demonstrates superior generalization and stability compared to existing methods, achieving high Dice and low surface distance across all modalities. Empowered by disentanglement-based augmentation and a two-stage training strategy, BrainExt provides a scalable, modality-agnostic foundation for unified brain extraction, establishing a strong basis for advancing neuroimaging research and clinical applications.
AIMS:Brain structural (SC) and functional connectivities (FC) are intrinsically coupled, and this relationship declines in mild cognitive impairment (MCI). However, the underlying mechanisms of this decoupling remain unclear. Moreover, the impact of focal lesions (white matter hyperintensities, WMH) on decoupling remains unclear. Brain network communication models provide insights into this decoupling mechanism in MCI by estimating neural signal propagation along SC. METHODS:Five communication measures (CM) representing diffusion, path-accessibility, and routing-based communication strategies were derived from SC to estimate CM-FC coupling in 186 MCI and 171 cognitively normal subjects. Comparisons were conducted across groups and tractography types. Lesion mapping and mediation analyses were conducted to investigate how WMH influences cognition through coupling. RESULTS:MCI subjects exhibited significant whole-brain CM-FC decoupling, prominently within the default mode, sensorimotor, and visual networks. This decoupling was driven predominantly by diffusion-based communication strategies (communicability and flow graph), whereas routing-based coupling was relatively preserved. Compared with intact-fiber networks, WMH-impaired networks showed reduced diffusion-based coupling but relatively enhanced path-accessibility-based coupling, suggesting compensatory rerouting of information flow. Lesion mapping implicated WMH in the corona radiata as critical for disrupting CM-FC coupling, and mediation analyses demonstrated that diffusion-based coupling partially mediated the association between WMH burden and memory deficits. CONCLUSIONS:Our findings revealed diffusion-based CM-FC decoupling in MCI and introduced a novel framework for assessing the lesion effects of WMH. These findings provide insights into brain network dynamics and their alterations by white matter injury in MCI.
While Positron Emission Tomography (PET) is considered the gold standard for early AD diagnosis, its availability is limited due to high cost and radiation exposure risk, making it less accessible compared to more widely available modalities like tabular data and MRI. Therefore, effectively synthesizing PET features from available modalities presents a promising alternative and is of significant interest. In this work, we propose a novel multi-modal framework for synthesizing PET features based on tabular data-enhanced alignment. Our model requires only tabular and MRI data during the inference stage, yet achieves substantial improvement using synthesized PET features. Specifically, our framework consists of two stages. In the first stage, the model is pre-trained using multimodal data, with a tabular data-guided contrastive learning scheme designed to align features across different modalities. In the second stage, tabular data-guided Transformer blocks are used to synthesize PET features from MRI and tabular data based on the aligned encoders trained in the first stage. The synthesized PET features with tabular data and MRI, are then integrated for early AD diagnosis. Experimental results show that our model outperforms related state-of-the-art methods. This approach holds great promise for enhancing diagnostic accuracy and efficiency in AD diagnosis.Our code is available at https://github.com/internbob/Tab-s-AD .
Detection of Amyloid- $\beta$ ($\mathrm{A} \beta$) deposition is critical for timely diagnosis and intervention in Alzheimer's Disease (AD). Although $\mathrm{A} \beta$-Positron Emission Tomography ($\mathrm{A} \beta$-PET) enables in vivo visualization of $\mathrm{A} \beta$ pathology, its limited accessibility hinders routine clinical use. In contrast, T1-weighted Magnetic Resonance Imaging (T1w-MRI) is widely available but lacks pathological sensitivity. Recent MRI-to-PET synthesis methods intend to bridge this gap. Still, they often overemphasize anatomical differences and fail to capture pathology-specific signals due to incomplete disentanglement of anatomical, pathological, and modality alignments. To address this, we propose TriAlign, a two-stage $\mathrm{A} \beta$-PET synthesis framework that explicitly decouples these three alignments. In the first stage, a multimodal variational autoencoder enforces latent consistency between complete and PET-masked MRI-PET pairs, capturing anatomical and A $\beta$-related pathological features. Meanwhile, the decoder ensures modality alignment through $\mathrm{A} \beta$-PET reconstruction. In the second stage, a pathology-aware latent diffusion model refines pathology-specific residual between latent features of complete and masked MRI-PET pairs, guided by multi-scale anatomical cues and demographic prompts. Experiments demonstrate that TriAlign outperforms three widely used models in terms of image quality and $\mathrm{A} \beta$ status classification.
Mild cognitive impairment (MCI) is widely recognized as a highly heterogeneous and critical prodromal stage of dementia. Emerging evidence reveals that MCI is a systemic condition, with pathological and metabolic alterations manifesting in both the brain and peripheral organs. However, current computer-aided diagnostic methods primarily use brain-only images and often overlook pathological alterations in peripheral organs associated with MCI progression. To address this limitation, we propose MOGAD-Net, a multi-organ guided alignment and distillation network, which uses the fused systemic pathological features from the brain, heart, and gut to guide MCI detection in the training stage, while using only brain images in the application stage. Specifically, a panoramic Swin Transformer (PanSwin) encoder is introduced across all phases of MOGAD-Net to effectively capture global features across diverse organ images. Building on this architecture, MOGAD-Net first uses a semi-supervised multi-organ collaborative framework to distinguish patients with MCI from individuals with normal cognition (NC). This framework aligns heart and gut feature representations with brain feature representations via a hierarchical feature alignment module, facilitating inter-organ feature fusion. Furthermore, to improve clinical applicability, MOGAD-Net uses a hierarchical-constraint knowledge distillation module to transfer diagnostic knowledge from the multi-organ network to a brain-only model. Experiments on seven datasets show that MOGAD-Net significantly outperforms state-of-the-art methods in distinguishing MCI from NC.
Integrating multimodal radiological images and clinical data is critical for survival prediction in rectal cancer. However, existing methods often lack sufficient consideration of 1) modality heterogeneity (caused by rectal peristalsis, noise artifacts, and missing modalities) and 2) site heterogeneity (caused by different imaging protocols and patient populations). These factors hinder the model from capturing reliable cross-modal relationships and adapting to distribution shifts across clinical sites. In this work, we propose UICSurv, a novel multimodal Survival prediction framework highlighted by Uncertainty-guided Iterative Contrastive fusion, to capture robust cross-site multimodal interactions while leveraging sample-level uncertainty to enhance fusion reliability. Specifically, UICSurv initializes a shared multimodal embedding and iteratively refines it by fusing each heterogeneous modality via the cross-attention mechanism. In each iteration, a novel Survival Contrastive Learning (SCL) strategy is designed to progressively enhance both cross-site alignment and survival discriminability of the multimodal embedding space. Moreover, we design an EvidenceHit module, which employs temporally consistent evidential learning to jointly estimate survival probabilities and uncertainty. The estimated uncertainty further guides the embedding alignment by reducing the interference of unreliable samples. All components operate synergistically within UICSurv to reinforce reliable survival prediction in rectal cancer. Extensive experiments on multimodal datasets of rectal cancer (collected from three sites) demonstrate the superiority of our method both in survival prediction and uncertainty estimation. The code is available open-source: https://github.com/ScorpioBao/UICSurv.
Accurate diagnosis of brain disorders (BDs) is challenging in clinical practice. Most existing deep learning-based methods perform diagnosis only in a one-step manner, ignoring the step-wise, multi-level diagnosis processes as performed by radiologists. This oversight often leads to a high risk of misdiagnosis, especially for long-tail or challenging BDs. In this work, we introduce a Hierarchical Prompt and Prototype Learning (HP2L) framework for BD diagnosis, which emulates multi-level diagnostic procedures. HP2L explicitly captures hierarchical relationships among 23 BDs and groups them into three diagnostic levels: coarse classes (e.g., vascular lesions), intermediate classes (e.g., hemorrhage), and fine-grained classes (e.g., chronic hemorrhage). HP2L integrates three key innovations: (1) Hierarchical Prompting Vision Transformer (ViT) backbone, which performs coarse-to-fine feature extraction for step-wise BD classification; (2) Prompt Learning, which employs optimizable prompt tokens that encode diagnostic knowledge, guiding the classification at each level of the hierarchy; (3) Prototype Learning, which enriches the prompt token with BD-specific prototypes by injecting diagnostic information to enhance diagnosis performance. Extensive evaluations on 54,360 subjects across six multi-center datasets show that HP2L consistently outperforms state-of-the-art methods, achieving a balanced accuracy of 88.43% for both common and long-tail BDs, 8.42 percentage points higher than the best-performing benchmark. Furthermore, HP2L improves interpretability by aligning its predictions and attention visualizations with the clinical hierarchical reasoning process. The code and a portion of data (more data will be released after the decision of the paper) are available under: code, data.
Lengthy acquisition time remains a key bottleneck for the widespread use of MRI in clinics. While accelerated MRI can reduce scan duration, it often introduces increased noise, compromising image quality and diagnostic reliability. In this study, we present a unified deep learning-based denoising model for multi-organ accelerated MRI, designed to operate directly on reconstructed images from commercial MRI systems. Our model was trained on a prospectively collected, large-scale real-world dataset comprising 148,930 noisy-clean image pairs from six clinical centers and four major MRI vendors, spanning six organs and 96 MRI protocols. On a test set of 20,143 real-world image pairs, our model consistently outperforms state-of-the-art denoising methods. Importantly, downstream evaluation using tissue segmentation demonstrates a 7.05% improvement in Dice score across multiple organs compared to noisy images. The model further generalizes effectively to 46,870 external clinical images from four independent cohorts, highlighting its robustness across various scanners and acquisition protocols. To assess clinical utility, two experienced radiologists conducted blinded evaluations across multiple organs, focusing on overall image quality, diagnostic confidence, and disease diagnosis. The denoised images retained high visual fidelity and yielded diagnostic performance equivalent to clean images even with acceleration factor of 3× compared to clinical scanning setup, such that many acquisitions can be completed within one minute. This unified MRI denoising model holds great potential for various clinical applications.
Lesion localization and medical report generation are two fundamental yet complementary tasks for modern healthcare systems, jointly underpinning accurate diagnosis and effective clinical decision-making. Although both tasks have been separately reviewed in the literature, their interconnection is not well studied. The advent of generative artificial intelligence (AI) offers transformative potential for linking both tasks. In this review, we conduct a comprehensive survey of the recent advances in lesion localization and automatic report generation. For lesion localization, we examine the evolution from non-generative approaches to state-of-the-art generative foundation models. For report generation, we focus on lesion-aware report generation and encapsulate the methodologies spanning knowledge injection, grounding, and reasoning. We further summarize the widely used datasets and evaluation metrics, and highlight the key challenges alongside potential research directions. This review offers an integrated perspective by framing lesion localization and report generation as interdependent tasks within the framework of generative AI. Future directions should integrate both tasks in one unified system for more reliable and interpretable clinical usage.
In orthognathic surgery, post-surgical facial appearance plays a critical role in both patient quality of life and mandibular function. However, existing methods typically rely on labor-intensive skeletal surgical plans to predict post-surgical facial shapes, rather than guiding the planning process using appearance-based objectives. To address this limitation, we propose the Face Model Assisted Regression Network (FMR-Net), which leverages the FLAME parametric face model to generate anatomically coherent facial representations. The FLAME model provides a topologically consistent mesh structure, enabling the network to learn structured facial information across subjects through effective loss computation. Furthermore, we introduce a region-weighted surgical loss that incorporates clinical prior knowledge from surgeons, and a Laplacian smooth loss to ensure surface smoothness. Experiments on skeletal Class II orthognathic surgery data demonstrate that our method outperforms state-of-the-art approaches in both quantitative accuracy and qualitative results for facial shape prediction.
Mild cognitive impairment (MCI) is the prodromal stage of dementia involving complex interactions between the brain and peripheral organs. Emerging evidence indicates that heart dysfunction and gut microbiota dysbiosis can contribute to MCI pathogenesis. Yet, these discoveries of cross-organ interactions have not been applied to assist MCI diagnosis. In this work, we propose a novel diagnostic framework that exploits the interactions of brain, heart, and gut using whole-body PET images to guide MCI diagnosis for scenarios when only brain MRI, PET, or PET&MRI are available. Specifically, we collected a multi-cohort, multi-modal dataset comprising 1,545 whole-body PET images, 6,010 brain MR images, and 2,446 brain PET images from eight data centers. Organ-specific image encoders are first pretrained for the brain, heart, and gut individually. Then, to effectively align and integrate brain, heart, and gut features, we introduce positional prompts to act as anatomical-level attention to highlight disease-relevant spatial regions, and further develop hierarchical Transformers to model brain-heart, brain-gut, and brain-heart-gut interactions. Finally, to achieve MCI diagnosis using only brain images, we transfer the above brain-heart-gut model to a brain-only model via an introduced multi-level knowledge distillation scheme, including sample-level contrastive distillation, group-level distribution alignment, and response-level supervision. Extensive experiments on multi-center data demonstrate the superiority of our method over the state-of-the-art methods by resorting to effective integration of heart and gut interactions for MCI diagnosis.
Accurate brain parcellation from structural MRI across the human lifespan is essential for advancing neuroimaging and neuroscience studies. However, existing methods often struggle to generalize owing to intensity and contrast variations across brain maturation, aging and differences in MRI acquisition protocols, limiting their clinical and research utility. Here we present BrainParc, a unified parcellation framework that leverages anatomical information invariant to intensity and contrast, enabling accurate, robust and longitudinally consistent parcellation across a heterogeneous dataset without the need for fine-tuning. Extensive experiments on both internal and external datasets demonstrate that BrainParc substantially outperforms state-of-the-art methods in delineating 106 brain regions. BrainParc consistently shows better performance across diverse populations and imaging conditions, both quantitatively and qualitatively. Beyond anatomical segmentation, we show that BrainParc enables reliable tracking of brain development and facilitates early diagnosis of neurological disorders, underscoring its potential as a robust and generalizable tool for large-scale neuroimaging studies and clinical translation.