Precise brain segmentation is fundamental for quantitative neuroimaging analysis. However, most existing methods lack generalization across the human lifespan and diverse imaging modalities, limiting their utility for Comprehensive Brain Segmentation (CBS) (i.e., tissue segmentation, parcellation, and lesion labeling). To address this, we propose BrainSeg, a novel unified framework, for CBS by using large-scale datasets spanning the entire lifespan, with adaptability to diverse uni- and multimodal input scenarios without the need for retraining or finetuning. Comprehensive experiments are conducted on lifespan data ranging from 14 gestational weeks to 100 years of age, consisting of 45,998 multimodal scans from 26 datasets, which are further augmented by our proposed synthesis strategy. Systematic validation and in-depth analysis demonstrate that our BrainSeg can achieve state-of-the-art performance across all three core CBS tasks, with the averaged Dice ratios reaching up to 96.94% for tissue segmentation, 94.25% for brain parcellation, and 91.06% for lesion labeling in the internal validations. It maintains similarly high accuracy in external validations, with averaged Dice ratios achieving 94.01% for tissue segmentation, and 91.20% for brain parcellation, underscoring its robustness and generalizability across diverse conditions. In summary, BrainSeg serves as a versatile foundation tool, providing flexible and reliable analysis for large-scale neuroimaging studies.
Automatic pancreas segmentation can facilitate diagnosis and treatment of pancreatic diseases. The combination of non-contrast, arterial, and venous phases of CT imaging can enhance differentiation of the pancreas from its surrounding structures. However, existing multi modal methods, which try to integrate the multimodal in formation in computer-aided pancreas segmentation, often overlook the inter-modal relationships and have a limited capability for information fusion. In this paper, we propose a multi-phase pancreas segmentation method for incorporating Feature Aggregation Module (FAM) and Modality Adaptive Transformer (MAT). Specifically, we use the venous phase as the primary modality, while the non-contrast and arterial phases serve as supplementary modalities, based on clinical prior knowledge. Our FAM integrates spatial information from the primary and supplementary modalities, while our MAT adaptively enhances feature representation and establishes long-range dependencies among modalities. Our method outperforms state-of-the art techniques on a large scale dataset. Based on the segmented pancreas region, We further perform a down-stream task focused on pancreatic volume calculation. The prediction accuracy is on par with manual segmentation, demonstrating effectiveness and potential application of our proposed method.
Coronary artery disease poses a significant public health threat, and coronary computed tomography angiography is the preferred imaging modality for diagnosis and risk assessment of coronary artery disease through plaque evaluation. However, understandings of how atherosclerotic characteristics vary by age and sex remains limited due to challenges in manual quantitative plaque assessment. Here, we conducted a retrospective, consecutive, multi-center Chinese cohort study of 16,300 patients undergoing clinically indicated coronary computed tomography angiography that revealed multi-level quantitative patterns of atherosclerosis stratified by age and sex. We found that females experienced a delayed atherosclerosis onset by approximately 20 years compared to males, with plaque burden increasing nonlinearly with age and accelerating more evidently after menopause. The built coronary atlas identified plaque clusters, primarily within proximal segments of major coronary arteries, slightly upstream side branch bifurcations. Our findings provide deeper insights into coronary atherosclerosis in the Chinese population, supporting more tailored prevention strategies.
Current hemodynamic analyses predominantly utilize rigid-wall coronary models derived from end-diastolic coronary computed tomography angiography (CCTA). However, the physiological pulsatility of coronary flow and dynamic geometric variations induced by cardiac motion raise important questions about the validity of single-phase approximations. This study systematically investigates how coronary hemodynamic metrics - particularly wall shear stress (WSS), a key determinant of endothelial function and vascular pathophysiology - vary when computed using multi-phase versus conventional single-phase CCTA reconstructions. We analyzed 20-phase CCTA datasets acquired throughout the cardiac cycle from 10 patients, representing a significant advancement over standard singlephase protocols. Following precise coronary artery segmentation, we performed comprehensive numerical simulations for each phase-specific model. Our analysis focused on six critical hemodynamic indices in atherosclerosis-prone regions, with particular attention to low-WSS areas. Key findings reveal substantial phase-dependent and inter-patient variability in hemodynamic assessments, with normalized root mean square errors ranging from 0.0168 to 1.0299 across phase groups. These results demonstrate that CCTA phase selection significantly impacts simulation outcomes, underscoring the importance of modeling methodology in coronary hemodynamic studies. These findings may serve as a theoretical references for guiding coronary modeling methods and clinical application of hemodynamic metrics.
Vessel dynamics simulation is vital in studying the relationship between geometry and vascular disease progression. Reliable dynamics simulation relies on high-quality vascular meshes. Most of the existing mesh generation methods highly depend on manual annotation, which is time-consuming and laborious, usually facing challenges such as branch merging and vessel disconnection. This will hinder vessel dynamics simulation, especially for the population study. To address this issue, we propose a deep learning-based method, dubbed as DVasMesh to directly generate structured hexahedral vascular meshes from vascular images. Our contributions are threefold. First, we propose to formally formulate each vertex of the vascular graph by a four-element vector, including coordinates of the centerline point and the radius. Second, a vectorized graph template is employed to guide DVasMesh to estimate the vascular graph. Specifically, we introduce a sampling operator, which samples the extracted features of the vascular image (by a segmentation network) according to the vertices in the template graph. Third, we employ a graph convolution network (GCN) and take the sampled features as nodes to estimate the deformation between vertices of the template graph and target graph, and the deformed graph template is used to build the mesh. Taking advantage of end-to-end learning and discarding direct dependency on annotated labels, our DVasMesh demonstrates outstanding performance in generating structured vascular meshes on cardiac and cerebral vascular images. It shows great potential for clinical applications by reducing mesh generation time from 2 h (manual) to 30 s (automatic).
Accurate 3D reconstruction of the small bowel skeleton is vital for understanding intestinal morphology, de-tecting structural abnormalities, and supporting diagnosis, yet limited resolution, organ adhesion, complex anatomy, and scarce annotations make continuous skeleton extraction from masks challenging. Voxel-based methods often struggle with the sparse topology and geometric directionality inherent in the small bowel skeleton, leading to inefficiency and high memory cost. To address these limitations, we propose a novel octree-based conditional diffusion model (i.e., Tree-Diffusion) that generates anatomically consistent small bowel skeletons guided by 3D segmentation masks. Specifically, we introduce two modules that captures structural priors from masks and topology characteristics from skeletons, ensuring cross-domain alignment and high-quality skeleton generation. Besides, we design a synthesis strategy to generate anatomically plausible skeleton-mask pairs, serving as topological priors to guide the diffusion model toward realis-tic structure predictions. To efficiently represent the elongated skeleton, we adopt an octree- based spatial encoding of hierarchical geometric features. Compared with baselines, our model achieves superior performance in anatomical fidelity, directional consistency, and inference efficiency. The code is available at: https://github.com/Small-Bowel-Skeleton-GenerationlCode
Accurate tissue segmentation of infant brain in magnetic resonance (MR) images is crucial for charting early brain development and identifying biomarkers. Due to ongoing myelination and maturation, in the isointense phase (6-9 months of age), the gray and white matters of infant brain exhibit similar intensity levels in MR images, posing significant challenges for tissue segmentation. Meanwhile, in the adult-like phase around 12 months of age, the MR images show high tissue contrast and can be easily segmented. In this paper, we propose to effectively exploit adult-like phase images to achieve robustmulti-view isointense infant brain segmentation. Specifically, in one way, we transfer adult-like phase images to the isointense view, which have similar tissue contrast as the isointense phase images, and use the transferred images to train an isointense-view segmentation network. On the other way, we transfer isointense phase images to the adult-like view, which have enhanced tissue contrast, for training a segmentation network in the adult-like view. The segmentation networks of different views form a multi-path architecture that performs multi-view learning to further boost the segmentation performance. Since anatomy-preserving style transfer is key to the downstream segmentation task, we develop a Disentangled Cycle-consistent Adversarial Network (DCAN) with strong regularization terms to accurately transfer realistic tissue contrast between isointense and adult-like phase images while still maintaining their structural consistency. Experiments on both NDAR and iSeg-2019 datasets demonstrate a significant superior performance of our method over the state-of-the-art methods.
Deep learning has shown great potential to automate abdominal organ segmentation and quantification. However, most existing algorithms rely on expert annotations and do not have comprehensive evaluations in real-world multinational settings. To address these limitations, we organised the FLARE 2022 challenge to benchmark fast, low- resource, and accurate abdominal organ segmentation algorithms. We first constructed an intercontinental abdomen CT dataset from more than 50 clinical research groups. We then independently validated that deep learning algorithms achieved a median dice similarity coefficient (DSC) of 900% (IQR 874-913%) by use of 50 labelled images and 2000 unlabelled images, which can substantially reduce manual annotation costs. The best-performing algorithms successfully generalised to holdout external validation sets, achieving a median DSC of 894% (852-913%), 900% (843-930%), and 885% (809-919%) on North American, European, and Asian cohorts, respectively. These algorithms show the potential to use unlabelled data to boost performance and alleviate annotation shortages for modern artificial intelligence models.
Survival analysis is paramount for cancer patients as it offers crucial prognostic insights for treatment planning. The performance of existing survival analysis methods is mainly limited by two factors: 1) inefficient extraction of features from multi-modal medical data, e.g., MR images and clinical diagnostic descriptions; and 2) inadequate focus on disease-relevant regions, e.g., primary tumor. To deal with these challenges, in this study, we propose a rectal cancer survival analysis model, dubbed as SurRecNet, which effectively fuse MR images and diagnostic descriptions and takes advantage of multi-task learning. Specifically, we introduce a cross-modality alignment module, aiming to precisely align diagnostic descriptions with MR images at a granular level and facilitate accurate survival analysis. Furthermore, SurRecNet simultaneously predicts tumor masks, relapse states, and survival outcomes by leveraging multi-task learning strategy, imitating the diagnostic process of radiologists to enhance prediction performance. Experimental results on a real clinical rectal multi-modal dataset demonstrate that our SurRecNet significantly outperforms the state-of-the-art methods.
Neural networks have found widespread application in medical image registration, although they typically assume access to the entire training dataset during training. In clinical scenarios, medical images of various anatomical targets, such as the heart, brain, and liver, may be obtained successively with advancements in imaging technologies and diagnostic procedures. The accuracy of registration on a new target may degrade over time, as the registration models become outdated due to domain shifts occurring at unpredictable intervals. In this study, we introduce a deep registration model based on continual learning to mitigate the issue of catastrophic forgetting during training with continuous data streams. To enable continuous network training, we propose a dynamic memory system based on a density-based clustering algorithm to retain representative samples from the data stream. Training the registration network on these representative samples enhances its generalization capabilities to accommodate new targets within the data stream. We evaluated our approach using the CHAOS dataset, which comprises multiple targets, such as the liver, left kidney, and spleen, to simulate a data stream. The experimental findings illustrate that the proposed continual registration network achieves comparable performance to a model trained with full data visibility.
PET-CT integrates metabolic information with anatomical structures and plays a vital role in revealing systemic metabolic abnormalities. Automatic segmentation of lesions from whole-body PET-CT could assist diagnostic workflow, support quantitative diagnosis, and increase the detection rate of microscopic lesions. However, automatic lesion segmentation from PET-CT images still faces challenges due to 1) limitations of single-modality-based annotations in public PET-CT datasets, 2) difficulty in distinguishing between pathological and physiological high metabolism, and 3) lack of effective utilization of CT's structural information. To address these challenges, we propose a threefold strategy. First, we develop an in-house dataset with dual-modalitybased annotations to improve clinical applicability; Second, we introduce a model called Latent Mamba U-Net (LM-UNet), to more accurately identify lesions by modeling long-range dependencies; Third, we employ an anatomical enhancement module to better integrate tissue structural features. Experimental results show that our comprehensive framework achieves improved performance over the state-of-the-art methods on both public and in-house datasets, further advancing the development of AI-assisted clinical applications. Our code is available at https://github.com/Joey-S-Liu/LM-UNet.
Automatic estimation of local vascular measurements, such as center-ness, radius, orientation, and vascular mask, could provide quantitative analysis of vascular diseases and support surgical procedures in clinical applications. However, developing a single model for vascular image computing poses additional challenges because vessels are widely spread but have thin and network-like structures. Traditionally, tubularity has been employed as the primary assumption to enhance the vascular features from images. Existing deep-learning-based approaches incorporate the assumptions of tubular structures in an implicit manner. This, however, imposes limitations on the potential for extensive exploration in the realm of effective feature extraction. In this paper, we propose a threefold strategy. First, we design a combined computing framework to estimate various local measurements of vessels. Second, we propose a tubular shape-guided convolution, i.e., orientational-deformable convolution, where the cuboid grids are rotated and deformed according to vascular orientations to capture the vascular features efficiently. Third, we deploy a tubular sampling strategy within the network to pinpoint the vascular center-ness accurately. Experimental results conducted on two vascular datasets containing coronary and cerebral vessels demonstrate that the method exhibits superior accuracy, particularly in identifying middle and distal vessels.
Medical images are generally acquired with limited field-of-view (FOV), which could lead to incomplete regions of interest (ROI), and thus impose a great challenge on medical image analysis. This is particularly evident for the learning-based multi-target landmark detection, where algorithms could be misleading to learn primarily the variation of background due to the varying FOV, failing the detection of targets. Based on learning a navigation policy, instead of predicting targets directly, reinforcement learning (RL)-based methods have the potential to tackle this challenge in an efficient manner. Inspired by this, in this work we propose a multi-agent RL framework for simultaneous multi-target landmark detection. This framework is aimed to learn from incomplete or (and) complete images to form an implicit knowledge of global structure, which is consolidated during the training stage for the detection of targets from either complete or incomplete test images. To further explicitly exploit the global structural information from incomplete images, we propose to embed a shape model into the RL process. With this prior knowledge, the proposed RL model can not only localize dozens of targets simultaneously, but also work effectively and robustly in the presence of incomplete images. We validated the applicability and efficacy of the proposed method on various multi-target detection tasks with incomplete images from practical clinics, using body dual-energy X-ray absorptiometry (DXA), cardiac MRI and head CT datasets. Results showed that our method could predict whole set of landmarks with incomplete training images up to 80% missing proportion (average distance error 2.29 cm on body DXA), and could detect unseen landmarks in regions with missing image information outside FOV of target images (average distance error 6.84 mm on 3D half-head CT). Our code will be released via https://zmiclab.github.io/projects.html.
Numerous unlabeled data is useful for supervised medical image segmentation, if the labeled data is limited. To leverage all the unlabeled images for efficient abdominal organ segmentation, we developed semi-supervised framework with cross supervision using siamese network, i.e., SemiSeg-CSSN. Cross supervision enables the two networks to optimize the network using pseudo-labels generated by the other. Moreover, we applied the cascade strategy for the task because of the large and uncertain locations of the abdomen regions. To validate the effects of unlabeled data, we employed an unlabeled image filtering strategy to select the unlabeled image and their pseudo label images with low uncertainty. On the FLARE2022 validation cases, with the help of unlabeled data, our method obtained the average dice similarity coefficient (DSC) of 77.7% and average normalized surface distance (NSD) of 82.0%, which is better than the supervised method. The average running time is 12.9 s per case in inference phase and maximum used GPU memory is 2052 MB.
Quantitative organ assessment is an essential step in automated abdominal disease diagnosis and treatment planning. Artificial intelligence (AI) has shown great potential to automatize this process. However, most existing AI algorithms rely on many expert annotations and lack a comprehensive evaluation of accuracy and efficiency in real-world multinational settings. To overcome these limitations, we organized the FLARE 2022 Challenge, the largest abdominal organ analysis challenge to date, to benchmark fast, low-resource, accurate, annotation-efficient, and generalized AI algorithms. We constructed an intercontinental and multinational dataset from more than 50 medical groups, including Computed Tomography (CT) scans with different races, diseases, phases, and manufacturers. We independently validated that a set of AI algorithms achieved a median Dice Similarity Coefficient (DSC) of 90.0\% by using 50 labeled scans and 2000 unlabeled scans, which can significantly reduce annotation requirements. The best-performing algorithms successfully generalized to holdout external validation sets, achieving a median DSC of 89.5\%, 90.9\%, and 88.3\% on North American, European, and Asian cohorts, respectively. They also enabled automatic extraction of key organ biology features, which was labor-intensive with traditional manual measurements. This opens the potential to use unlabeled data to boost performance and alleviate annotation shortages for modern AI models.
Significant breakthroughs in medical image registration have been achieved using deep neural networks (DNNs). However, DNN-based end-to-end registration methods often require large quantities of data or adequate annotations for training. To leverage the intensity information of abundant unlabeled images, unsupervised registration methods commonly employ intensity-based similarity measures to optimize the network parameters. However, finding a sufficiently robust measure can be challenging for specific registration applications. Weakly supervised registration methods use anatomical labels to estimate the deformation between images. High-level structural information in label images is more reliable and practical for estimating the voxel correspondence of anatomic regions of interest between images, whereas label images are extremely difficult to collect. In this paper, we propose a two-stage semi-supervised learning framework for medical image registration, which consists of unsupervised and weakly supervised registration networks. The proposed semi-supervised learning framework is trained with intensity information from available images, label information from a relatively small number of labeled images and pseudo-label information from unlabeled images. Experimental results on two datasets (cardiac and abdominal images) demonstrate the efficacy and efficiency of this method in intra- and inter-modality medical image registrations, as well as its superior performance when a vast amount of unlabeled data and a small set of annotations are available. Our code is publicly available at https://github.com/jdq818/SeRN .
Developing efficient vessel-tracking algorithms is crucial for imaging-based diagnosis and treatment of vascular diseases. Vessel tracking aims to solve recognition problems such as key (seed) point detection, centerline extraction, and vascular segmentation. Extensive image-processing techniques have been developed to overcome the problems of vessel tracking that are mainly attributed to the complex morphologies of vessels and image characteristics of angiography. This paper presents a literature review on vessel-tracking methods, focusing on machine-learning-based methods. First, the conventional machine-learning-based algorithms are reviewed, and then, a general survey of deep-learning-based frameworks is provided. On the basis of the reviewed methods, the evaluation issues are introduced. The paper is concluded with discussions about the remaining exigencies and future research.
Registration networks have shown great application potentials in medical image analysis. However, supervised training methods have a great demand for large and high-quality labeled datasets, which is time-consuming and sometimes impractical due to data sharing issues. Unsupervised image registration algorithms commonly employ intensity-based similarity measures as loss functions without any manual annotations. These methods estimate the parameterized transformations between pairs of moving and fixed images through the optimization of the network parameters during training. However, these methods become less effective when the image quality varies, e.g., some images are corrupted by substantial noise or artifacts. In this work, we propose a novel approach based on a low-rank representation, i.e., Regnet-LRR, to tackle the problem. We project noisy images into a noise-free low-rank space, and then compute the similarity between the images. Based on the low-rank similarity measure, we train the registration network to predict the dense deformation fields of noisy image pairs. We highlight that the low-rank projection is reformulated in a way that the registration network can successfully update gradients. With two tasks, i.e., cardiac and abdominal intra-modality registration, we demonstrate that the low-rank representation can boost the generalization ability and robustness of models as well as bring significant improvements in noisy data registration scenarios.
BackgroundComputed tomography angiography (CTA) is a non-invasive technique to image coronary arteries and evaluate coronary artery diseases (CAD). The diagnosis of CAD requires modeling anatomical structures and analyzing the function and pathology of the coronary arteries. Therefore, a robust and automated method for extracting reliable coronary artery centerlines is valuable in clinical practice.MethodWe extracted coronary centerlines using the directional fast marching (DFM) method and improved DFM with a multi-model strategy. The method comprises model guidance, the application of vessel direction, and a multi-model strategy: (1) coronary models are constructed using registration techniques and then used as prior knowledge of the vessels; (2) the vessel direction, modified from the eigenvectors of the Hessian matrix and vesselness, is used to guide the search for the vessel points during fast marching; and (3) the multi-model strategy is applied to identify suboptimal results from the overall outcome as in multi-atlas segmentation. Overlap and accuracy metrics are used to assess the segmentation. The authors evaluated the performance of the proposed method on 32 CT cardiac angiography datasets from the Rotterdam Coronary Artery Algorithm Evaluation Framework (RCAAEF). The authors also studied the effect of models on DFM.ResultsFor the quantitative evaluation, DFM improved the average overlap (OV) from 43.6% of a method without model information to 77.8%. In addition, with the ground truth delineated by experts, multi-model DFM (MM-DFM) obtained 83.5% average overlap (OV) in the training datasets and 86.6% in the test datasets.ConclusionThe authors propose a novel approach to extract coronary centerlines from CTA using DFM and further extend DFM to a multi-model strategy. DFM effectively applies the prior shape of the coronary vessels and vascular features within the target image and has the potential to achieve clinically relevant results.
Extracting centerlines of coronary arteries is challenging but important in clinical applications of cardiac computed tomography angiography (CTA). Since manual annotation of coronary arteries is time-consuming, labor-intensive and subject to intra- and inter-variations, we propose a new method to fully automatically extract the coronary centerlines. We first develop a new image filter which generates pixels with salient vessel features within a given window. This filter hence can capture sparsely distributed but important vessel points, enabling the minimal path (MP) process to track the key centerline points at different resolution of the images. Then, we reformulate the filter for multi-resolution fast marching, which not only can speed up the coronary tracking process, but also can help the front propagation to step over the indistinct segments of the coronary artery such as at the locations of stenosis. We embed this scheme into the MP framework to develop a multi-resolution multi-model approach (MMP), where the extracted centerlines from low-resolution MP serve as prior and constraints for the high-resolution process. We evaluated the performance of this method using the Rotterdam CTA training data and the coronary artery algorithm evaluation framework. The average inside of our extraction was 0.51 mm and the overlap was 72.9