Background:Breast ultrasound (BUS) is widely used for breast cancer (BC) screening and diagnosis, yet accurate breast lesion segmentation remains challenging. Although You Only Look Once (YOLO) and its variants have shown strong performance in object segmentation, their effectiveness on BUS lesion segmentation has not been systematically explored. This study aims to benchmark twelve YOLO variants from four families (YOLOv5, YOLOv8, YOLOv9, and YOLO11) for BUS lesion segmentation under same-database and cross-database settings. Methods:Twelve YOLO variants spanning nano to extra-large scales were fine-tuned and evaluated on two public BUS datasets, the Breast Ultrasound Images (BUSI) dataset (n=647) and the breast ultrasound lesion segmentation dataset from the University of Castilla-La Mancha (BUS-UCLM) dataset (n=264). Each dataset was split into training (80%), validation (10%), and testing (10%) subsets with stratified random partitioning, and experiments were repeated across eight random seeds. Six evaluation metrics, including Dice coefficient, intersection over union (IoU), precision, recall, F1 score (F1S), and mean average precision at IoU threshold 0.5 (mAP@0.5), were used. In addition, U-Net and DeepLabV3+ were compared under the same protocol. Results:Under same-database evaluation, all variants achieved strong performance, with mean Dice scores ≥0.81 on BUSI and ≥0.87 on UCLM. On BUSI, yolov5s achieved the highest mean Dice (0.93±0.012) and IoU (0.88±0.014). On UCLM, yolov8m attained the highest mean Dice (0.98±0.006) and IoU (0.96±0.007). In cross-database evaluation, however, performance degraded substantially, with Dice scores decreasing by approximately 0.20 or more. For BUSI→UCLM, yolov5s achieved the highest mean Dice (0.71±0.032); and for UCLM→BUSI, yolov9c achieved the highest mean Dice (0.60±0.038). Among all variants, yolo11n demonstrated competitive performance across both same- and cross-database evaluations (BUSI Dice 0.91±0.014, and BUSI→UCLM Dice 0.69±0.034; UCLM Dice 0.96±0.009, and UCLM→BUSI 0.59±0.039) while maintaining low computational cost (training time <290 s; inference latency ~64 ms). Notably, yolo11n substantially outperformed U-Net and DeepLabV3+ under both same- and cross-database settings. Conclusions:Fine-tuned YOLO variants achieve strong same-database performance for BUS lesion segmentation, with yolo11n offering the most favorable balance among segmentation accuracy, cross-database competitiveness, and computational efficiency. However, the substantial performance degradation in cross-database settings highlights the critical need for improving domain generalization performance in future work.
Background and Objective:The scarcity of high-quality annotated data is a primary bottleneck in developing artificial intelligence (AI) for chest X-ray (CXR) interpretation. Foundation models (FMs), trained on broad datasets via self-supervision, present a transformative solution. This narrative review analyzes the current state of vision and vision-language foundation models (VLFMs) specifically for CXR, evaluating their potential to bridge research and clinical practice through a novel analytical framework. Methods:Following PRISMA 2020 guidelines, we conducted a narrative review of literature from 2021-2025. We introduced and applied a novel four-pillar analytical taxonomy to structure the field: (I) core model architecture; (II) training framework; (III) clinical application adaptation; and (IV) quality assurance (QA). This framework guided the synthesis of evidence from over 150 studies, enabling a structured analysis of model designs, training paradigms, adaptation strategies, and evaluation metrics. Key Content and Findings:The four-pillar taxonomy provides a critical lens to deconstruct the FM development pipeline. Our analysis reveals distinct maturity levels across pillars: while architectural innovation and training strategies are advanced, rigorous QA and adaptation for real-world heterogeneity (e.g., portable vs. stationary imaging, device/site shifts) are underdeveloped. We identify domain-specific challenges, including bias propagation from limited public datasets, label noise in report-mined supervision, and a significant gap between retrospective benchmark performance and prospective clinical utility. Evidence-based recommendations are provided for model selection, efficient adaptation via parameter-efficient fine-tuning (PEFT), and the implementation of stratified evaluation. Conclusions:CXR-FMs offer a path toward more generalizable and data-efficient diagnostic AI. However, successful clinical translation is contingent upon coordinated advances across all four pillars of the proposed framework, moving beyond architectural scaling to prioritize data diversity, robustness auditing, and seamless workflow integration. This taxonomy provides researchers and clinicians with an actionable framework to develop and evaluate models whose impact will be determined by demonstrated fairness, reliability, and utility in diverse clinical environments.
Medical image segmentation plays a crucial role in computer-aided diagnosis and treatment planning. Unsupervised segmentation methods that can effectively leverage unlabeled data bring significant promise in clinical application. However, they remain a challenging task in maintaining anatomical structure topological consistency that often produces anatomical structure breaks, connectivity errors, or boundary discontinuities. To address these issues, we propose a novel Unsupervised Topological-Aware Diffusion Condensation Network (UTADC-Net) for medical image segmentation. Specifically, we design a diffusion condensation-based framework that achieves structural consistency in segmentation results by effectively modeling long-range dependencies between pixels and incorporating topological constraints. First, to effectively fuse local details and global semantic information, we employ a pixel-centric patch embedding module by simultaneously modeling local structural features and inter-region interactions. Second, to enhance the topological consistency of segmentation results, we introduce an adaptive topological constraint mechanism that guides the network to learn anatomically aligned structural representations through pixel-level topological relationships and corresponding loss functions. Extensive experiments conducted on three public medical image datasets demonstrate that our proposed UTADC-Net significantly outperforms existing unsupervised methods in terms of segmentation accuracy and topological structure preservation. Notably, our method demonstrates segmentation results with excellent anatomical structural consistency. These results indicate that our framework provides a novel and practical solution for unsupervised medical image segmentation.
Deformable image registration plays an important role in medical image analysis. Recent deep learning-based methods typically adopt either U-Net or pyramid architectures. However, U-Net-based approaches often struggle with large deformations, while pyramid methods—despite better handling of large deformations—suffer from irreversible detail loss and accumulated interpolation errors due to multi-scale deformation field composition and repeated sampling operations. To address these limitations, we propose WaveTrans, a wavelet pyramid registration framework that decomposes images into wavelet coefficients, fuses multi-scale frequency features through a Swin Transformer backbone, and reconstructs deformation fields via inverse wavelet transform. Experiments on two diverse registration tasks demonstrate that WaveTrans improves registration accuracy over spatial-domain baselines and markedly outperforms prior wavelet-based approaches on large-deformation cases.
Objective.Complicated deformation remains a critical challenge in medical image registration (MIR). Although deep learning-based spatial domain registration networks have achieved improved accuracy and efficiency, they still suffer from irreversible information loss, the 'small objects move fast' problem-where small objects are lost at low resolutions and reappear with large motions at finer resolutions-and accumulated deformation errors propagating through coarse-to-fine architectures.Approach.We propose Wave-Reg, a spatial-frequency registration framework built on a wavelet pyramid architecture. By integrating the discrete wavelet transform (DWT), the encoder extracts spatial-frequency features via DWT-guided ConvNet to minimize detail loss, while the decoder reconstructs displacement vector field (DVF) via inverse DWT-guided Swin Transformer to mitigate the 'small objects move fast' problem. A cross-scale self-correction module based on Heun's predictor-corrector method further refines deformation fields across scales to tackle accumulated errors.Main results.Experiments on three datasets demonstrate substantial gains in registration accuracy for both large deformation and multi-modality registration tasks.Significance.Wave-Reg demonstrates that spatial-frequency feature learning and predictor-corrector refinement offer an effective solution to longstanding challenges in MIR.
Limited-angle cone-beam computed tomography (LA-CBCT) enables rapid imaging and reduced radiation exposure, but its severely incomplete projection data lead to ill-posed reconstructions with prominent artifacts, limiting clinical applicability. Recent advances in 3D Gaussian Splatting (3D-GS) have shown promise for efficient tomographic reconstruction, yet its performance remains highly sensitive to initialization. In this work, we present SPARK (Structurally-Informed Projection-Accelerated Reconstruction), a two-stage framework that introduces a generative, structurally informed initialization for 3D-GS. In the first stage, a geometry-conditioned network directly predicts complete 3D Gaussian parameters from a sparse subset of projections, embedding learned anatomical priors to mitigate artifact propagation. In the second stage, the generated scene is refined through physics-based 3D-GS optimization, yielding high-fidelity reconstructions consistent with measured projections. Extensive experiments on public datasets demonstrate that SPARK substantially improves both image quality and convergence speed, achieving superior PSNR/SSIM in severely limited-angle scenarios compared with analytical, iterative, and deep learning baselines. Moreover, SPARK reconstructions provide enhanced inputs for downstream post-processing networks, further boosting image fidelity. These results suggest that SPARK is a promising prior-informed 3D-GS framework for simulated LA-CBCT reconstruction under limited angular coverage, providing an effective bridge between data-driven anatomical priors and physics-based projection-domain refinement.
Purpose:Mammography screening is less sensitive in dense breasts, where tissue overlap and subtle findings increase perceptual difficulty. We present MammoColor, an end-to-end framework with a Task-Driven Chromatic Encoding (TDCE) module that converts single-channel mammograms into TDCE-encoded views for visual augmentation. Materials and Methods:MammoColor couples a lightweight TDCE module with a BI-RADS triage classifier and was trained end-to-end on VinDr-Mammo. Performance was evaluated on an internal test set, two public datasets (CBIS-DDSM and INBreast), and three external clinical cohorts. We also conducted a multi-reader, multi-case (MRMC) observer study with a washout period, comparing (1) grayscale-only, (2) TDCE-only, and (3) side-by-side grayscale+TDCE. Results:On VinDr-Mammo, MammoColor improved AUC from 0.7669 to 0.8461 (P=0.004). Gains were larger in dense breasts (AUC 0.749 to 0.835). In the MRMC study, TDCE-encoded images improved specificity (0.90 to 0.96; P=0.052) with comparable sensitivity. Conclusion:TDCE provides a task-optimized chromatic representation that may improve perceptual salience and reduce false-positive recalls in mammography triage.
Reconstructing three-dimensional (3D) anatomy from routine X-ray imaging remains a long-standing challenge, promising high accessibility and minimal radiation exposure compared to computed tomography (CT). We propose X2Shape, a deep learning framework that enables direct 3D multi-organ reconstruction from orthogonal biplanar X-rays, without requiring CT priors. X2Shape combines a geometry-aware volumetric backprojection with a cross-view fusion module based on state-space modeling, achieving accurate 2D-to-3D mapping and robust integration of complementary views. To overcome the scarcity of paired training data, we devise a hybrid deformation-based augmentation strategy that generates anatomically diverse, realistic samples, markedly improving model generalization. Across two challenging thoracic benchmarks, X2Shape substantially outperforms existing methods, reaching Dice scores of 88.98% and 75.62% on TotalSegmentator-Subset and LCTSC datasets, respectively. Beyond accuracy, it demonstrates strong cross-dataset generalization, reconstructing diverse organ structures with efficiency and robustness. By eliminating dependence on CT while preserving volumetric fidelity, X2Shape establishes a scalable paradigm for low-cost, low-radiation 3D imaging, with potential to broaden clinical access to personalized diagnostics, surgical planning, and image-guided interventions. The code is available at https://github.com/pzhhhhh2263/X2Shape.
Left ventricular non-compaction (LVNC) is a rare cardiomyopathy with distinctive myocardial morphology. Due to its low prevalence, LVNC cases are rarely included in public cardiac datasets, hindering the development of specialized deep learning segmentation models. This study aims to systematically evaluate the generalization capability of state-of-the-art models, trained on public, multi-disease cardiac datasets, to the challenging LVNC cohort. Beyond performance benchmarking, we investigate the impact of critical data factors, including training set size and pathology composition, to derive actionable insights for building effective models in data-scarce, rare disease scenarios. We benchmarked state-of-the-art segmentation models from four architectural categories (CNN-based, Transformer-based, Mamba-based, and pre-trained foundation models). Trained on the M M dataset, the models were tested on an independent LVNC dataset, with performance evaluated using Dice score, mIoU, ASD, 95
Traditional Chinese acupuncture methods often face controversy in clinical practice due to their high subjectivity. Additionally, current intelligent-assisted acupuncture systems have two major limitations: slow acupoint localization speed and low accuracy. To address these limitations, a new method leverages the excellent inference efficiency of the state-space model Mamba, while retaining the advantages of the attention mechanism in the traditional DETR architecture, to achieve efficient global information integration and provide high-quality feature information for acupoint localization tasks. Furthermore, by employing the concept of residual likelihood estimation, it eliminates the need for complex upsampling processes, thereby accelerating the acupoint localization task. Our method achieved state-of-the-art (SOTA) accuracy on a private dataset of acupoints on the human back, with an average Euclidean distance pixel error (EPE) of 7.792 and an average time consumption of 10.05 milliseconds per localization task. Compared to the second-best algorithm, our method improved both accuracy and speed by approximately 14%. This significant advancement not only enhances the efficacy of acupuncture treatment but also demonstrates the commercial potential of automated acupuncture robot systems. Access to our method is available at https://github.com/Sohyu1/RT-DEMT.
Background and Purpose: Precise delineation of pelvic organs-at-risk (OARs) is crucial for high-dose-rate brachytherapy (HDR-BT) in cervical cancer treatment. While deep learning methods have shown promise in automatic delineation, substantial and complex organ deformations pose significant challenges. This study presents a novel approach to address these issues. Materials and Methods: We introduce a coarse-to-refine strategy for annotation, utilizing limited existing data to expedite the process. Combined with deformation-based data augmentation, we incorporate this information into a three-dimensional attention U-Net (C2FAU-Net). The study included 100 cervical cancer patients, with OARs annotated by experienced oncologists. The dataset was divided into 80 patients for training, 10 for validation, and 10 for testing. To assess the delineation performance, we employed the volumetric dice similarity coefficient (DSC), 95th percentile Hausdorff Distance (HD95), average symmetric surface distance (ASSD), precision, and recall. We compared the time consumed by manual delineation versus artificial intelligence (AI)-assisted delineation. Dosimetric parameters were compared using different contours to evaluate the clinical impact of the automated approach. Results: Our method achieved an average DSC of 89.7%, HD95 of 3.61 mm, and an ASSD of 1.02 mm in the test cohort. The AI-assisted method significantly reduced the manual delineation time from 17.85 ± 3.84 min to 7.54 ± 4.95 min. No significant difference was observed in ΔD2cc, ΔD1cc, ΔD0.1cc, and ΔDmax for bladder, rectum, and sigmoid when comparing contours generated by C2FAU-Net to those created manually. Conclusion: We introduce an effective automatic delineation framework for pelvic OARs, enhancing efficiency within the HDR-BT workflow and potentially improving treatment outcomes.
Unsupervised deformable multimodal medical image registration often confronts complex scenarios, which include intermodality domain gaps, multi-organ anatomical heterogeneity, and physiological motion variability. These factors introduce substantial grayscale distribution discrepancies, hindering precise alignment between different imaging modalities. However, existing methods have not been sufficiently adapted to meet the specific demands of registration in such complex scenarios. To overcome the above challenges, we propose SynMSE, a novel multimodal similarity evaluator that can be seamlessly integrated as a plug-and-play module in any registration framework to serve as the similarity metric. SynMSE is trained using random transformations to simulate spatial misalignments and a structure-constrained generator to model grayscale distribution discrepancies. By emphasizing spatial alignment and mitigating the influence of complex distributional variations, SynMSE effectively addresses the aforementioned issues. Extensive experiments on the Learn2Reg 2022 CT-MR abdomen dataset, the clinical cervical CT-MR dataset, and the CuRIOUS MR-US brain dataset demonstrate that SynMSE achieves state-of-the-art performance. Our code is available on the project page https://github.com/MIXAILAB/SynMSE.
In medical multimodal image registration, images from different modalities exhibit significant differences in intensity distributions. These differences hinder traditional methods from detecting robust feature points and computing accurate descriptors. This study employs modality translation to address these challenges, transforming multimodal registration into a unimodal task and enabling robust feature detection across modalities. Specifically, CycleGAN is utilized to convert CT images to pseudo-MRI, reducing inter-modality discrepancies. The significance of this approach lies in extending feature-pointbased methods, such as R2D2, from unimodal to multimodal scenarios, with experiments showing enhanced feature matching and rigid registration accuracy on converted images.
Obstructive Sleep Apnea and Hypopnea (OSAH) is the most prevalent sleep-related breathing disorder, which has a significant impact on human health. Currently, polysomnography (PSG) devices, as the gold standard for diagnosing this disease, have several drawbacks, including high costs, discomfort during use, and unsuitability for home monitoring. In contrast, radar devices offer a non-contact and convenient monitoring approach, with significant potential for application. However, manual annotation of nighttime radar data is often time-consuming and labor-intensive, which makes the apnea detection task even more challenging. To address this, we propose a deep learning method based on contrastive self-supervised learning for detecting OSAH from radar signals. We evaluated the proposed method using a public dataset with different training-test splits. The experimental results demonstrate that our method outperforms existing state-of-the-art methods, achieving the highest classification accuracy of 98.97%. Notably, even with only 10% labeled data, the method still achieved an accuracy of 86.00%. These results indicate that the proposed method effectively enables monitoring of OSAH from radar signals, while significantly reducing the workload of manual annotation.
Automatic clinical tumor volume (CTV) delineation is pivotal to improving outcomes for interstitial brachytherapy cervical cancer. However, the prominent differences in gray values due to the interstitial needles bring great challenges on deep learning-based segmentation model. In this study, we proposed a novel interstitial-guided segmentation network termed advance reverse guided network (ARGNet) for cervical tumor segmentation with interstitial brachytherapy. Firstly, the location information of interstitial needles was integrated into the deep learning framework via multi-task by a cross-stitch way to share encoder feature learning. Secondly, a spatial reverse attention mechanism is introduced to mitigate the distraction characteristic of needles on tumor segmentation. Furthermore, an uncertainty area module is embedded between the skip connections and the encoder of the tumor segmentation task, which is to enhance the model's capability in discerning ambiguous boundaries between the tumor and the surrounding tissue. Comprehensive experiments were conducted retrospectively on 191 CT scans under multi-course interstitial brachytherapy. The experiment results demonstrated that the characteristics of interstitial needles play a role in enhancing the segmentation, achieving the state-of-the-art performance, which is anticipated to be beneficial in radiotherapy planning.
Cone-Beam Computed Tomography (CBCT) holds significant clinical value in image-guided radiotherapy (IGRT). However, CBCT images of low-density soft tissues are often plagued with artifacts and noise, which can lead to missed diagnoses and misdiagnoses. We propose a new unsupervised CBCT image artifact correction algorithm, named Spatial Convolution Diffusion (ScDiff), based on a conditional diffusion model, which combines the unsupervised learning ability of generative adaptive networks (GAN) with the stable training characteristics of diffusion models. This approach can efficiently and stably achieve CBCT image artifact correction, resulting in clear, realistic CBCT images with complete anatomical structures. The proposed model can effectively improve the image quality of CBCT. The obtained results can reduce artifacts while preserving the anatomical structure of CBCT images. We compared the proposed method with several GAN- and diffusion-based methods. Our method achieved the highest corrected image quality and the best evaluation metrics.
Medical image registration is a critical problem in medical image analysis, enabling the spatial alignment of anatomical structures across different imaging modalities. However, existing algorithms often struggle with local registration of large deformations and exhibit limited feature extraction capabilities. To address these challenges, we propose a multi-level transformation progressive registration algorithm (MTPR). This method incorporates the concept of multi-level transformations, the model performs four progressive registration steps, predicting the deformation field from coarse to fine. Initially, the model applies an enhancement process, introducing a hybrid filtering enhancement module based on wavelet transform and improved guided filtering to enhance image edges. During the registration phase, we propose a pyramid-shared weight enhancement network, which precisely extracts multimodal image features and implements a progressive deformation field prediction strategy. In order to increase the feature extraction capability of the model, we propose spatial feature fusion module in skip connection of encoder and decoder, which combines multi-scale information into a spatial feature representation with rich context information. Additionally, we introduce a dual-similarity metric to enhance model capability for local organ registration by incorporating structural similarity, increasing the model's attention to local organs. Experiments conducted on publicly available datasets (OASIS, LPBA40) and clinical CT/MR data, achieved an average dice similarity coefficient (DSC) of 0.822, average average symmetric surface distance (ASSD) of 0.741 mm, average standard deviation of jacobian determinant (Std.Jacobian) of 0.247 in the clinical CT/MR data. The Wilcoxon signed-rank test statistical analysis shows that the evaluation indicators of the MTPR algorithm are significantly better than other baseline methods (P < 0.05). The MTPR model's multi-scale information aggregation capability effectively handles large deformations, demonstrating excellent registration accuracy and generalization performance.
Breast cancer is a significant global health issue, and the diagnosis of breast cancer through imaging remains challenging. Mammography images are characterized by extremely high resolution, while lesions often occupy only a small portion of the image. Down-sampling in neural networks can easily lead to the loss of microcalcifications or subtle structures. To tackle these challenges, we propose a Context Clustering based triple information fusion framework. First, in comparison to CNNs or transformers, we observe that Context clustering methods are (1) more computationally efficient and (2) better at associating structural or pathological features. This makes them particularly well-suited for mammography in clinical settings. Next, we propose a triple information fusion mechanism that integrates global, feature-based local, and patch-based local information. The proposed approach is rigorously evaluated on two public datasets, Vindr-Mammo and CBIS-DDSM, using five independent data splits to ensure statistical robustness. Our method achieves an AUC of 0.828 ±0.020 on Vindr-Mammo and 0.805 ±0.020 on CBIS-DDSM, outperforming the second best method by 3.5
Respiratory motion artifacts severely compromise the diagnostic utility of Cone-Beam CT (CBCT). Motion-compensated (MoCo) reconstruction faces a critical trade-off: fast deformation vector field (DVF) estimation from prior scans can be inaccurate against anatomical changes, while accurate iterative methods are too slow for clinical workflows. We propose a learning-based MoCo framework to resolve this dilemma. It utilizes a deep neural network, trained patient-specifically on a prior 4D-CT, to rapidly and directly infer DVFs from 2D projections. This non-iterative approach leverages a learned motion model to ensure both efficiency and robustness against inter-fractional changes. These DVFs then guide a reconstruction pipeline that corrects motion on a perprojection basis by warping each back-projection. Experiments on phantom and clinical data demonstrate that this approach effectively corrects motion artifacts, yields superior reconstruction quality. Our framework yields high-quality static 3D images and enables full 4D-CBCT synthesis, paving the way for advanced adaptive radiotherapy workflows.
The deformable registration is a critical task in medical image processing. Due to significant differences in texture patterns and intensity information between modalities, current multimodal registration algorithms fail to extract multimodal features accurately and lack effective similarity measures in complex regions. To address these challenges, we propose a multilevel pyramid large deformation multimodality registration elastic network (MPLD). The framework adopts a global-to-local strategy for registration and is divided into three stages: level0 and level2 stage (global registration stage) and level1 stage (local registration stage). We propose an accurate similarity measurement evaluator to measure the spatial difference between two images, this method combines morphological and deep learning, and then optimize registration by minimizing the errors predicted by the evaluator. In addition, we propose a pyramid multi-level registration network (PM-Net), the module includes two independent encoders to extract image features of different modes, and share the same decoder, using progressive deformation field estimation in the decoder. The proposed method was validated on publicly available datasets LPBA40, OASIS, and hospitals clinical CT/MR data. In clinical data registration, our method achieved an average DSC of 0.816 +/- 0.016, average ASSD of 0.894 +/- 0.128 mm, average Std. Jacobian of 0.289 +/- 0.012. Our algorithm achieved a higher registration accuracy compared with state-of-the-art registration methods. This method adopts a coarse-to-fine strategy for progressive deformation field prediction and leverages multi-scale feature aggregation to enhance feature extraction capability. It effectively handles large deformation registration tasks, and comparative experiments confirm the superior registration performance and generalization ability.