To develop and validate APEX-NET for early diagnosis and severity stratification of acute pancreatitis (AP) using non-contrast CT (NCCT), by leveraging contrast-enhanced CT (CECT) feature learning. This five-center retrospective and prospective study included 3383 patients, comprising AP and Non-AP (abdominal pain patients and healthy individuals) patients. APEX-NET was trained and evaluated to perform pancreas segmentation, AP diagnosis (AP vs Non-AP), and severity prediction (mild, moderately severe, or severe per the revised Atlanta classification) using 3 internal and 2 external cohorts. A feature mapping module was employed to derive simulated CECT features from NCCT based on paired NCCT-CECT feature learning. The model was further evaluated with subgroup analyses, and a reader study was conducted by comparing its performance with six radiologists of varying experiences. Evaluation metrics included the Dice similarity coefficient, area under the receiver operating characteristic curve (AUC), and accuracy. For AP diagnosis, APEX-NET achieved AUCs of 0.949, 0.958, 0.981, and 0.955 in the validation, internal, and two external testing cohorts, respectively. For severity prediction, APEX-NET significantly outperformed the NCCT model (p < 0.05), with macro-average AUCs of 0.873 (validation) and 0.872 (internal testing). The advantage of APEX-NET had been demonstrated in almost all the age, gender, and etiology subgroups. In the reader study, APEX-NET performed comparably to senior radiologists and superior to junior radiologists (p < 0.05). APEX-NET enables accurate NCCT-based diagnosis and early severity stratification of AP, demonstrating strong potential for clinical integration to overcome the inherent delay of CECT-based assessment. Question The absence of an accurate method for predicting AP severity from early NCCT, the initial diagnostic scan, thus forgoing the critical intervention window. Findings Achieve accurate severity prediction for AP by incorporating contrast-enhanced feature learning. Demonstrate robust performance across diverse demographic groups, etiologies, and imaging parameters. Clinical relevance The APEX-NET, an integrated deep learning framework using NCCT, accelerated the diagnosis and severity stratification of AP, demonstrating performance comparable to senior radiologists and direct potential for clinical workflow integration by reducing reliance on delayed contrast-enhanced scans.
Deformable image registration remains a fundamental task in clinical practice, yet solving registration problems involving complex deformations remains challenging. Current deep learning-based registration methods employ continuous deformation to model large deformations, which often suffer from accumulated registration errors and interpolation inaccuracies. We propose a novel field refinement framework (FiRework) to enhance the efficiency and accuracy of deformable image registration. Unlike existing continuous deformation frameworks, FiRework treats the network as a deformable field refiner, directly incorporating the moving image and the previous deformation field as additional inputs. This allows the network to reassess and optimize the previous deformation, thereby mitigating error accumulation and interpolation issues. Notably, our FiRework requires only one level of recursion during training and supports continuous inference, offering improved accuracy compared to continuous deformation frameworks. The efficacy of our FiRework is evaluated using two brain MRI datasets, LPBA and Mindboggle. For the Mindboggle dataset, FiRework-enhanced network achieves a Dice similarity coefficient (DSC) of 62.8% and an average symmetric surface distance (ASSD) of 1.30 mm, while for the LPBA dataset, it attains a DSC of 70.5% and an ASSD of 1.71 mm. These results show that FiRework consistently outperforms other methods across both datasets, demonstrating favorable improvements in registration accuracy. The code is publicly available at https://github.com/ZAX130/FiRework.
Accurate segmentation of ultrasound images plays a critical role in disease screening and diagnosis. Recently, neural network-based methods have garnered significant attention for their potential in improving ultrasound image segmentation. However, these methods still face significant challenges, primarily due to inherent issues in ultrasound images, such as low resolution, speckle noise, and artifacts. Additionally, ultrasound image segmentation encompasses a wide range of scenarios, including organ segmentation (e.g., cardiac and fetal head) and lesion segmentation (e.g., breast cancer and thyroid nodules), making the task highly diverse and complex. Existing methods are often designed for specific segmentation scenarios, which limits their flexibility and ability to meet the diverse needs across various scenarios. To address these challenges, we propose a novel Localized and Globalized Frequency Fusion Model (LGFFM) for ultrasound image segmentation. Specifically, we first design a Parallel Bi-Encoder (PBE) architecture that integrates Local Feature Blocks (LFB) and Global Feature Blocks (GLB) to enhance feature extraction. Additionally, we introduce a Frequency Domain Mapping Module (FDMM) to capture texture information, particularly high-frequency details such as edges. Finally, a Multi-Domain Fusion (MDF) method is developed to effectively integrate features across different domains. We conduct extensive experiments on eight representative public ultrasound datasets across four different types. The results demonstrate that LGFFM outperforms current state-of-the-art methods in both segmentation accuracy and generalization performance.
Accurate breast tumor segmentation in dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is vital for diagnosis and treatment planning. Despite advances in deep learning, its performance remains constrained by the need for extensive voxel-wise annotations. To mitigate this burden, we propose an annotation-efficient framework that jointly optimizes data selection, unlabeled data utilization, and data augmentation under limited annotation budgets. A diversity-aware uncertainty query (DUQ) strategy guides the annotation by jointly modeling data representativeness and informativeness through a representative candidate selector (RCS) and an uncertainty-based decision maker (UDM), ensuring efficient and targeted labeling. To leverage unlabeled data, a cross-decoder consistency regularization (CDCR) mechanism enforces prediction consistency between two decoders with distinct attention mechanisms, enhancing robustness and confidence. Furthermore, a lesion transplant augmentation (LTA) technique synthesizes anatomically valid pseudo samples by transplanting lesion regions from labeled to unlabeled images, effectively expanding training diversity. Experiments were conducted on two DCE-MRI datasets with biopsy-proven breast cancers, one as internal dataset containing 676 subjects and the other as external dataset with 344 subjects. Comparative and ablation results demonstrate that our framework consistently outperforms state-of-the-art semi-supervised and active learning methods, providing a simple yet effective annotation-efficient solution for breast cancer segmentation in DCE-MRI. The code is publicly available at https: //github.com/zouquanling/DUQ_and_CDCR.
The segmentation of pelvic fracture fragments in CT and X-ray images is crucial for trauma diagnosis, surgical planning, and intraoperative guidance. However, accurately and efficiently delineating the bone fragments remains a significant challenge due to complex anatomy and imaging limitations. The PENGWIN challenge, organized as a MICCAI 2024 satellite event, aimed to advance automated fracture segmentation by benchmarking state-of-the-art algorithms on these complex tasks. A diverse dataset of 150 CT scans was collected from multiple clinical centers, and a large set of simulated X-ray images was generated using the DeepDRR method. Final submissions from 16 teams worldwide were evaluated under a rigorous multi-metric testing scheme. The top-performing CT algorithm achieved an average fragment-wise intersection over union (IoU) of 0.930, demonstrating satisfactory accuracy. However, in the X-ray task, the best algorithm achieved an IoU of 0.774, which is promising but not yet sufficient for intra-operative decision-making, reflecting the inherent challenges of fragment overlap in projection imaging. Beyond the quantitative evaluation, the challenge revealed methodological diversity in algorithm design. Variations in instance representation, such as primary-secondary classification versus boundary-core separation, led to differing segmentation strategies. Despite promising results, the challenge also exposed inherent uncertainties in fragment definition, particularly in cases of incomplete fractures. These findings suggest that interactive segmentation approaches, integrating human decision-making with task-relevant information, may be essential for improving model reliability and clinical applicability.
Medical image challenges have played a transformative role in advancing the field, catalyzing innovation and establishing new performance benchmarks. Image registration, a foundational task in neuroimaging, has similarly advanced through the Learn2Reg initiative. Building on this, we introduce the Large-scale Unsupervised Brain MRI Image Registration (LUMIR) challenge, a next-generation benchmark for unsupervised brain MRI registration. Previous challenges relied upon anatomical label maps, however LUMIR provides 4,014 unlabeled T1-weighted MRIs for training, encouraging biologically plausible deformation modeling through self-supervision. Evaluation includes 590 in-domain test subjects and extensive zero-shot tasks across disease populations, imaging protocols, and species. Deep learning methods consistently achieved state-of-the-art performance and produced anatomically plausible, diffeomorphic deformation fields. They outperformed several leading optimization-based methods and remained robust to most domain shifts. These findings highlight the growing maturity of deep learning in neuroimaging registration and its potential to serve as a foundation model for general-purpose medical image registration.
Accurate automated segmentation of bone fractures from computed tomography (CT) requires large amounts of annotated data to train deep learning models. However, obtaining such annotations presents unique challenges, as the process demands expert knowledge to identify diverse fracture patterns, assess severity, and account for individual anatomical variations. This makes the annotation process highly time-consuming and expensive. Although semi-supervised learning methods can utilize unlabeled data, existing approaches often struggle with the complexity and variability of fracture morphologies, as well as limited generalizability across datasets. To address these challenges, we propose an effective training strategy based on masked autoencoder (MAE) for accurate bone fracture segmentation in CT. The proposed method combines MAE-based anatomical structure learning from unlabeled data with cross-scale cascaded attention (CSCA)-enhanced fine-tuning that focuses on fracture-specific detail preservation. The method is evaluated on two CT datasets: 180 tibial plateau fractures (TPF) and 103 hip fractures (HF). It outperforms representative semi-supervised baselines, achieving high segmentation accuracy with only 20 annotated cases and demonstrating strong cross-site generalization. These results demonstrate that the proposed training strategy effectively improves the efficiency and accuracy of bone fracture segmentation while maintaining strong generalizability across anatomically distinct datasets. This makes the method well-suited for real-world clinical deployment, particularly in data-scarce or resource-constrained environments. The code is publicly available at https://github.com/yuepeiyan/ GeneralizableFractureSeg.
Text-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation results initial deviation between frames; discrepancy in semantic feature tokens between denoising network blocks gradually accumulates as the frame count grows, leading to greater deviations; attention mechanisms struggle to capture global relationships across distant frames in long videos. To address these, we propose FreeMem, a tuning-free framework leveraging hierarchical memory update and injection: the noise memory stabilizes consistency by manipulating low and high frequency components in the initial noise space; the token memory combats inconsistency through adaptive fusion of historical and current semantic feature tokens between denoising network blocks; and the attention memory establishes persistent cache to model long-range relationships within self attention layers. Evaluated on VBench, FreeMem improves subject and background consistency matrics across various methods, offering a practical solution for low-cost, high-consistency long video generation.
Visual context is essential for point cloud semantic segmentation. The contextual information captures the semantic relationship between 3-D points, providing helpful hints for reasoning the category labels of points. Most current methods harness the internal context from the parts of the same object (or from the things within the same scene). In contrast, we propose the external point-set context (EPSC), allowing a richer context of external points learned across various objects/scenes to assist the segmentation task. We employ an external memory with multiple sets to store the EPSC representations learned from the training data. Each representation is a cluster feature, which captures the relationship between adjacent 3-D points for recognizing the semantic category of the center point. During the inference phase, the external memory releases the EPSC representations, providing rich and relevant context for segmenting the target point cloud. We extensively evaluate our method on Stanford Large-Scale 3-D Indoor Spaces (S3DIS), ScanNetv2, and ShapeNetPart datasets, where we achieve the result of effective improvement.
The development of robust deep learning models for breast ultrasound (BUS) image analysis is significantly constrained by the scarcity of expert-annotated data. To address this limitation, we propose a clinically controllable generative framework for synthesizing BUS images. This framework integrates clinical descriptions with structural masks to generate tumors, enabling fine-grained control over tumor characteristics such as morphology, echogencity, and shape. Furthermore, we design a semantic-curvature mask generator, which synthesizes structurally diverse tumor masks guided by clinical priors. During inference, synthetic tumor masks serve as input to the generative framework, producing highly personalized synthetic BUS images with tumors that reflect real-world morphological diversity. Quantitative evaluations on six public BUS datasets demonstrate the significant clinical utility of our synthetic images, showing their effectiveness in enhancing downstream breast cancer diagnosis tasks. Furthermore, visual Turing tests conducted by experienced sonographers confirm the realism of the generated images, indicating the framework's potential to support broader clinical applications.
We introduce MM-Mixing, a multi-modal mixing alignment framework for 3D understanding. MM-Mixing applies mixing-based methods to multi-modal data, preserving and optimizing cross-modal connections while enhancing diversity and improving alignment across modalities. Our proposed two-stage training pipeline combines feature-level and input-level mixing to optimize the 3D encoder. The first stage employs feature-level mixing with contrastive learning to align 3D features with their corresponding modalities. The second stage incorporates both feature-level and input-level mixing, introducing mixed point cloud inputs to further refine 3D feature representations. MM-Mixing enhances intermodality relationships, promotes generalization, and ensures feature consistency while providing diverse and realistic training samples. We demonstrate that MM-Mixing significantly improves baseline performance across various learning scenarios, including zero-shot 3D classification, linear probing 3D classification, and cross-modal 3D shape retrieval. Notably, we improved the zero-shot classification accuracy on ScanObjectNN from 51.3% to 61.9%, and on Objaverse-LVIS from 46.8% to 51.4%. Our findings highlight the potential of multi-modal mixing-based alignment to significantly advance 3D object recognition and understanding while remaining straightforward to implement and integrate into existing frameworks.
Deformable image registration (DIR) plays a pivotal role in medical imaging, which is essential for disease diagnosis, treatment planning, and longitudinal studies. Despite the success of pyramid registration networks in modeling coarse-to-fine deformations, current state-of-the-art (SOTA) methods face critical limitations. First, while transformers excel in capturing long-range feature similarities, their complex windowing mechanisms violate the equivariance assumption required for pyramid encoders. Second, motion decomposition strategies ignore image feature interactions, which restricts their ability to achieve reasonable deformations. To address these challenges, we propose a Neighborhood Attention-based Pyramid (NAP) registration network. Our framework introduces three key innovations: (1) A Convolution and Neighborhood Attention (CoNa) hybrid module integrates convolutional and neighborhood attention for equivariance and discriminative representation. (2) A Neighborhood Motion Enhanced Flow Estimator (NMFE) decoder synergizes cross-attention motion patterns and residual features, enabling hierarchical refinement of deformation fields by leveraging both local texture and motion cues. (3) A Test Time Recursion (TTR) strategy further enhances the registration performance during inference. Extensive experiments on brain and cardiac magnetic resonance imaging (MRI) datasets substantiate the efficacy of our proposed approach, achieving SOTA performance in all these tasks, effectively addresses complex deformable registration challenges in clinical practice. The code is publicly available at https://github.com/ZAX130/NAP.
Echocardiography, a vital cardiac imaging modality, faces challenges due to limited annotated data, impeding the application of deep learning. This paper introduces EchoCardMAE, a customized masked video autoencoder framework designed to leverage unlabeled echocardiography data and enhance performance across diverse cardiac tasks. EchoCardMAE addresses key challenges in echocardiogram analysis through three innovations built upon masked video modeling (MVM): (1) Key Area Masking, which concentrates feature learning on the diagnostically relevant sector of the image; (2) Temporal-Invariant Alignment Loss, promoting feature consistency across different clips of the same echocardiogram; and (3) Reconstruction Denoising, improving robustness to speckle noise inherent in echocardiography. We comprehensively evaluated EchoCardMAE on three public datasets, demonstrating stateof-the-art results in ejection fraction (EF) estimation, Myocardial infarction (MI) prediction, and cardiac segmentation. For example, on the EchoNet-Dynamic dataset, EchoCardMAE achieved an EF estimation MAE of 3.78 and a left ventricular segmentation mDice of 92.96, surpassing existing methods. The code is available at https://github.com/ m1dsolo/EchoCardMAE.
Computer-aided diagnosis (CAD) has become an essential solution for breast ultrasound (BUS) image analysis; however, the development of CAD systems is hindered by high-quality data scarcity and annotation challenges. We propose a novel clinical prior-guided tumor generation method that allows precise control over tumor characteristics, such as size, shape, and texture, using clinical knowledge from textual descriptions and structural masks. Additionally, our method enables cross-domain data generation, enhancing the adaptability of the synthetic data across different imaging conditions. Experiments on three public BUS datasets demonstrate the favorable generation quality and effective cross-domain adaptation of our method. Moreover, the improved accuracy in downstream classification and segmentation tasks further show the clinical utility and practical effectiveness of our synthetic images in supporting breast cancer diagnosis. The code is available at https:// github.com/Violetphy/Clinical-Prior- Tumor-Generation.
Accurate segmentation of cancerous regions in breast dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is crucial for the diagnosis and prognosis assessment of high-risk breast cancer. Deep learning methods have achieved success in this task. However, their performance heavily relies on large-scale fully annotated training data, which are time-consuming and labor-intensive to acquire. To alleviate the annotation effort, we propose a simple yet effective bounding box supervised segmentation framework, which consists of a primary network and an ancillary network. To fully exploit the bounding box annotations, we initially train the ancillary network. Specifically, we integrate abounding box encoder into the ancillary network to serve as a naive spatial attention mechanism, thereby enhancing feature distinction between voxels inside and outside the bounding box. Additionally, we convert uncertain voxel-wise labels inside bounding box into accurate projection labels, ensuring a noise-free initial training process. Subsequently, we adopt an alternating optimization scheme where self-training is performed to generate voxel-wise pseudo labels, and a regularized loss is optimized to correct potential prediction error. Finally, we employ knowledge distillation to guide the training of the primary network with the pseudo labels generated by the ancillary network. We evaluate our method on an in-house DCE-MRI dataset containing 461 patients with 561 biopsy-proven breast cancers (mass/non-mass: 319/242). Our method attains a mean Dice value of 81.42%, outcompeting other weakly- supervised methods in our experiments. Notably, for the non-mass-like lesions with irregular shapes, our method can still generate favorable segmentation with an average Dice of 79.31%. The code is publicly available at https://github.com/Abner228/weakly_box_breast_cancer_seg.
Prostate cancer is a leading cause of cancer-related mortality in men. The registration of magnetic resonance (MR) and transrectal ultrasound (TRUS) can provide guidance for the targeted biopsy of prostate cancer. In this study, we propose a salient region matching framework for fully automated MR-TRUS registration. The framework consists of prostate segmentation, rigid alignment and deformable registration. Prostate segmentation is performed using two segmentation networks on MR and TRUS respectively, and the predicted salient regions are used for the rigid alignment. The rigidly-aligned MR and TRUS images serve as initialization for the deformable registration. The deformable registration network has a dual-stream encoder with cross-modal spatial attention modules to facilitate multi-modality feature learning, and a salient region matching loss to consider both structure and intensity similarity within the prostate region. Experiments on a public MR-TRUS dataset demonstrate that our method achieves satisfactory registration results, outperforming several cutting-edge methods. The code is publicly available at https://github.com/mock1ngbrd/salient-region-matching.
Lychee detection and maturity classification are crucial for yield estimation and harvesting. In densely packed lychee clusters with limited training samples, accurately determining ripeness is challenging. This paper proposes a new transformer model incorporating a Kolmogorov–Arnold Network (KAN), termed GhostResNet (GRN)–KAN–Transformer, for lychee detection and ripeness classification in dense on-tree fruit clusters. First, within the backbone, we introduce a stackable multi-layer GhostResNet module to reduce redundancy in feature extraction and improve efficiency. Next, during feature fusion, we add a large-scale layer to enhance sensitivity to small objects and to increase polling of the small-scale feature map during querying. We further propose a multi-layer cross-fusion attention (MCFA) module to achieve deeper hierarchical feature integration. Finally, in the decoding stage, we employ an improved KAN for the classification and localization heads to strengthen nonlinear mapping, enabling a better fitting to the complex distributions of categories and positions. Experiments on a public dataset demonstrate the effectiveness of GRN-KANformer. Compared with the baseline, GFLOPs and parameters of the model are reduced by 8.84% and 11.24%, respectively, while mean Average Precision (mAP) metrics mAP50 and mAP50–95 reach 94.7% and 58.4%, respectively. Thus, it lowers computational complexity while maintaining high accuracy. Comparative results against popular deep learning models, including YOLOv8n, YOLOv12n, CenterNet, and EfficientNet, further validate the superior performance of GRN-KANformer.
Medical image registration is critical for clinical applications, and fair benchmarking of different methods is essential for monitoring ongoing progress in the field. To date, the Learn2Reg 2020-2023 challenges have released several complementary datasets and established metrics for evaluations. Building on this foundation, the 2024 edition expands the challenge’s scope to cover a wider range of registration scenarios, particularly in terms of modality diversity and task complexity, by introducing three new tasks, including large-scale multi-modal registration and unsupervised inter-subject brain registration, as well as the first microscopy-focused benchmark within Learn2Reg. The new datasets also inspired new method developments, including invertibility constraints, pyramid features, keypoints alignment and instance optimisation.
Deformable image registration plays a crucial role in medical imaging, aiding in disease diagnosis and image-guided interventions. Traditional iterative methods are slow, while deep learning (DL) accelerates solutions but faces usability and precision challenges. This study introduces a pyramid network with the enhanced motion decomposition Transformer (ModeTv2) operator, showcasing superior pairwise optimization (PO) akin to traditional methods. We re-implement ModeT operator with CUDA extensions to enhance its computational efficiency. We further propose RegHead module which refines deformation fields, improves the realism of deformation and reduces parameters. By adopting the PO, the proposed network balances accuracy, efficiency, and generalizability. Extensive experiments on three public brain MRI datasets and one abdominal CT dataset demonstrate the network’s suitability for PO, providing a DL model with enhanced usability. The code is publicly available at https://github.com/ZAX130/ModeTv2.
Prostate cancer (PCa) is a leading cause of cancer-related mortality in men, and accurate identification of clinically significant PCa (csPCa) is critical for timely intervention. Transrectal ultrasound (TRUS) is widely used for prostate biopsy; however, its low contrast and anisotropic spatial resolution pose diagnostic challenges. To address these limitations, we propose a novel hybrid-view attention (HVA) network for csPCa classification in 3D TRUS that leverages complementary information from transverse and sagittal views. Our approach integrates a CNN-transformer hybrid architecture, where convolutional layers extract fine-grained local features and transformer-based HVA models global dependencies. Specifically, the HVA comprises intra-view attention to refine features within a single view and cross-view attention to incorporate complementary information across views. Furthermore, a hybrid-view adaptive fusion module dynamically aggregates features along both channel and spatial dimensions, enhancing the overall representation. Experiments are conducted on an in-house dataset containing 590 subjects who underwent prostate biopsy. Comparative and ablation results prove the efficacy of our method. The code is available at this https URL.