
ABSTRACT Quantitative assessment of neuromuscular disorders relies on accurate musculoskeletal ultrasound (MSKUS) image segmentation. However, speckle noise and ambiguous anatomical boundaries pose continuous challenges. Existing convolutional networks often exhibit limitations in balancing detail preservation with noise suppression. Furthermore, traditional multi‐scale fusion is prone to transmitting low‐level artifacts, causing false positive misclassifications. To address these issues, a robust segmentation network termed CADG‐UNet is proposed. The architecture incorporates two tailored modules to process specific acoustic interference. In the encoding path, the Coordinate Attention and Dynamic Spatial Attention (CADSA) module is designed, combining directional positional encoding with content‐adaptive dynamic convolutions. This design suppresses spatially heterogeneous noise while precisely anchoring muscle boundaries. In the decoding path, the Dynamic Gated Fusion (DGF) module is designed for skip connections, employing a post‐gating local residual refinement mechanism to filter high‐frequency reverberation artifacts from shallow features before multi‐scale fusion. The proposed method is evaluated on the MUST and DeepACSA datasets. CADG‐UNet achieves mean intersection over union (mIoU) scores of 83.31% and 90.97%, respectively. Statistical significance tests indicate stable cross‐dataset performance compared to leading baseline methods. Furthermore, the network yields improved boundary and clinical metrics, including reduced 95th percentile Hausdorff Distance (HD95), lower Average Surface Distance (ASD), and cross‐sectional area (CSA) estimations. It also demonstrates significant robustness under multiplicative speckle noise conditions. These results suggest that CADG‐UNet provides a reliable approach for automated clinical morphological measurements.
ABSTRACT Recent advancements in breast cancer detection using deep CNNs and vision transformers (ViTs) on breast ultrasound images (BUSI) have demonstrated strong performance. However, model complexity and variations in contrast, texture, and morphology continue to limit effectiveness. This study introduces the CB‐Res‐RBCMT, a hybrid framework that combines customized residual CNNs with ViT components for detailed BUSI cancer analysis. The RBCMT integrates stem convolution blocks and CNN‐Meet‐Transformer (CMT) modules, followed by regional‐boundary (RB) feature operations. These operations apply the Laplacian of Gaussian (LoG) filter to enhance homogeneity, reduce speckle noise, and highlight structural transitions, while boundary operations capture malignant morphological changes. The CMT module utilizes multi‐head attention for global context interactions, improving computational efficiency. New inverse residual blocks and stem CNNs further extract malignant texture information and address vanishing gradients. A multiscale channel fusion and attention (MSCFA) block enhances feature representation by combining global context with boundary‐aware cues. The Channel‐Boosted (CB) strategy fuses RBCMT and residual CNN feature maps to increase feature diversity for the limited BUSI dataset. Finally, a spatial attention block refines channels for optimal pixel selection and reduced redundancy. The CB‐Res‐RBCMT achieves 95.63% accuracy, 95.57% F1‐score, 96.42% precision, and 94.79% sensitivity, outperforming existing CNN and ViT‐based approaches. Furthermore, cross‐dataset generalization on mammographic images validates the model's robust clinical applicability across diverse screening scenarios.
ABSTRACT Early and reliable classification of benign and malignant pulmonary nodules remains challenging because nodule appearances are highly heterogeneous and most deep learning models provide limited diagnostic transparency. To address these issues, we propose Cross Fusion‐CapNets, an interpretable model that combines a Multi‐scale Fusion Network (MSFNet) with a capsule‐based explanation module. MSFNet integrates wavelet‐based local features and attention‐based global features through a Multi‐feature Fusion (MFF) Block, while the capsule module learns attribute‐oriented representations associated with eight radiological characteristics from LIDC‐IDRI. Experiments on LIDC‐IDRI achieved 95.4% accuracy and 99.2% AUC for benign‐malignant nodule classification. As a cross‐modality image‐level assessment on the LC25000 dataset, five‐fold stratified cross‐validation and independent partition evaluation yielded 99.96% ± 0.04% validation accuracy and 99.97% ± 0.03% evaluation accuracy. These results show that the proposed model combines accurate classification with attribute‐level interpretability.
ABSTRACT The phonocardiogram (PCG) and electrocardiogram (ECG) are widely used for early detection and prevention of cardiovascular diseases (CVDs) due to their noninvasive acquisition and accurate representation of heart function from different perspectives. However, it is still challenging to extract discriminative characteristics without losing essential information, and few studies have successfully integrated PCG and ECG for CVD screening. Therefore, this research introduces a novel multimodal fused feature‐based CVD classification framework that uniquely integrates advanced preprocessing, modality‐specific feature extraction, and optimized classification. Raw ECG and PCG signals are preprocessed using multispectral adaptive wavelet denoising (MAWD) for noise removal, a U‐Net‐based sequence‐to‐sequence fully convolutional network (U‐TSS) for precise segmentation, and a frequency transformation layer (FTL) for frequency domain conversion. Novelty lies in the tailored feature extraction strategy employing a parallel convolutional neural network (PCNN) for PCG and HCR‐Net for ECG, followed by a Multi‐scale Contextual Feature Fusion (MC2F) based feature fusion mechanism that preserves complementary information across modalities. Classification is performed using the proposed Deep Maxout Fuzzy EfficientNet (DMFE), enabling highly discriminative decision boundaries and robust generalization. Experimental results demonstrate that the proposed ECG–PCG fusion approach surpasses existing methods, achieving 99.18% accuracy, 99.09% sensitivity, and 99.03% specificity. These results highlight the method's strong potential for practical medical deployment, offering high performance, flexibility, and comprehensive characterization of cardiac health.
ABSTRACT Accurate segmentation of prostate tumors on T2‐weighted magnetic resonance imaging (MRI) scans is essential for diagnosis and treatment planning. However, this remains challenging due to the heterogeneous appearance of tumors, their indistinct boundaries, and the frequent occurrence of small lesions. We propose AE‐VMUNet, an attention‐enhanced Vision Mamba U‐Net that performs stage‐wise feature refinement within a state‐space encoder‐decoder framework. The model integrates the following: (i) an Adaptive State‐space Enhanced Feature (ASF) block in the encoder, which couples long‐range dependency modelling with channel‐wise recalibration; (ii) an Attention‐Gate‐based Skip Connection Fusion (SCF) module, which performs spatially selective, multi‐scale aggregation; and (iii) an Efficient State‐space Guided (ESG) block in the decoder, which preserves global context while enabling detail‐aware reconstruction. Experiments on the public PI‐CAI 2022 dataset and an independent, private clinical cohort demonstrate that the AE‐VMUNet model achieves Dice scores of 59.51% and 48.83% respectively, with HD95 values of 14.21 and 18.96 mm. These results outperform those of representative CNN‐, attention‐ and Mamba‐based segmentation baselines. Further evaluations on PROMISE12 and ACDC demonstrate strong cross‐dataset and cross‐anatomy generalization, achieving Dice scores of 92.19% and 91.64%, respectively. Overall, these results suggest that combining state‐space modelling with lightweight attention mechanisms can lead to precise and robust medical image segmentation.
Colon polyp segmentation is important in the early detection of colorectal cancer. Advances in artificial intelligence and computer-aided diagnosis systems have dramatically improved the capabilities of colonoscopy by enhancing the segmentation and detection of colonic polyps. Accurate segmentation is essential, as it allows for the differentiation of polyps from surrounding tissues and aids in assessing their potential malignancy. This paper proposed a method for segmenting colon polyp images by improving the Recalling U-Net model based on the original U-Net and U-Net++ models. The new model aims to take full advantage of multiscale features by introducing full-scale skip connections, combining low-level details with high-level semantics from full-scale feature maps, but with fewer parameters. The model can develop deep supervision to learn hierarchical representations from full-scale aggregated feature maps in order to optimize the combined loss function to enhance organ boundaries. Specifically, the proposed method improves four major changes compared to Recalling U-Net: separate down sampling process, improved encoding path, upgraded bridge and improved recalling block layers. These changes help improve the learning ability and efficiency of the proposed method, resulting in more accurate segmentation results. The proposed method is tested on the Kvasir-SEG, CVC-ClinicDB and CVC-ColonDB datasets with evaluation metrics such as Dice coefficient, Intersection over Union and accuracy. Experimental results presented that the proposed method achieved an accuracy of 98.992%, 97.628%, and 97.695% on these datasets, respectively, and outperforms the results of several recent methods on the same dataset.
Retinal vessel segmentation is a critical step for the diagnosis and monitoring of ocular and cardiovascular diseases. Despite advances in medical imaging, achieving accurate segmentation across varying fundus image resolutions remains challenging. Conventional convolutional neural networks (CNNs) often fail to capture both fine and large vascular structures due to repeated pooling operations and the limitations of static convolution kernels. To address these issues, we introduce in this work a robust and efficient CNN architecture specifically designed for retinal vessel segmentation. The key innovations of our approach consist of: (1) providing multi-scale input strategy applied across all downsampling blocks to mitigate pooling drawbacks, (2) adopting a dual encoding mechanism with heterogeneous kernel processing to enrich feature extraction, and (3) fusing the extracted features that combine complementary encoder representations before decoding, thereby preserving spatial detail and enhancing discriminative features. The proposed method is evaluated on the DRIVE and HRF databases, which differ substantially in image resolution. The proposed network demonstrates strong performance, achieving a mean accuracy and sensitivity in the order of 97.62% and 97.22% on the DRIVE database, and in the order of 83.43% and 84.76% on the HRF database, respectively. Notably, the best fold attains sensitivity rates of 88.8% on DRIVE and 87.46% on HRF, highlighting the robustness of the method across both databases and their evaluation folds.
Accurate classification of breast ultrasound images (BUSI) is crucial for the early diagnosis of breast cancer. Current deep learning models still have limitations in this task, primarily due to the insufficient global modeling capabilities of convolutional neural networks (CNNs) and the lack of local detail and spatial inductive bias in vision transformers (ViTs). Additionally, how to effectively fuse multiscale features to address lesions of varying sizes remains a major challenge in this field. To address these issues, this paper proposes a novel multiscale convolution retentive transformer (MS-CRTU) network. The network first uses a CNN module to extract key local texture features from ultrasound images. Subsequently, to synergistically model local and global information from shallow to deep levels, the feature maps are fed into the core processing stage composed of our designed convolution-retentive transformer block (CRT block). This block interacts with information through convolutional operations and a novel Manhattan self-attention (MaSA) mechanism. Finally, to dynamically aggregate the most informative features, a selective fusion module (MSF) integrates the multiscale features from all stages for the final classification task. On the public BUSI and our B-UCLM datasets, the MS-CRTU model achieves accuracies of 95.38% and 93.85% respectively, outperforming baseline models in F1 score and area under the curve (AUC). This study confirms that the proposed MS-CRTU enhances the accuracy and robustness of BUSI classification, offering a new approach for intelligent diagnosis of BUSIs.
Timely diagnosis and treatment of skin cancer is crucial for improving patient prognosis. Convolutional neural networks (CNNs) have been applied to skin cancer image classification, but they are difficult to fully capture feature pose information. To address it, a dual-channel Attention Routing capsule network (DC-AR-CapsNet) is established to take advantage of capsule networks for skin cancer image classification. To overcome the limitations associated with inadequate feature extraction capabilities, a hybrid domain feature-weighted DCS module is proposed. By combining with multi-scale feature fusion, a two-branch parallel architecture is established to enhance the ability of capsule network in capturing features from skin cancer images. Then a self-attention mechanism is incorporated to reduce redundancy of the primary capsules and inhibit the homogenization of capsule features by calculating the cosine similarity values of capsules in the same layer. Furthermore, a dynamic loss function is constructed to adaptively adjust the decision boundary based on the classification outcomes. Experimental results indicate that the designed DC-AR-CapsNet model achieves superior image classification accuracy, despite the increase in the number of parameters. It outperforms existing models and achieves excellent accuracy and robustness in skin lesion classification.
ABSTRACT Feature extraction techniques such as Wavelet Transforms (WT), Histogram of Oriented Gradients (HOG), Gray Level Co‐occurrence Matrix (GLCM), Gabor filters, and Local Binary Pattern (LBP) are widely used in medical imaging for dimensionality reduction and enhanced image analysis. However, reconstructing the original image from these extracted features is crucial for interpretability, validation, and ensuring data integrity. This study aims to propose a combined approach for reconstructing images using wavelet coefficients and feature extraction techniques, ultimately improving the accuracy and reliability of image reconstruction in medical imaging applications. The wavelet coefficients are processed through Inverse Wavelet Transform (IWT), while a regression model is employed to map extracted features to image details. The combination of multiple reconstruction models—WHOG, WGLCM, WGabor, and WLBP—are compared based on performance metrics such as Peak Signal‐to‐Noise Ratio (PSNR), Mean Square Error (MSE), Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Structural Similarity Index (SSIM). Specifically, WHOG refers to the combination of wavelet coefficients and HOG, WGLCM is the combination of WT and GLCM, WGabor integrates WT and Gabor filters, and WLBP combines wavelet coefficients with LBP. The results from three different imaging modalities (Computed Tomography scan for a breast image, Magnetic Resonance Imaging for a brain image, and X‐ray for a chest image) show that LBP‐based reconstruction delivers the highest fidelity, making it the most robust method for feature‐based image reconstruction.
ABSTRACT Clustering is a foundational paradigm in data mining and pattern recognition aimed at grouping data and uncovering meaningful clusters. This data being clustered can be bounded and exhibit non‐Gaussian characteristics in feature spaces, where traditional approaches may contend challenges. Moreover, the presence of irrelevant features can obscure latent structure, undermining both cluster quality and downstream decision‐making. In this work, we address these challenges by proposing a Bayesian Libby‐Novick Beta mixture model (BLNBMM) with integrated feature selection. To enable posterior inference in our proposed hierarchical model, we develop a variational inference (VI) framework that provides uncertainty quantification. Our model flexibly captures relevant features using the Libby‐Novick Beta distribution. Experiments on medical datasets with varying complexity demonstrate that BLNBMM effectively captures complex class distributions in bounded data domains.
Lung cancer, the largest cause of cancer deaths globally, requires effective diagnostic imaging to identify and define pulmonary abnormalities. Sparse view projections lower CT radiation but diminish image quality and make diagnosis challenging. This research introduces a comprehensive approach that systematically assesses seven sparse projection levels (10-512 views) for lung cancer detection, uniquely integrating generative adversarial networks (GANs)-based reconstruction with diagnostic performance evaluation. In contrast to previous research, our study offers a systematic framework for generating and assessing sparse-view datasets over the entire clinical range. Given the limited availability of sparse-view projection datasets across clinical ranges, we systematically generated sparse-view images at multiple levels: 10, 16, 32, 64, 128, 256, and 512 projections. Three GAN architectures, standard GAN, conditional GAN (CGAN), and Pix2Pix, were employed to reconstruct high-quality images from degraded sparse-view data, enhancing structural integrity and diagnostic utility. A Visual Geometry Group 16-layer (VGG16) deep learning (DL) model evaluated the diagnostic efficacy of both generated and original sparse-view images. Results demonstrate superior performance of the Pix2Pix model, achieving structural similarity index measure values of 79.385% and 81.265% for 10- and 16-view projections, respectively. Classification performance using VGG16 yielded exceptional metrics: 98.82% accuracy, 99.18% precision, 99.18% recall, 99.18% F1-score, and 99.86% area under the receiver operating characteristic curve for 10-view projections. This GAN-DL integration offers a clinically viable approach for sparse-view imaging, ensuring diagnostic reliability while minimizing radiation exposure and enhancing computational efficiency, particularly critical for cancer patients requiring reduced radiation doses.
As modern medical imaging technology advances, resting‐state functional magnetic resonance imaging (rs‐fMRI) is becoming a preferred method for studying brain activity and identifying autism spectrum disorder (ASD) because of its cost‐effectiveness and noninvasive approach. To better utilize the temporal and spatial dimensions of the rs‐fMRI signals, this paper proposes a spatiotemporal feature fusion adaptive learning graph neural network (SF 2 AL‐GNN) for ASD diagnosis. SF 2 AL‐GNN first creates a functional connectivity (FC) matrix for every individual. Then, combining gated recurrent units (GRU), transformer, and graph convolution, it develops a spatiotemporal local feature learning module to extract temporal features from 1D time series and the FC matrix–based spatial features. Subsequently, these features are fused to construct a global subject graph using multimodal information. A self‐adjusting global feature learning (SGFL) module adds adaptive weights during GCN updates to better obtain ASD‐related feature embeddings. Finally, an MLP is used for classification. Training and evaluation of the model were conducted using the ABIDE I dataset, outperforming the latest advanced approaches.
Portable near-infrared (NIR) brain imaging may support point-of-care triage when computed tomography (CT) is unavailable or delayed, but its practical value depends on whether novice, non-specialist operators can use the device quickly and maintain scan quality after brief training. Thirty-two right-handed adults with no prior NIR experience received a brief standardized training session of approximately 2 min on the ArcheOptix NIRD device, which detects intracranial hemorrhage by tracking hemoglobin absorption during guided scalp scans. Operators then completed two full-head scans on a healthy volunteer: an initial competency assessment and a follow-up assessment after 1 day without refresher training. Performance metrics included total scan time, repeat scans prompted by loss of contact or light leaks, and mean scanpath time as an index of handling efficiency and consistency. Scan quality was evaluated using loss of probe contact and noise during acquisition. User experience was measured after each scan with the 10-item System Usability Scale. All operators completed both sessions. Median performance improved from the initial to the follow-up assessment, with total scan time decreasing from 5 min 27 s to 2 min 53 s. The proportion of excellent scans completed in under 5 min increased from 50% to 84%, while poor scans taking more than 10 min fell to zero. Repeat scans declined from 38 to 24, and mean scanpath time shortened while becoming more consistent. Lift on dark and Noise on dark remained stable, indicating no degradation in signal quality as operators worked faster. System Usability Scale scores improved from 69.4 to 76.5, reflecting greater perceived ease of use and confidence. After minimal training, novice operators achieved rapid and reliable NIR scans with faster performance, fewer repeat scans, stable signal quality, and improved usability ratings, supporting the feasibility of rapid training and operation of portable NIR devices in emergency, sideline, and remote head trauma settings.
As one of the most typical clinical manifestations of fluorosis, accurate segmentation of dental areas is quite significant for early disease assessment, grading diagnosis, and clinical intervention. However, traditional manual examination is highly subjective and often leads to misdiagnosis or missed diagnosis. Existing deep learning segmentation methods face challenges such as the variable morphology of dental fluorosis, indistinct edges, and the lack of explicit geometric constraints. To overcome these challenges, we propose a dental fluorosis segmentation model, DSSL-UNet, based on the U-Net backbone, integrating dynamic snake convolution (DSConv) and a learnable shape prior module (LSPM). DSConv introduces a learnable offset field that adaptively deforms convolutional kernels along curvilinear tooth contours, thereby strengthening feature extraction for indistinct and tortuous edges. The LSPM consists of self-updating blocks (SUB) and cross-updating blocks (CUB), which model long-range dependencies and local shape priors to improve the continuity and accuracy of segmentation masks. Comprehensive experiments on both public and custom dental fluorosis datasets demonstrate that DSSL-UNet outperforms mainstream models across all metrics, providing more precise and complete segmentation of dental fluorosis. This model provides robust technical support for the automated and accurate diagnosis of dental fluorosis.
Pneumonia classification from chest x-ray images remains a challenging task because different pneumonia categories often exhibit highly similar visual manifestations, while subtle lesion regions require both fine-grained local detail perception and effective global contextual modeling. To address these challenges, we propose PneumoMamba, a dual-path network for computer-aided pneumonia diagnosis that combines a convolutional neural network branch for local texture extraction with a Mamba-based branch for efficient long-range dependency modeling. In the proposed framework, the convolutional branch focuses on capturing detailed local structures, whereas the Mamba-based branch is designed to model global contextual information with linear computational complexity. To better preserve spatial continuity for sequence modeling, we further introduce an eight-directional scan strategy (8DScan), which converts two-dimensional feature maps into multidirectional scan sequences for more comprehensive spatial dependency modeling. In addition, a Multi-Scale Asymmetric Convolution (MSAConv) and a focused feature module (FFM) are incorporated to enhance the representation of subtle pneumonia-related patterns. We evaluate the proposed method on the Pneumonia-CXR-Database, which contains four classes of chest x-ray images, under a 6:2:2 train/validation/test split. Experimental results show that PneumoMamba achieves a test accuracy of 94.94%, together with strong F1-score, sensitivity, and precision, outperforming representative comparison methods. These results demonstrate the effectiveness of the proposed framework for computer-aided pneumonia classification.
Lung cancer is one of the cancers with the highest incidence and mortality rates worldwide. The large volume of CT images and the limited resources of radiologists have highlighted the demand for computer-aided diagnostic (CAD) systems. Among existing methods, interpretable capsule networks face prominent challenges in pulmonary nodule classification, such as insufficient feature extraction and high computational costs associated with routing algorithms. To address these issues, this study proposes the BiCaps-VG model, which comprises three components: dual-branch feature extraction, capsule vector routing, and multi-task learning. One branch incorporates a receptive field adaptation module that employs multi-receptive field parallel heterogeneous convolutional paths to extract scale-sensitive nodule features. The other branch is a hierarchical semantic recovery module, which first adopts an encoder-decoder architecture and then leverages a spatial attention mechanism to model global contextual semantics. These two branches are fused to form a unified semantic representation space, thereby enhancing the representational capacity of nodule features. The capsule routing module adopts a variance-guided capsule routing algorithm that calculates activation values based on the variance between predicted low-level capsules and high-level capsules, thus avoiding iterative routing and reducing computational overhead. In the multi-task module, the extracted nodule attribute capsules are used for malignancy prediction, attribute grade classification, and mask reconstruction. Experimental results demonstrate that BiCaps-VG achieves an accuracy of 94.88% in pulmonary nodule malignancy grading, representing a 1.58% improvement over the baseline model. Combined with its inherent interpretability, the model exhibits considerable potential for clinical application. The code is available at https://github.com/fzhou924anpan-prog/BiCaps-VG.
Uncertainty quantification is also an area of classification of Alzheimer's Disease (AD) that has been largely underexplored using multimodal deep learning. Although there have been significant successes in the detection of AD, the factor of predictive uncertainty has not been widely addressed, especially in relation to the diagnosis of AD. The precision of uncertainty estimation is vital for assessing model reliability, as acquiring multimodal data is inherently complex and diagnostic errors may result in very high clinical risks. The inclusion of uncertainty measures can make models more interpretable and indicate predictions that may require additional clinical evaluation. This paper proposes a new density-based evidential network, the Kernel Density Estimation Evidential Deep Learning Network (KDE-EDLNet), that captures aleatoric uncertainty and epistemic uncertainty. Multimodal inputs to the model were tested using the Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Open Access Series of Imaging Studies (OASIS) databases. The results of the experiment indicate that the given approach offers superior predictive performance than current techniques and underscore the importance of quantifying uncertainty to improve the reliability and clinical validity of model predictions.
In the trial CareForColon2015 (CFC2015) we included more than 2000 colon cancer screening participants for colon capsule endoscopy (CCE). We detected distinct fluctuations in the rate of complete investigations defined as CCE with both complete transit and adequate bowel cleansing. We aimed to investigate possible contributing factors to incomplete investigations and compare fluctuations in manual and artificial intelligence (AI) generated cleansing evaluations. We retrieved CCE videos and relevant information from CFC2015. We considered the following factors as possible contributors to completion rate fluctuations: different outpatient clinics for CCE, individual nurses instructing participants, number of participants at the information session, timing of the last dose of bowel preparation, timing of capsule ingestion, time gap between the last dose of bowel preparation and capsule ingestion, manual CCE reader-cleansing evaluations. During the trial, the prokinetic prucalopride was added to the regimen. We compared the completion rates between different time periods of the trial. Monthly completion rates ranges from 56.6% to 75.7%. The fluctuations were caused by variation in cleansing quality and not in the proportion of complete transit. Early capsule ingestion showed increased odds of a complete investigation. The timing of the last dose of bowel preparation was associated with cleansing quality. Fluctuations in cleansing quality were detected by manual reading in all colonic segments but by the AI algorithm predominantly in the right colon. Timing of bowel preparation and capsule ingestion is an important factor for the bowel cleansing quality in CCE. Using AI for cleansing quality monitoring in a CCE cohort can be an important measure for early identification of changes.Trial Registration: ClinicalTrials.gov: NCT04049357