Abstractive summarization remains a computationally intensive task, posing significant challenges for the efficient deployment of large-scale encoder-decoder models. In this paper, we propose a complexity-aware synergistic distillation framework that balances model compactness with semantic fidelity. First, we introduce a teacher enhancement utilizing a triple-objective optimization scheme to ensure the teacher model provides robust and semantically enriched supervision without requiring external signals. Subsequently, the framework employs a complexity-aware layer recommendation mechanism that quantifies dataset-specific attributes into a complexity score to dynamically determine the optimal student depth. Furthermore, we adopted a multi teacher layer overlapping alignment and collaboration (MLOAC) strategy to synergistically aggregate knowledge from multiple adjacent teacher layers via learnable collaboration weights, thereby optimizing cross-layer information flow. Experiments across diverse benchmarks demonstrate that our framework consistently preserves near-teacher performance under significant compression ratios. The integration of complexity-based adaptability and overlapping knowledge transfer yields an efficient summarization solution that is highly robust and complementary to existing inference-time acceleration methods.
In recent years, unmanned aerial vehicles (UAVs) have been extensively employed in traffic monitoring and intelligent surveillance. However, challenges such as dense targets and intricate backgrounds in aerial imagery have hindered the performance of existing detectors in this domain. To overcome these limitations, we introduce STCM-YOLO, a UAV-based small object detection algorithm. First, to mitigate the issue of insufficient local feature capture, we propose the stabilizer trim control module (STCM), which integrates the Swin Transformer with a lightweight C3 module. Second, the convolutional block attention module (CBAM) is incorporated into the backbone network to enhance the model's robustness in complex backgrounds. CBAM not only improves the saliency of small objects in complex environments but also effectively mitigates feature dilution and spatial localization challenges. Furthermore, we introduce the concept of a minimal enclosing region in the intersection over union loss function, addressing optimization issues related to non-overlapping bounding boxes and thereby improving detection accuracy. Finally, in order to enhance the model's ability to identify small-scale items, a fourth detection layer is incorporated in addition to the original three detection layers. Experimental results demonstrate that on the VisDrone2019 dataset, STCM-YOLO achieves a 5.5% improvement in mAP50 over YOLOv7. Moreover, on the CARPK dataset, STCM-YOLO surpasses YOLOv7 by 3.3% in mAP50, underscoring its effectiveness in UAV-based small object detection.
The performance and stability of electronic equipment may be impacted by a variety of flaws in PCBs (Printed Circuit Boards) that arise throughout the industrial production process. These defects include soldering problems, short circuits, open circuits and physical damage on the board. Regarding the above problems, this paper proposes ODC-YOLO, an improved PCB defect detection method based on YOLOv5s, which firstly introduces the dynamic convolution module (ODConv) in the backbone network to enhance the network's dynamic perception of multi-scale features; secondly, integrates the CBAM attention mechanism in the feature fusion stage to strengthen the localization and identification of key defective regions; and lastly, employs the MDD_ PCB dataset for comparison experiments to verify the advantages of the improved algorithm in detection accuracy and speed. The experimental results show that the mAP@0.5 of the improved model can reach 97.1%, which is 1.1% higher than that of the original model, thus demonstrating that ODC-YOLO performs better in defect detection with higher recall and lower false detection rate, and also proving its feasibility and practicability in real industrial environments.
In rice pest management, accurate pest detection is critical for intelligent agricultural systems, yet challenges like limited dataset availability, pest occlusion, and insufficient small object detection accuracy hinder effective monitoring. To address the aforementioned challenges, this study presents YOLO-PEST, an innovative detection approach based on the YOLOv5s architecture to address these issues. YOLO-PEST collects rice pest images from multiple channels and images are randomly cropped to occlude detection boxes, effectively simulating pest overlapping scenarios. During the feature fusion process, the ConvNeXt module is integrated to improve the detection accuracy for small objects via multiscale feature extraction. Additionally, the CoTAttention mechanism is incorporated to enhance the model's robustness under complex environmental conditions. Comparative experiments show that the YOLO-PEST approach achieves a 97% of mAP@0.5, representing a 1.4-point improvement compared with previous methods, thus verifying its effectiveness in rice pest management.
Pavement defect detection greatly affects pavement service life and vehicle operation safety. Current pavement defect detection models encounter difficulties in accurately detecting minor defects, handling imbalanced class samples, and maintaining speed. To overcome these issues, we propose a fast and improved pavement surface defect detection model named MMS-YOLOv10, which is based on YOLOv10n. This model includes three significant improvements. First, we incorporate the multidimensional collaborative attention (MCA) mechanism into the C2f module of the backbone network to enhance adaptability to objects of different scales and improve the feature extraction ability. Second, we design a multilevel feature fusion (MFF) module to enhance semantic and detailed information and improve the feature expression ability of the model with different levels of features. Third, the sample correlated weighting loss function is introduced during network training to solve the issue of sample imbalance through the dynamic weight mechanism. The performance of the MMS-YOLOv10 model is assessed through experiments conducted on the famous RDD2022 dataset. The qualitative and quantitative results show that the proposed model can lead to promising improvements in detection accuracy. Through further ablation experiments, the components of the proposed model are validated to achieve performance improvements.
Spectral computed tomography (CT) provides the potential to generate attenuation maps at varying spectral bins, which can subsequently be employed for the resolve of tissue materials. However, the majority of reconstruction algorithms employed for every projection energy bin typically exhibit a considerable degree of noise. Recently, a series of techniques have been designed in spectral CT reconstruction based on traditional iterative models or deep learning (DL) methods. However, these works are independent or simply coupled, often neglecting the dependency relationships of spatial and spectrum. Additionally, the interpretability and generalization capabilities remain as formidable challenges for current methods. To effectively address these challenges, we initially introduce a spatial-spectral convolutional sparse coding (SS-CSC) framework, which aims to jointly represent spatial and spectral features in a unified manner. Subsequently, we devise a novel deep unfolding network, inspired by SS-CSC, for spectral CT image representation. We refer to this component as the SS-CSC module. Furthermore, we expand the iterative reconstruction scheme for constructing an interpretable deep SS-CSC iterative reconstruction network (SS-CSC-Net) model for spectral CT imaging. The SS-CSC-Net is composed of several iteration blocks, with each block containing two modules: image reconstruction and SS-CSC modules. The spectral CT image is updated using a deep learning-based spatial-spectral prior constraint. Experimental results from both qualitative results and quantitative analyses indicate that the proposed SS-CSC-Net excels in noise reduction and in maintaining tissue edge integrity, delivering superior overall performance.
Purpose Numerous techniques based on deep learning have been utilized in sparse view computed tomography (CT) imaging. Nevertheless, the majority of techniques are instinctively constructed utilizing state-of-the-art opaque convolutional neural networks (CNNs) and lack interpretability. Moreover, CNNs tend to focus on local receptive fields and neglect nonlocal self-similarity prior information. Obtaining diagnostically valuable images from sparsely sampled projections is a challenging and ill-posed task. Method To address this issue, we propose a unique and understandable model named DCDL-GS for sparse view CT imaging. This model relies on a network comprised of convolutional dictionary learning and a nonlocal group sparse prior. To enhance the quality of image reconstruction, we utilize a neural network in conjunction with a statistical iterative reconstruction framework and perform a set number of iterations. Inspired by group sparsity priors, we adopt a novel group thresholding operation to improve the feature representation and constraint ability and obtain a theoretical interpretation. Furthermore, our DCDL-GS model incorporates filtered backprojection (FBP) reconstruction, fast sliding window nonlocal self-similarity operations, and a lightweight and interpretable convolutional dictionary learning network to enhance the applicability of the model. Results The efficiency of our proposed DCDL-GS model in preserving edges and recovering features is demonstrated by the visual results obtained on the LDCT-P and UIH datasets. Compared to the results of the most advanced techniques, the quantitative results are enhanced, with increases of 0.6-0.8 dB for the peak signal-to-noise ratio (PSNR), 0.005-0.01 for the structural similarity index measure (SSIM), and 1-1.3 for the regulated Fréchet inception distance (rFID) on the test dataset. The quantitative results also show the effectiveness of our proposed deep convolution iterative reconstruction module and nonlocal group sparse prior. Conclusion In this paper, we create a consolidated and enhanced mathematical model by integrating projection data and prior knowledge of images into a deep iterative model. The model is more practical and interpretable than existing approaches. The results from the experiment show that the proposed model performs well in comparison to the others.
In practical scenarios, sparse view computed tomography (CT) is a successful method for reducing the X-ray radiation dose and imaging time. Generating high-quality images with undersampled projection data using conventional techniques is difficult. To reconstruction high quality CT image from the situation of undersampling, we proposed a weighted sparse constraint reconstruction (WSCR) framework in this work. Motivated by the concept of deep convolutional neural network, we expand the iterative reconstruction scheme for constructing reweighted sparse representation statistically, with a specific number of iterations for training based on data-driven approach, resulting in the development of a reconstruction network based on WSCR. The WSCR network is composed of several iteration blocks, with each block containing two modules: image reconstruction and sparse representation network modules. The CT image is updated using a deep learning-based prior constraint in the image reconstruction module. To further restrict the process of reconstructing the image, the reweighted feature map and filter updating are implemented using a sparse representation network. The results of the experiment showed that our suggested WCSR network was able to achieve superior performance in removing artifacts and preserving edges.
To prevent and control agricultural diseases and insect pests, the timely detection and accurate identification of crop diseases and insect pests are significant. Studies have shown that pests on plant surfaces are challenging to detect because of their small size and strong camouflage. Therefore, to better detect pests on citrus leaves, a citrus disease and insect pest detection method based on a double backbone network is proposed. The double backbone network-improved Single Shot MultiBox Detector (SSD) model was used to detect citrus images. The accuracy and recall rate of the neural network target detection were evaluated, and the robustness was verified by analyzing the detection results. The experimental results showed that the trained network's mean average precision (mAP) on the test dataset was 72.54%. In addition, the model showed high robustness on citrus pest datasets, with mAP reaching 86.01%. The results showed that the method was accurate and efficient compared with other target detection methods and could be applied to detect and control citrus pests and diseases.
针对计算机断层扫描(Computed Tomography,CT)中因采用低剂量扫描方式,导致图像噪声伪影干扰,尤其是不同部位噪声和伪影强度存在较大差异这一问题,提出了一种基于动态可控残差的卷积神经网络(DC-ResNet)算法.DC-ResNet的主要思想是在常规残差网络连接中添加一个图像质量指导的控制变量,以允许获取残差的加权和,从而实现残差特征的动态可控.所设计的DC-ResNet网络是一种包括两个子网络的组合型结构,一个是作为主干网络的基础子网络,在该子网络中使用全局动态残差块和局部动态残差块来实现低剂量CT图像质量的提高;另一个是作为辅助网络的条件子网络,用来生成基础子网络中不同动态可控残差块的权值,辅助基础子网络的学习.通过Mayo与UIH数据实验验证,其视觉结果表明:处理后的不同部位CT图像噪声伪影均能够得到较好的抑制,并能有效地保留结构细节及组织纹理;量化结果表明:处理后的CT图像峰值信噪比(Peak-Signal to Noise Ratio,PSNR)和结构相似性(Structure Similarity,SSIM)均优于对比方法.
Low-dose computed tomography (LDCT) is desirable due to ionizing radiation, but the resulting images suffer from serious streak artifacts and spot noise. Recently, deep learning (DL)-based methods have emerged as promising alternatives for medical image processing. However, most DL-based methods are built intuitively and lack interpretability, and it is difficult to effectively separate the artifacts and noise in LDCT images. Obtaining diagnostically useful images, especially when using a low-dose scanner protocol, remains an open challenge. To improve the quality of LDCT images, we developed a novel processing network called the sparse autorepresentation U-Net (SureUnet). First, inspired by multilayer convolutional sparse coding (CSC), we constructed a sparse autorepresentation encoder to sufficiently capture and represent hierarchical image features. Then, we chose the widely used U-Net model for sparse autorepresentation block applications and designed SureUnet by adding a feature decoding block. Therefore, every module has well-defined interpretability in our network. Additionally, hybrid loss functions were specifically designed, including the mean absolute error, edge loss and perceptual loss. Through the cooperation of multiple loss functions, the noise artifact suppression effect of the network was improved. The visual results obtained on the MAYO and UIH datasets show that the proposed method’s noise artifact suppression effect was more significant. The quantitative results showed promising improvement levels compared to those of the other state-of-the-art methods. The SureUnet model significantly outperformed the compared methods on two datasets, with margins of 0.4 dB for the PSNR, 0.007 for the SSIM, and 1.6 for the FID on the MAYO dataset and margins of 0.5 dB for the PSNR, 0.004 for the SSIM and 2.9 for the FID on the UIH dataset. This work paves the way for sparse autorepresentation in DL for processing LDCT images. Experimental results have demonstrated the competitive performance of SureUnet in terms of noise suppression, structural fidelity and visual impression improvement.
Imaging in the field of low-dose computed tomography (LDCT) tend to be rather noisy and artificial but is diagnostically useful. One approach to improve the quality of LDCT images is to use deep learning (DL) techniques. DL-based methods produce state-of-the-art performance in low-level medical image restoration tasks but remain defect to interpret due to their black-box constructions. In this paper, we present a simple yet effective LDCT image denoising model by combining the advantages of a residual strategy and a multilayer convolutional analysis-based sparse encoder (CASE). Inspired by convolutional sparse coding (CSC), we constructed a multilayer CASE to sufficiently capture and represent hierarchical image features and designed CASE-net to achieve improved LDCT noise artifact suppression. Moreover, a hybrid loss function, e.g. mean absolute error (MAE) loss, edge loss and perceptual loss, was used to achieve better denoising effects. Experiments on the MAYO and UIH datasets demonstrated the performance of our framework. The results prove that the proposed approach can restrain noise and artifacts and maintain tissue structure during the LDCT imaging.
Background and objectiveLow-dose computed tomography (LDCT) has become increasingly important for alleviating X-ray radiation damage. However, reducing the administered radiation dose may lead to degraded CT images with amplified mottle noise and nonstationary streak artifacts. Previous studies have confirmed that deep learning (DL) is promising for improving LDCT imaging. However, most DL-based frameworks are built intuitively, lack interpretability, and suffer from image detail information loss, which has become a general challenging issue.MethodsA multiscale reweighted convolutional coding neural network (MRCON-Net) is developed to address the above problems. MRCON-Net is compact and more explainable than other networks. First, inspired by the learning-based reweighted iterative soft thresholding algorithm (ISTA), we extend traditional convolutional sparse coding (CSC) to its reweighted convolutional learning form. Second, we use dilated convolution to extract multiscale image features, allowing our single model to capture the correlations between features of different scales. Finally, to automatically adjust the elements in the feature code to correct the obtained solution, a channel attention (CA) mechanism is utilized to learn appropriate weights.ResultsThe visual results obtained based on the American Association of Physicians in Medicine (AAPM) Challenge and United Image Healthcare (UIH) clinical datasets confirm that the proposed model significantly reduces serious artifact noise while retaining the desired structures. Quantitative results show that the average structural similarity index measurement (SSIM) and peak signal-to-noise ratio (PSNR) achieved on the AAPM Challenge dataset are 0.9491 and 40.66, respectively, and the SSIM and PSNR achieved on the UIH clinical dataset are 0.915 and 42.44, respectively; these are promising quantitative results.ConclusionCompared with recent state-of-the-art methods, the proposed model achieves subtle structure-enhanced LDCT imaging. In addition, through ablation studies, the components of the proposed model are validated to achieve performance improvements.
Small objects in traffic scenes are difficult to detect. To improve the accuracy of small object detection using images taken by unmanned aerial vehicles (UAV), this study proposes a feature-enhancement detection algorithm based on a single shot multibox detector (SSD), named composite backbone single shot multibox detector (CBSSD), which uses a composite connection backbone to enhance feature representation. First, to enhance the detection effect of small objects, the lead backbone network, VGG16, is kept constant, and ResNet50 is added as an assistant backbone network, and the residual structure in ResNet50 is used to obtain lower feature information. The obtained lower feature information is then fused to the lead network through feature fusion, allowing the lead network to retain rich lower feature information. Finally, the lower feature information in the prediction layer increases. The experimental results show that CBSSD has a significantly higher recognition rate and a lower false detection rate than conventional algorithms, and it still maintains a good detection effect under low illumination. This is of great significance to small object detection using images taken by UAVs in traffic scenes. Furthermore, a method to improve the SSD algorithm is proposed.
With the explosion of deep learning algorithms, big data, and high-performance computing, deep learning has flourished in the fields of medical analysis and image processing. In this paper, we present a simple yet effective model for low dose computed tomography (CT) image processing procedure, by combining with the advantages of residual convolution network and convolutional sparse coding (DRCSC). Through the learned iterative shrinkage threshold algorithm (LISTA), we extend convolutional sparse coding to its convolutional learning from and entirely following the residual convolution network structure, which improves the network's interpretability. The network workflow consists of three components: input feature maps prepare, recursive manner for feature maps learning by convolutional sparse coding, and high-frequency information recover. Within the residual learning strategy, the deep network training become easier and preservation more detail feature. Experiments on AAPM datasets has shown the efficacy of our method. Network testing results identify that the proposed method can restrain of artifacts and noise oscillations for low dose CT imaging.
Low light image enhancement is a challenging task, and it has been a hot research topic. Inspired by retinex theory and U-Net network, we propose a U-Net-based multiscale feature preserving method for low light image enhancement, which can realize the extraction of low-level features and high-level semantic features. Before feature extraction, we carry out multiscale pre-extraction processing on the image to improve the feature extraction ability of the network. Considering the discontinuity between low-level features and high-level semantic features, we propose a spatial consistency method to maintain the global feature correlation. Finally, we propose a new multiscale structure calculation method, which greatly alleviates the phenomenon of uneven illumination and color deviation after enhancement and makes the enhancement results more consistent with human visual perception. Extensive experiments demonstrate that compared with other advanced enhancement methods, our method has better enhancement effect and can retain more details. The enhanced image not only has good visual perception but also is better than other methods in objective evaluation. (C) 2021 SPIE and IS&T
The potential radiation damage in CT scans have been receiving increasing attention. However, reducing the scan dose will degrade the image quality and affect the diagnosis results. Aiming at addressing the above problems, a three-dimensional (3D) reconstruction algorithm combining convolutional sparse coding and gradient L-0-norm is proposed herein. The proposed algorithm uses the frequency decomposition reconstruction form to perform unsupervised multiscale online convolution sparse-coding constraints on high-frequency components, and gradient L-0-norm constraints on low-frequency components to achieve the suppression and organization of noise artifacts in low-dose CT imaging keep the details. Moreover, three different scales of 3D filter sets are used in convolutional sparse coding, which can effectively adapt to the feature information at different scales and improve the coding ability. The experimental results of abdominal CT simulation data and real-time scan data show that the proposed algorithm can obtain fewer noise artifacts, high contrast in structural details, and better imaging results in the reconstruction process of 25% conventional dose.