
ABSTRACT Recently, images have been used widely in a variety of fields, including social media, the military and medicine. Sensitive images are often transmitted over insecure networks and chaotic maps offer effective cryptographic properties for securing these transmissions. One‐dimensional (1D) chaotic maps are commonly employed in data encryption due to their simplicity, high randomness properties, excellent chaotic behaviours and high security levels. However, numerous chaos‐based image encryption algorithms have drawbacks that are generally related to the chaotic systems used and their encryption structures. To overcome these limitations, this paper first proposes a hybrid chaotic system based on 1D‐sine‐logistic‐sine‐tent (1D‐SLST) chaotic maps. Using 1D‐SLST chaotic generators, we produced a new 1D chaotic map exhibiting enhanced chaotic properties compared to the conventional maps. To examine its application in multimedia security, an effective image encryption algorithm is also introduced that leverages the chaotic behaviour of 1D‐SLST. Our algorithm utilises a variety of confusion operations, including zigzag transform, image rotation and pixel confusion. A pixel adaptive diffusion operation based on modulo arithmetic is also introduced. Simulation experiments and security performance evaluations confirm the effectiveness of the proposed scheme, which gives information entropy, NPCR and UACI values of 7.998866, 99.60876% and 33.46990%, respectively, which are close to ideal values. Comparative simulation results indicate that the proposed image encryption algorithm ensures a high level of security and fast encryption speed, achieving superior performance compared to several state‐of‐the‐art methods.
ABSTRACT Blind image watermarking plays a major role in copyright protection and digital media authentication. Maintaining robustness against rotation attacks remains a challenging problem. This paper proposes a saliency‐guided adaptive watermarking framework based on discrete wavelet transform, with improved robustness against rotation. The proposed method employs a pre‐trained U 2 ‐Net heatmap to guide both adaptive subband selection and embedding strength. In adaptive subband selection, the primary embedding subband is selected from the gradient direction and watermark bits are then distributed across selected subbands using an alternating embedding. In addition, heatmap‐driven adaptive embedding strength is applied. For watermark extraction, the proposed framework integrates prediction‐based residual estimation and correlation‐based detection, with four extraction modes: prediction, weighted prediction, correlation and weighted detection. A rotation correction algorithm is introduced to counter rotation attacks. The algorithm evaluates candidate rotation angles and selects the optimal orientation based on the confidence score derived from detector responses. The proposed method is validated on standard test images and additional benchmark datasets under several attacks, including rotation, JPEG compression, scaling and noise perturbation. Empirical results show that the method maintains acceptable visual quality, with an average peak signal‐to‐noise ratio of 34.25 dB and a structural similarity index measure of 0.9951 and achieves robustness with a bit error rate of 0.0182 and a normalised correlation (NC) of 0.9635. These results indicate that the proposed method improves extraction robustness under rotation and JPEG compression, while the rotation correction process leads to additional computational cost.
ABSTRACT This article addresses the critical issue of fingerprint template protection and feature stability by putting forward a novel gradient‐oriented learned contourlet transform (GOLCT) framework for a robust biometric representation. Existing fingerprint hashing methods based on contourlet, curvelet and related multiscale transforms rely on fixed directional filter banks, which fail to adapt to the inherently non‐stationary ridge orientation patterns of fingerprints, leading to reduced discriminability under rotation, elastic deformation and partial impressions. To overcome these limitations, the proposed GOLCT learns adaptive wedge parameters that enable orientation‐aware and data‐driven multi‐scale decomposition based on the orientation statistics of fingerprints, in contrast to traditional contourlet‐ or curvelet‐based methods that utilise fixed directional filters. Furthermore, the subbands are stabilised through gradient‐response subspace stabilisation (GRSS) for local rotational alignment, while orientation‐weighted discriminant aggregation (OWDA) enhances inter‐subject feature distinctiveness by exploiting orientation‐consistent spectral responses. Finally, a subsampled randomised Hadamard transform (SRHT) is applied for secure dimensionality reduction and key‐dependent hashing. Experimental results on standard fingerprint databases (FVC2002 and FVC2004) demonstrate that the proposed GOLCT achieves improved performance than traditional contourlet‐based fingerprint hashing methods, achieving lower equal error rate (EER), higher peak signal‐to‐noise ratio (PSNR) and improved entropy preservation, thereby confirming its effectiveness for secure and cancelable fingerprint template protection.
ABSTRACT Knowledge distillation has become an important paradigm for transferring knowledge from high‐capacity teacher models to compact student models, yet its behaviour in object detection differs substantially from classification due to the joint optimisation of localisation and classification, multiscale feature representations, and complex spatial reasoning. This survey provides a systematic and evidence‐driven analysis of knowledge distillation techniques developed specifically for object detection based on a curated collection of research studies. A multi‐dimensional taxonomy is constructed by categorising methods according to the type of knowledge transferred and the architectural context in which distillation operates, including feature‐level, response‐level, localisation‐aware, relational, attention‐based, and cross‐modal strategies across convolutional, transformer‐based, three‐dimensional, and multi‐modal detection frameworks. The study further examines how distillation is integrated within detection pipelines such as backbone representations, feature aggregation modules, instance‐level predictions, and query‐based reasoning mechanisms, and synthesises consistent observations regarding strengths, limitations, and recurring challenges, including multi‐scale learning complexity, foreground–background imbalance, and teacher–student architectural mismatch. Without introducing new algorithms, the survey consolidates existing findings into a structured reference and outlines research directions supported by the surveyed literature, providing guidance for the analysis and design of future distillation‐based detection systems.
ABSTRACT The widespread use of generated media based on generative AI models, raises the risks of copyright violation and ownership conflicts. It has also introduced challenges such as deepfake misuse and unauthorised distribution of digital content. Watermarking methodologies, including embedding and tracing, have emerged as solutions for protecting ownership of digital media. This paper introduces a novel cryptographically secure deep learning‐based watermarking methodology. Unlike conventional methods where the given watermark is embedded uniformly across a digital image and recent deep learning‐based approaches that rely on end‐to‐end encoder‐decoder architectures operating over the entire image, this paper is the first to propose a method embedding the watermark in the specific non‐contiguous blocks in the image that were identified using convolutional neural network (CNN). This non‐linear, image‐specific watermark distribution introduces a dynamic unique ‘saliency‐aware’ embedding map for every individual image, significantly increasing the complexity to predict the watermark locations without knowing the right model weights; which improves the resistance to attacks. To address security limitations, the proposed method encrypts the watermark using RSA‐4096 and embeds the encrypted watermark in the selected blocks using the multi‐bit LSB technique. This approach combines the adaptability of deep learning with the robustness of cryptographic security and not only preserves the visual quality of the input image but also achieves a peak PSNR of 55.40 dB and extraction of the watermark with NCC deviation of zero. These results demonstrate the practicality of this proposed method over state‐of‐the‐art methodologies. The encryption of the given watermark before embedding, secures it and shows precise tamper detection (NPCR > 99%, UACI ≈ 33%). In addition to assuring the authenticity of the digital images, the introduced method provides a robust solution for applications requiring data integrity and provenance tracking such as misinformation detection, digital forensics, intellectual property protection, and content management systems.
ABSTRACT Coal classification is essential for industrial production, quality control and intelligent sorting systems. Traditional laboratory‐based chemical analysis methods are often destructive and time‐consuming, while manual visual inspection remains subjective and inconsistent. To address these challenges, this paper proposes DP‐KDNet (dynamic perception and knowledge distillation network), an efficient coal image classification network that integrates a dynamic perception attention mechanism (DPAM) with an online knowledge distillation strategy. The proposed DPAM enhances feature representation by dynamically fusing global statistical features, local variance information and multi‐scale contextual dependencies, enabling adaptive emphasis on discriminative coal surface characteristics such as lustre, texture and pore distribution. Meanwhile, the online distillation framework transfers knowledge from deeper network stages to shallow representations, improving classification robustness while maintaining computational efficiency. Experiments on a self‐constructed coal image dataset containing anthracite, bituminous coal and lignite demonstrate that DP‐KDNet achieves 91% classification accuracy, outperforming the baseline ResNet50 by 7.8 percentage points (83.2%→91.0%). The improvements are consistent across all metrics (precision, recall, F1‐score and ROC‐AUC), confirming the robustness and practical significance of the proposed method. Further analysis confirms improved feature separability and competitive computational efficiency, suggesting the proposed framework is well‐suited for practical deployment in industrial environments.
ABSTRACT Accurate detection of the neodymium rare‐earth molten salt flame is critical for automating neodymium metal production, as it directly determines the addition of neodymium oxide raw material. Thus, this study proposes LAD‐YOLOv8n, an enhanced YOLOv8‐based model tailored for neodymium rare‐earth molten salt flame detection. The method incorporates an image pre‐processing strategy based on a chromaticity‐space transformation, constructing a colour model tailored to the burning characteristics of molten salt to accurately extract its regions. With YOLOv8n serving as the backbone, the model integrates the asymptotic feature pyramid network for multi‐scale feature fusion, the large separable kernel attention mechanism for shape‐aware focus, and the DySample upsampler for enhanced flame localization. These modifications effectively mitigate challenges such as flame variability and background noise, thereby significantly enhancing the robustness of the detection process. Ablation experiments validate the individual contributions of each component and demonstrate their synergistic effects. Comparative assessments against SVM, LSTM, ResNet‐50, RT‐DETR‐l and other YOLO variants reveal that LAD‐YOLOv8n achieves a mean average precision of 77.0%—a 3.6% improvement over the original YOLOv8n. This confirms that the proposed deep learning recognition algorithm is a viable and effective approach for detecting rare‐earth molten‐salt flames.
ABSTRACT Accurate breast lesion segmentation in ultrasound (BUS) remains challenging due to low contrast and high variability. Current methods focus primarily on feature extraction but often fail to model the topological structure, leading to spurious disconnected predictions and inconsistent boundaries. We propose TCB‐Net, a novel architecture designed to ensure topological and geometric consistency through three key contributions. First, the lesion‐conditioned boundary attention gate (LCBAG) implements a coarse‐to‐fine feedback loop, suppressing background noise using predicted mask priors. Second, the boundary‐interior decoupled decoder (BIDD) utilises morphologically derived supervision to separate interior and boundary learning into distinct gradient signals. Third, the cross‐prediction geometric consistency loss (CPGCL) enforces spatial gradient identity between predictions and incorporates a differentiable soft Euler‐number regulariser to penalise topological errors. Evaluated on BUSI, UDIAT and BUS‐BRA datasets, TCB‐Net achieves Dice scores of 81.97%, 89.06% and 80.58% and HD95 of 12.67, 9.34 and 15.23 pixels, respectively. Notably, boundary‐sensitive metrics show even larger relative gains than Dice: TCB‐Net reduces HD95 by 15.1% relative to the strongest baseline (UMA‐Net) on BUSI, indicating that the proposed topology‐constrained design yields disproportionately large improvements in clinically relevant boundary precision. It outperforms seven state‐of‐the‐art baselines, including Attention U‐Net and HAU‐Net. Ablation studies confirm that our combined components yield a +7.87% Dice improvement over a ResNet34 U‐Net baseline on BUSI, demonstrating its effectiveness in producing clinically reliable, topologically sound segmentations.
ABSTRACT In this paper, we propose an automatic image segmentation (AIS) framework for medical imaging that integrates a heatmap attention mechanism (HMAM) into two deep neural network (DNN) architectures: U‐Net and generative adversarial networks (GANs). This HMAM guides the segmentation process by incorporating targeted spatial information that directs the networks toward clinically relevant regions. Heatmaps are generated using a modified VGG16 network, and the learned filters are subsequently transferred to conventional DNNs, whose activations provide complementary spatial guidance during segmentation. Furthermore, we develop two GAN‐based segmentation models that employ a U‐Net architecture and a U‐shaped modified VGG16 architecture as generator networks. In these models, the HMAM is incorporated exclusively into the generator network (GN), rather than the discriminator network (DN), because the GN directly produces the pixel‐wise segmentation mask and is therefore expected to benefit more from spatial guidance. The proposed methods are evaluated using COVID‐19 lung computed tomography (CT) images. Experimental results demonstrate consistent improvements over matched conventional baselines, particularly under limited‐data conditions. For example, the mean dice score of the U‐Net architecture increased from 61.05% to 70.49% on the larger data split and from 73.80% to 84.12% on the smaller split. Similarly, the proposed GAN‐based models achieved substantial performance gains, with the mean dice score increasing from 55.87% to 67.90% for the U‐Net‐based generator and from 58.11% to 78.61% for the VGG16‐based generator.
ABSTRACT Transformer‐based visual tracking has achieved strong representation capability, yet balancing tracking accuracy, temporal robustness and real‐time efficiency remains difficult under limited computational resources. The challenge is particularly pronounced in dynamic scenes, where lightweight feature extraction can weaken spatial discrimination, while inaccurate cross‐frame associations may accumulate into tracking drift. To address this problem, this paper proposes a dynamic position‐calibrated lightweight Transformer (DPCLT) that organizes efficient feature modelling, temporal representation calibration and adaptive inference within a unified tracking process. The framework first models spatial structure and channel semantics through a compact dual‐correlation mechanism, preserving discriminative information without relying on computationally expensive global attention. It then establishes a position–memory calibration process in which foreground and background prototypes are used to refine historical representations, and the calibrated temporal information is further converted into spatial guidance for current‐frame localization. This bidirectional interaction improves cross‐frame correspondence and reduces the influence of unreliable memory updates. In addition, scene complexity and target motion are incorporated into inference‐path selection, enabling the tracker to adjust its computational cost across frames. Experiments on LaSOT, GOT‐10k and TrackingNet show that DPCLT achieves a competitive accuracy–efficiency trade‐off and improves tracking stability in challenging conditions including occlusion, fast motion and background interference.
ABSTRACT This paper presents a perceptually aligned framework for color‐composition‐based image similarity by representing images as graphs and learning their structures with graph convolutional networks (GCNs). Images are first segmented into homogeneous color regions, which are converted into nodes characterized by CIELAB color features and region statistics. Spatial adjacency among regions is encoded as edges, enabling the model to capture both global color distribution and localized relationships that are difficult to represent in conventional RGB‐based convolutional neural network (CNN) approaches. The proposed method processes each graph through a shared GCN‐based feature extraction network, followed by a similarity prediction network that estimates the similarity between images. Experimental results show that the model achieves higher correspondence with perceptual similarity judgments than the existing CNN‐based method. The graph representation also provides improved stability under geometric transformations such as rotations and reflections, and facilitates the preservation of subtle yet perceptually salient accent colors. These findings demonstrate the effectiveness of integrating CIELAB color features with graph‐based modeling for holistic analysis of color composition. The proposed framework offers a promising direction for applications such as content‐based image retrieval, visual design support, and color scheme analysis, contributing to more perceptually meaningful assessments of image similarity.
ABSTRACT Although significant progress has been made in object detection, detecting objects in foggy conditions remains a challenging task. Fog reduces image clarity, thereby affecting the detection performance of the model. In order to address this challenge, we enhance the model's detection performance in foggy scenes through a task‐driven adaptive image enhancement method and a feature enhancement strategy (DuoEnhance). We first analyse image enhancement techniques and propose an adaptive image enhancement method for dehazing foggy images, leveraging physical priors and gamma correction. In addition, we introduce dilated‐aware weighted convolution, which enhances the model's feature extraction capability through a dynamic multi‐scale feature weighting strategy. To validate the generality and effectiveness of the proposed method, we further conduct extensive experiments on both YOLO‐based and DETR‐based detection frameworks. Due to the complexity of objects in foggy weather, we not only validate the effectiveness of the model on the synthetic dataset foggy cityscapes, but also validate model generalization on the real‐world task‐driven testing set. The experimental results show that the model is able to significantly improve mean average precision (mAP), mAP@0.5 and recall ( R ) with low floating point operations (FLOPs) on all two datasets, reducing the risk of missed or false object detections. Notably, the small model achieves a 1.5% improvement in mAP and a 2.5% increase in R , with only 24.6G FLOPs on the foggy cityscapes dataset.
ABSTRACT Accurate glioma segmentation from multimodal MRI is essential, yet existing methods struggle with inconsistencies, noise, and high tumour variability. To address this, we propose AIRP‑UNet, an asymmetric hybrid 2D network. Its residual‑block‑based encoder stably extracts hierarchical features and alleviates gradient vanishing; the atrous spatial pyramid pooling bottleneck captures multi‑scale context via atrous convolutions; the attention‑gated inverted residual decoder suppresses fusion redundancy, enhancing boundary adherence and small‑region consistency. Our contribution is design‑analytical: we deliberately break the encoder–decoder symmetry conventionally assumed in U‑shaped networks and report controlled evidence for the three design decisions. On BraTS2021 with patient‑level five‑fold cross‑validation, we achieve an average Dice of 90.10% ± 0.91% and HD95 of 2.45 ± 0.43 mm, demonstrating a competitive accuracy–efficiency trade‑off and serving as a reproducible reference for 2.5D/3D extensions.
Recently, data-driven methods for image deraining have achieved remarkable progress. However, the scarcity of accurately paired real-world rainy and clean images poses considerable challenges, which consequently hinders the generalization capability of existing networks to real-world rainy scenarios. To tackle these problems, we propose a variational Bayesian inference deraining network (VBDNet) that integrates model-driven and data-driven approaches within a unified variational Bayesian framework. Specifically, we introduce a latent variable grouping strategy to enhance the expressive capacity of the prior and variational posterior for modelling complex rain distributions. To maintain distributional consistency across groups, a residual distribution mechanism is designed to stabilize the Kullback-Leibler divergence during optimization. Moreover, we redesign BNet with a multi-scale feature fusion network, which enables it to better model the diverse and complex structures of rain streaks. VBDNet consists of BNet for background inference, RNet for rain streak generation with grouped latent variables and DNet as a discriminator for adversarial learning. Extensive experiments on both synthetic datasets and real-world SPA-Data demonstrate that VBDNet achieves superior deraining performance and exhibits stronger generalization capability compared with state-of-the-art methods.
Convolutional neural networks (CNNs) have achieved strong performance in medical image analysis; however, their deployment in edge computing environments remains challenging. This paper introduces an architecture-aware block-level pruning framework that progressively removes implementation-level blocks from sequential (VGG16, VGG19) and branched (ResNet50, InceptionV3) CNN architectures. Each pruning step is followed by training under one of three paradigms-training from scratch, feature extraction, or fine-tuning-and performance is evaluated through the accuracy-efficiency trade-off and complementary metrics. To exploit the diverse and architecture-dependent representations learned by the best-performing models in each regime, a stacking ensemble strategy is employed. Experiments on the PatchCamelyon dataset reveal a distinctly architecture- and regime-dependent pruning behavior: sequential models trained from scratch and fine-tuned exhibit monotonic accuracy-efficiency degradation, while under feature extraction they show non-monotonic patterns; branched architectures display non-monotonic behavior across all regimes due to their multi-path or residual block redundancy. Under moderate pruning, pruned Inception-V3 model outperforms prior state-of-the-art (SOTA) with a 5.47% accuracy improvement; under strong pruning, pruned Inception-V3 and ResNet-50 models exceed the literature with 4.37% and 5.38% accuracy enhancement, respectively. In aggressive pruning, two different pruned configuration of Inception-V3 achieve superior performance relative to SOTA with 1.21% and 10% accuracy improvements. The stacking ensembles further yield 1.78%-6.61% accuracy gains, 2%-5% AUC improvements, and 4%-6% increases in F1-score over individual CNN backbones. These findings demonstrate that the notable improvements obtained via ensemble learning reflect the strong capability of the pruned networks, which maintain excellent predictive behavior even with their compressed architectural design.
Kidney tumours (KTs) constitute a significant global health burden, ranking as the fourteenth most prevalent tumour type among men and women worldwide. Early detection is critical for reducing mortality, enabling timely preventive measures, and improving treatment outcomes; however, traditional diagnostic methods are often time-consuming, labour-intensive and reliant on specialist interpretation. Deep learning (DL) based automated detection systems offer a promising alternative by improving accuracy, reducing costs and alleviating radiologists' workload. Despite extensive research, limitations such as inadequate datasets and sub-optimal detection techniques have hindered progress. This study proposes a deep learning (DL) based framework for the automated detection and classification of KTs using computed tomography (CT) images. A dataset of 7770 images from 111 patients was preprocessed and augmented to enhance model generalization. For the task of tumour detection, six DL architectures were evaluated, with a custom CNN12 model achieving the highest performance (accuracy: 99.43%, precision: 0.99, recall: 1.00, 1-score: 0.99), outperforming several baseline and transfer learning models. For the subsequent classification of detected tumours into benign and malignant types, a custom CNN11 model attained superior results (accuracy: 97.17%, precision: 0.95, recall: 0.97, 1-score: 0.96), demonstrating robust classification capability. These results indicate a significant improvement over many existing approaches and validate the effectiveness of the proposed two-phase framework. The findings demonstrate that the proposed framework can accurately distinguish KTs from normal tissues and further classify their malignancy with high precision. This approach has the potential to serve as a decision-support tool in clinical settings, assisting radiologists by reducing diagnostic time, minimizing the risk of misdiagnosis and facilitating earlier intervention to improve patient outcomes.
ABSTRACT Nowadays, it is crucial to provide a high quality of experience (QoE) for video streaming, and this is particularly challenging under fluctuating network conditions. Adaptive bitrate (ABR) algorithm is the core mechanism to optimize QoE under variable network bandwidth. However, existing ABRs may fall into sub‐optimization without considering video quality enhancement, which is incrementally employed to improve QoE after downloading. In this work, we observe that the degree of video quality enhancement (DVQE) varies across different video chunks with different bitrates, and thus propose the DVQE sensitivity‐aware adaptive bitrate algorithm (DSABR). Specifically, DSABR quantifies the DVQE sensitivity of each chunk bitrate by jointly considering its DVQE value and the corresponding chunk size, and accordingly selects the chunk bitrate with higher DVQE sensitivity. In this way, DSABR would achieve high QoE with minimal network bandwidth occupancy. Experimental results demonstrate the performance advantage of DSABR, especially under the conditions of limited network bandwidth and videos with large average DVQE and large DVQE variance. Specifically, under the real‐world network bandwidth dataset PUFFER, DSABR improves QoE by 94.0%, 133.7%, and 26.8% compared to the heuristic algorithms BOLA, FESTIVE, and RobustMPC, respectively. Furthermore, compared to the learning‐based algorithms COMYCO and MERINA, DSABR improves the QoE by 24.4% and 20.5%.
ABSTRACT Diffusion models (DM) have revolutionized the field of image dehazing, further narrowing the gap between image quality and human perceptual preferences. In recent years, DM‐based image dehazing has attracted widespread attention and numerous works have emerged. In this survey, we comprehensively review more than 90 research works conducted from 2023 to 2026. First, we introduce the relevant background of DM and image dehazing. Next, we describe three types of classical DM as well as their improvements and optimizations. Then, we focus on recent advances in the application of DM in image dehazing and the all‐in‐one image restoration (AiOIR) (including dehazing). We also compare the performance of seven different methods on both synthetic and real‐world haze images across various scenarios. Finally, we offer unique insights into enhancing the capabilities of DM‐based image dehazing methods and possible future development directions. In summary, this survey represents the first systematic and comprehensive overview of DM‐based image dehazing, aiming to provide a valuable guide for future researchers and stimulate continued progress in this field. The paper and corresponding code link for the image dehazing and AiOIR method based on DM can be found at https://github.com/ZhuLiangyu123/Awesome‐Diffusion‐Model‐for‐Image‐Dehazing .
Accurate calibration of multi-camera systems is essential for the deployment of vision-based mobile robots and autonomous vehicles. Conventional checkerboard-based calibration methods typically require full field-of-view overlap of the calibration pattern across all cameras-a condition often difficult to satisfy in complex multi-camera configurations. To address this limitation, we propose a multi-camera calibration method based on joint optimization of adjacent camera nodes. Our approach establishes local overlapping constraints between neighboring cameras and estimates the intrinsic parameters of individual cameras along with the extrinsic parameters between adjacent cameras in a staged manner. These extrinsic parameters are then recursively transformed via intermediate nodes to achieve globally consistent parameter alignment within a unified coordinate system. Furthermore, by constructing a camera topology graph and incorporating a confidence evaluation function, our method quantitatively assesses the quality of calibration contributions from each camera. This enables dynamic selection of the optimal extrinsic parameter propagation path. The introduction of soft global consistency regularization and a covariance-aware path selection mechanism effectively mitigates cumulative errors arising from chain propagation. The proposed method successfully addresses the challenges of calibrating multi-camera systems with non-overlapping fields of view and uneven calibration quality. Experimental results demonstrate that our algorithm achieves high calibration accuracy.