Artificial intelligence technology based on deep learning has been widely used in key fields such as automatic driving, medical diagnosis and financial risk control. These applications also bring more and more serious security problems. In particular, as the means of attack continue to evolve, well-designed countermeasures seriously threaten the reliability of the model and the security of the system. In order to deal with this risk, defensive confrontation samples have become the core task of AI security research, playing a key role in improving the security and credibility of the model. Aiming at the problems of unclear concepts and overlapping standards in previous classification methods, this paper proposes a clearer and unified classification framework, combs and defines the contents of existing research, and solves the inconsistencies. The framework systematically divides the existing countermeasures and defense methods into three categories: detection, purification and optimization. This classification will help researchers understand the actual effects of different methods in the face of various attacks more clearly. This paper also analyzes the tradeoffs between accuracy, robustness, operational efficiency and generalization capability of various defense mechanisms, and reveals how they balance the calculation cost and actual deployment requirements. In addition, the paper points out the main challenges facing the current research, and puts forward the future research directions, including developing more efficient, adaptive, and cross modal defense methods to comprehensively improve the security of AI systems. The purpose of this review is to help researchers understand the development process of anti sample defense technology and provide a reference path for building a stable and reliable AI system.
Generator-based adversarial attack methods aim to fool deep neural networks (DNNs) by training a generator for crafting adversarial examples (AEs). However, as DNNs evolve from Convolutional Neural Networks (CNNs) to Transformers, the existing generator-based methods can hardly achieve satisfactory attack performance against different target model architectures in semi-whitebox attack scenarios. In addition, the generated AEs are susceptible to various distortions (especially for JPEG compression with low quality factors), which deteriorate the attack ability and increase the unreliability. To address these issues, we propose a dual-branch guided generative model called Heterogeneous Encoding Network (HENet) to form a robust generator-based adversarial attack framework. Specifically, our HENet introduces an Adaptive Feature Fusion Module (AFFM) to solve the dimensions and representativeness contradictions between CNNs and Transformers, which steers the perturbation generation based on a richer latent space and achieves better general attack ability. To further improve the robustness against JPEG compression, we design and integrate a Dynamic Differentiable JPEG Simulator (DDJS), which introduces an adaptive quantization mask to determine the flow of the gradient backpropagation in each frequency position. Extensive experiments prove the proposed method achieves a better attack success rate, lower perturbation magnitude, and higher robustness for various target network architectures under compressed, distorted, and lossless scenarios. Our codes will be made publicly available.
JPEG image forensics has become a critical area of research due to the increasing prevalence of compressed images and their vulnerability to tampering. This paper provides a comprehensive survey of three major tasks within JPEG image forensics: JPEG Compression Detection (JCD), JPEG Quantization Step Estimation (JQE), and the Application of JPEG Features (AJF). We explore the key features and techniques for detecting compression traces and estimating quantization parameters, emphasizing their applicability in both forensic analysis and anti-forensic strategies. JCD focuses on identifying JPEG compression artifacts, while JQE estimates the specific quantization steps used during compression. AJF, with its broad scope, supports diverse applications such as tampering detection, image recovery, and anti-forensics. We also examine the interrelationships between these tasks and discuss the challenges that hinder the field’s progress, including issues related to multi-compression scenarios, quality factor dependence, and the growing need for robust, generalizable methods. Finally, we propose benchmarks for evaluating the robustness of forensic models against adversarial attacks and multi-format compression schemes, while highlighting emerging trends and future research directions, including the integration of multi-modal information and advancements in deep learning-based solutions.
JPEG compression, a ubiquitous operation on social media platforms, disrupts the embedding distribution of stego images, impairs the extraction of secret information, and increases their detectability. Existing robust image steganographic methods typically approximate JPEG compression using differentiable noise layers, which enable gradient propagation but model compression mainly as a one-way perturbation. To address these issues, we propose a JPEG Compression Simulation Invertible Neural Network (JCS-INN). JCS-INN employs invertible encoding layers to simulate JPEG compression and transforms the lost high-frequency information into auxiliary latent variables constrained by an isotropic Gaussian distribution, thereby enabling high-quality reconstruction through invertible decoding. Based on JCS-INN, we design a dual-domain security-enhanced steganographic network (DDSESNet) that integrates both spatial-domain and JPEG-domain steganalyzers. Through joint adversarial training, the generator achieves end-to-end compression-aware optimization and learns to produce compression-resistant stego images. Furthermore, JCS-INN can serve as a plug-and-play module to replace the noise layers in other steganographic frameworks, improving both their robustness against JPEG compression and training efficiency. Extensive experimental results demonstrate that the proposed method outperforms other state-of-the-art approaches in both security and the quality of extracted secret images under compressed transmission environments.
Synthetic aperture radar (SAR) imagery plays a vital role in maritime surveillance, yet it faces severe challenges such as speckle noise, cluttered backgrounds, and the small scale and low contrast of ship targets. To overcome these issues, we propose MBE-Net, a lightweight and high-accuracy SAR ship detection network featuring multibranch edge-guided feature enhancement. First, the edge-enhanced wavelet backbone (EEWB) is designed to enhance edge information and suppress detail loss by integrating multibranch edge guidance and wavelet pooling, thereby improving small-ship perception in complex scenes. Second, the lightweight decoupled detail-enhanced head (LD-DEH) decouples classification and regression tasks while embedding a detail-enhanced convolution module to recover high-resolution features, significantly boosting localization precision under strong interference. Finally, the adaptive contextual focus feature pyramid network (ACF-FPN) introduces context-guided fusion to improve multiscale feature alignment and robustness, especially in low-contrast and cluttered environments. Extensive experiments on the HRSID and SSDD datasets validate that MBE-Net achieves excellent detection accuracy with minimal computational cost, making it highly suitable for real-time SAR ship detection.
The forensic examination of AIGC(Artificial Intelligence Generated Content) faces poses a contemporary challenge within the realm of color image forensics. A myriad of artificially generated faces by AIGC encompasses both global and local manipulations. While there has been noteworthy progress in the forensic scrutiny of fake faces, current research primarily focuses on the isolated detection of globally and locally manipulated fake faces, thus lacking a universally effective detection methodology. To address this limitation, we propose a sophisticated forensic model that incorporates a dual-stream framework comprising quaternion RGB and PRNU(Photo Response Non-Uniformity). The PRNU stream extracts the “camera fingerprint” feature by discerning the non-uniform response of the image sensor under varying lighting conditions, thereby encapsulating the overall distribution characteristics of globally manipulated faces. The quaternion RGB stream leverages the inherent nonlinear properties of quaternions and their informative representation capabilities to accurately describe changes in image color, background, and spatial structure, facilitating the meticulous capture of nuanced local distinctions between locally manipulated faces and real faces. Ultimately, we integrate the two streams to establish the exchange of feature information between PRNU and quaternion RGB streams. This strategic integration fully exploits the complementarity between two streams to amalgamate local and global features effectively. Experimental results obtained from diverse datasets underscore the advantages of our method in terms of accuracy, achieving a detection accuracy of 96.81%.
Medical image segmentation, i.e., labeling structures of interest in medical images, is crucial for disease diagnosis and treatment in radiology. In reversible data hiding in medical images (RDHMI), segmentation consists of only two regions: the focal and nonfocal regions. The focal region mainly contains information for diagnosis, while the nonfocal region serves as the monochrome background. The current traditional segmentation methods utilized in RDHMI are inaccurate for complex medical images, and manual segmentation is time-consuming, poorly reproducible, and operator-dependent. Implementing state-of-the-art deep learning (DL) models will facilitate key benefits, but the lack of domain-specific labels for existing medical datasets makes it impossible. To address this problem, this study provides labels of existing medical datasets based on a hybrid segmentation approach to facilitate the implementation of DL segmentation models in this domain. First, an initial segmentation based on a 3 x 3 kernel is performed to analyze identified contour pixels before classifying pixels into focal and nonfocal regions. Then, several human expert raters evaluate and classify the generated labels into accurate and inaccurate labels. The inaccurate labels undergo manual segmentation by medical practitioners and are scored based on a hierarchical voting scheme before being assigned to the proposed dataset. To ensure reliability and integrity in the proposed dataset, we evaluate the accurate automated labels with manually segmented labels by medical practitioners using five assessment metrics: dice coefficient, Jaccard index, precision, recall, and accuracy. The experimental results show labels in the proposed dataset are consistent with the subjective judgment of human experts, with an average accuracy score of 94% and dice coefficient scores between 90%- 99%. The study further proposes a ResNet-UNet with concatenated spatial and channel squeeze and excitation (scSE) architecture for semantic segmentation to validate and illustrate the usefulness of the proposed dataset. The results demonstrate the superior performance of the proposed architecture in accurately separating the focal and nonfocal regions compared to state-of-the-art architectures. Dataset information is released under the following URL: https:// www.kaggle.com/lordamoah/datasets (accessed on 31 March 2025).
It is crucial to detect double JPEG compression images in digital image forensics. When detecting recompressed images, most detection methods assume that the quantization table in the JPEG header is safe. The method fails once the quantization table in the header file is tampered with. Inspired by this phenomenon, this paper proposes a double JPEG compression anti-detection method based on the generative adversarial network (GAN) by modifying the quantization table of JPEG header files. The proposed method draws on the structure of GAN to modify the quantization table by gradient descent. Also, our proposed method introduces adversarial loss to determine the direction of the modification so that the modified quantization table can be used for cheat detection methods. The proposed method achieves the aim of anti-detection and only needs to replace the original quantization table after the net training. Experiments show that the proposed method has a high anti-detection rate and generates images with high visual quality.
Adversarial examples have been shown to deceive Deep Neural Networks (DNNs), raising widespread concerns about this security threat. More seriously, as different DNN models share critical features, feature-level attacks can generate transferable adversarial examples, thereby deceiving black-box models in real-world scenarios. Nevertheless, we have theoretically discovered the principle behind the limited transferability of existing feature-level attacks: Their attack effectiveness is essentially equivalent to perturbing features in one step along the direction of feature importance in the feature space, despite performing multiple perturbations in the pixel space. This finding indicates that existing feature-level attacks are inefficient in disrupting features through multiple pixel-space perturbations. To address this problem, we propose a P2FA that efficiently perturbs features multiple times. Specifically, we directly shift the perturbed space from pixel to feature space. Then, we perturb the features multiple times rather than just once in the feature space with the guidance of feature importance to enhance the efficiency of disrupting critical shared features. Finally, we invert the perturbed features to the pixels to generate more transferable adversarial examples. Numerous experimental results strongly demonstrate the superior transferability of P2FA over State-Of-The-Art (SOTA) attacks.
Face Recognition (FR) systems, while widely used across various sectors, are vulnerable to adversarial attacks, particularly those based on deep neural networks. Despite existing efforts to enhance the robustness of FR models, they still face the risk of secondary adversarial attacks. To address this, we propose a novel approach employing “strengthened face” with preemptive defensive perturbations. Strengthened face ensures original recognition accuracy while safeguarding FR systems against secondary attacks. In the white-box scenario, the strengthened face utilizes gradient-based and optimization-based methods to minimize feature representation differences between face pairs. For the black-box scenario, we propose Shielded Gradient Sign Descent (SGSD) to optimize the gradient update direction of strengthened faces, ensuring the transferability and effectiveness against unknown adversarial attacks. Experimental results demonstrate the efficacy of strengthened faces in defending against adversarial faces without compromising the performance of FR models or face image visual quality. Moreover, SGSD outperforms conventional methods, achieving an average performance improvement of 4% in transferability across different attack intensities.
Image-to-image (I2I) translation has emerged as a valuable tool for privacy protection in the digital age, offering effective ways to safeguard portrait rights in cyberspace. In addition, I2I translation is applied in real-world tasks such as image synthesis, super-resolution, virtual fitting, and virtual live streaming. Traditional I2I translation models demonstrate strong performance when handling similar datasets. However, when the domain distance between two datasets is large, translation quality may degrade significantly due to notable differences in image shape and edges. To address this issue, we propose Long-Domain Search GAN (LDSGAN), an unsupervised I2I translation network that employs a GAN structure as its backbone, incorporating a novel Real-Time Routing Search (RTRS) module and Sketch Loss. Specifically, RTRS aids in expanding the search space within the target domain, aligning feature projection with images closest to the optimization target. Additionally, Sketch Loss retains human visual similarity during long-domain distance translation. Experimental results indicate that LDSGAN surpasses existing I2I translation models in both image quality and semantic similarity between input and generated images, as reflected by its mean FID and LPIPS scores of 31.509 and 0.581, respectively.
Deep neural networks have achieved excellent performance in various research and applications, but they have proven to be susceptible to adversarial examples. Generating adversarial examples can help identify the vulnerability of the deep neural networks and further enhance the robustness and reliability of these models. However, the existing adversarial attacks can hardly achieve the balance between robustness and imperceptibility, which is not trustworthy in social networks. To solve these problems, we propose adaptive adversarial perturbation (AAP) to improve the universal robustness of the adversarial examples while ensuring imperceptibility. To optimize the imperceptibility of the perturbation, we design a noise visibility function (NVF) to reflect the features of the original images based on the human visual system (HVS). By further calculating a coefficient matrix based on the NVF, the perturbation intensity of different pixels can be adjusted dynamically to improve the robustness. The experimental results prove that the proposed method alleviates the trade-off between robustness and imperceptibility, and outperforms existing attack methods in both one-step and iterative ways. Our method makes the adversarial attack more reliable and applicable in social networks.
The quantization step is a crucial parameter in JPEG compression, that can reveal the compression history of a JPEG image. Estimating the quantization steps for single compressed and recompressed images is attracting considerable interest in the field of image forensics and steganalysis. Several effective methods have been proposed, but the performance of these methods still needs to be improved on small-sized and low-quality images. To solve the above problems, feature enrichment is performed on images in the frequency domain, resulting in clustering discrete cosine transform (DCT) coefficients of the same frequency. Then, we construct a hierarchical connection within the residual blocks of the network to represent multi-scale features, enabling the network to learn deep features of the image. At the same time, we use multiple small-sized convolution kernels instead of one large-sized convolution kernel to minimize the impact of block artifacts. Based on the above two ideas, we construct a network model, Res2Net-C, to discover information about the quantization steps in the frequency domain. The integration of multi-channel information of color images is achieved by multi-channel convolution, and the quantization steps of the chrominance and luminance channels of the color images are estimated. The experimental results show that the accuracy of the proposed method for estimating the quantization steps is 29.97% better than that of the existing algorithm with a single compressed dataset and 4.87% better than that of the existing algorithm with a recompressed image dataset. In addition, the method has good performance with mixed datasets that contain both single compressed and recompressed images.
Invisible watermarking can be used as an important tool for copyright certification in the Metaverse. However, with the advent of deep learning, Deep Neural Networks (DNNs) have posed new threats to this technique. For example, artificially trained DNNs can perform unauthorized content analysis and achieve illegal access to protected images. Furthermore, some specially crafted DNNs may even erase invisible watermarks embedded within the protected images, which eventually leads to the collapse of this protection and certification mechanism. To address these issues, inspired by the adversarial attack, we introduce Invisible Adversarial Watermarking (IAW), a novel security mechanism to enhance the copyright protection efficacy of watermarks. Specifically, we design an Adversarial Watermarking Fusion Model (AWFM) to efficiently generate Invisible Adversarial Watermark Images (IAWIs). By modeling the embedding of watermarks and adversarial perturbations as a unified task, the generated IAWIs can effectively defend against unauthorized identification, access, and erase via DNNs and identify the ownership by extracting the embedded watermark. Experimental results show that the proposed IAW presents superior extraction accuracy, attack ability, and robustness on different DNNs, and the protected images maintain good visual quality, which ensures its effectiveness as an image protection mechanism.
With the continuous development of the internet, an increasing number of images have been tampered with on the network, accompanied by a growing range of techniques to cover up tampering traces.However, most current detection models neglect the impact of image post-processing on tamper detection algorithms, limiting their real-life applications.To address these issues, a general image tampering location model based on enhanced samples and the pseudo-twin network was proposed.The pseudo-twin network enabled the model to learn tampering features in real images.On one hand, by applying convolution constraints, the image content was suppressed, allowing the model to focus more on residual trace information of tampering.The two-branch structure of the network facilitated the comprehensive utilization of image feature information.By utilizing enhanced samples, the model could dynamically generate the most crucial pictures for learning tamper types, enabling targeted training of the model.This approach ensured that the model converged in all directions, ultimately obtaining the global optimal model.The idea of data enhancement was employed to automatically generate abundant tampered images and corresponding masks, effectively resolving the limited tampering dataset issue.Extensive experiments were conducted on four datasets, demonstrating the feasibility and effectiveness of the proposed model in pixel-level tamper detection.Particularly on the Columbia dataset, the algorithm achieves a 33.5% increase in F1 score and a 23.3% increase in MCC score.These results indicate that the proposed model harnesses the advantages of deep learning models and significantly improves the effectiveness of tamper location detection.
Traditional encrypted communication technologies have been easily detected and have struggled to meet the needs of secure communication. Steganography, capable of hiding information by modifying the carrier, has been utilized to realize covert communication. However, the potential for steganography to be employed in illegal acts has directed increasing attention towards steganalysis for the detection of steganography, thereby bestowing great research significance upon it. Deep learning has yielded numerous research achievements in fields such as computer vision, pattern recognition, and natural language processing, introducing new opportunities and challenges to steganalysis. These advancements have propelled the generation of new ideas and methods in steganalysis. Currently, color images constitute the mainstream carrier in the process of internet transmission. Nevertheless, existing steganalysis features for color images primarily rely on manual design and often treat the color image as three independent grayscale images, without fully considering the internal relationships between the three color channels, thus necessitating an improvement in detection capabilities for encrypted images. The application of deep learning in the field of color image steganalysis remains in its preliminary stage. The concepts, classifications, and research significance of steganography and steganalysis were introduced, along with an outline of their current research status. Several key techniques for the steganalysis of color images were introduced, compared, summarized, and their development trends were analyzed.
Existing deep learning-based steganography detection methods utilize convolution to automatically capture and learn steganographic features, yielding higher detection efficiency compared to manually designed steganography detection methods. Detection methods based on convolutional neural network frameworks can extract global features by increasing the network's depth and width. These frameworks are not highly sensitive to global features and can lead to significant resource consumption. This manuscript proposes a lightweight steganography detection method based on multiple residual structures and Transformer (ResFormer). A multi-residuals block based on channel rearrangement is designed in the preprocessing layer. Multiple residuals are used to enrich the residual features and channel shuffle is used to enhance the feature representation capability. A lightweight convolutional and Transformer feature extraction backbone is constructed, which reduces the computational and parameter complexity of the network by employing depth-wise separable convolutions. This backbone integrates local and global image features through the fusion of convolutional layers and Transformer, enhancing the network's ability to learn global features and effectively enriching feature diversity. An effective weighted loss function is introduced for learning both local and global features, BiasLoss loss function is used to give full play to the role of feature diversity in classification, and cross-entropy loss function and contrast loss function are organically combined to enhance the expression ability of features. Based on BossBase-1.01, BOWS2 and ALASKA#2, extensive experiments are conducted on the stego images generated by spatial and JPEG domain adaptive steganographic algorithms, employing both classical and state-of-the-art steganalysis techniques. The experimental results demonstrate that compared to the SRM, SRNet, SiaStegNet, CSANet, LWENet, and SiaIRNet methods, the proposed ResFormer method achieves the highest reduction in the parameter, up to 91.82%. It achieves the highest improvement in detection accuracy, up to 5.10%. Compared to the SRNet and EWNet methods, the proposed ResFormer method achieves an improvement in detection accuracy for the J-UNIWARD algorithm by 5.78% and 6.24%, respectively.
Research on adversarial attacks mainly focuses on reducing the amplitude of disturbances, increasing the success rate of attacks, and improving attack efficiency. However, the adversarial examples are all in the form of raw (float matrix or lossless compressed pictures). Due to the difficulty of storage and transmission caused by a large amount of redundant information, raw pictures rarely exist in the real world. In this work, a convolution operation is used to simulate the JPEG decompression process, forming a decompression module independent of the classification model to obtain the gradient information about the JPEG stream. We establish a process for generating JPEG adversarial examples. To improve the decoded image's visual quality and reduce the perturbation amplitude, we select the complex texture area according to the human visual model and select only the most effective points in each DCT block to embed adversarial noise. A large number of tests on the ImageNet dataset show the effectiveness of the proposed method.
Detection of aligned double Joint Photographic Experts Group (JPEG) compressed images is a crucial area of research within the field of digital image forensics. The detection tasks for aligned double JPEG compression can be categorized into two sub-tasks, namely detecting double JPEG images with the same quantization matrix (DJSQM) or double JPEG images with different quantization matrices (DJDQM). Existing methods for one of these sub-tasks may not be effective for the other. To address this issue, a novel approach is proposed by recompressing both DJDQM and DJSQM using modified quantization coefficients. The perturbation in the recompression process results in a perturbed error image, which is valid for both DJDQM and DJSQM. Subsequently, the relative change rate is used to combine the perturbed error image, the original error image, and the quantization error to derive the interference error and the interference quantization error. The interference error and interference quantization error further expand the difference between single and double compressed images by preserving the general validity of the original image information. Furthermore, the recompression process of DJDQM and DJSQM results in the conversion of truncation and rounding errors at the pixel level, which can be represented by the pixel state map. The pixel state map characterizes the differing transformation relationships between single and double compressed images and provides additional valid features, thereby enhancing the performance of the proposed method. The empirical results demonstrate that the proposed method outperforms existing methods on detecting aligned double JPEG compressed images.
Detection of color images that have undergone double compression is a critical aspect of digital image forensics. Despite the existence of various methods capable of detecting double Joint Photographic Experts Group (JPEG) compression, they are unable to address the issue of mixed double compression resulting from the use of different compression standards. In particular, the implementation of Joint Photographic Experts Group 2000 (JPEG2000) as the secondary compression standard can result in a decline or complete loss of performance in existing methods. To tackle this challenge of JPEG+JPEG2000 compression, a detection method based on quaternion convolutional neural networks (QCNN) is proposed. The QCNN processes the data as a quaternion, transforming the components of a traditional convolutional neural network (CNN) into a quaternion representation. The relationships between the color channels of the image are preserved, and the utilization of color information is optimized. Additionally, the method includes a feature conversion module that converts the extracted features into quaternion statistical features, thereby amplifying the evidence of double compression. Experimental results indicate that the proposed QCNN-based method improves, on average, by 27% compared to existing methods in the detection of JPEG+JPEG2000 compression.