The reversibility of robust reversible watermarking (RRW) strictly depends on the premise that the watermarked image has not been attacked. However, in practical applications, multimedia content often suffers from various attacks, and existing RRW technologies completely lose their reversible recovery capabilities. Therefore, this paper proposes a self-recovery robust watermarking (SRRW) algorithm, which provides a new ability to recover the watermarked and attacked image more closely to the original image while the watermark maintains strong robustness to those common signal processing operations and geometric attacks. First, the watermark information is inserted into low-order Zernike moments of a cover image by using quantization watermarking technique so that the watermark is resistant to those additive noise-like operations (like JPEG compression and Gaussian noises) and invariant to geometric transforms (like rotation and scaling). Then, the scaled difference vectors between the cover vectors and its quantized watermarked vectors are luminously and ingeniously computed as the distortion compensation information added back to the quantized watermarked vectors for restoration of the cover image. Finally, at the receiver side, based on the watermark bits extracted from the watermarked image or its distorted versions, those Zernike moments embedded data can be restored by the known scaling factor of the difference vector. Furthermore, a self-recovery image resembling the original image more closely can be reconstructed. Experimental results show that the proposed SRRW scheme can effectively restore those distorted images due to the 128-bit watermark embedding, JPEG compression with the quality factor 60 or JPEG2000 compression with the compression ratio 9, while providing strong robustness performance to various kinds of attacks. Compared with the attacked watermarked images, the restored images show PSNR improvements of 1.66 dB, 3.53 dB, and 1.72 dB under JPEG compression (Q = 90), JPEG2000 compression (R = 3), and AWGN (sigma = 0.0001), respectively.
Reversible data hiding (RDH) for JPEG images, particularly those focusing on DCT coefficient modification, has garnered significant attention in recent years. Existing methods primarily select coefficients valued $\pm 1$ for expansion embedding to avoid significant file size increases caused by modifying zero-valued DCT coefficients. However, zero-valued coefficients, which constitute the majority of DCT coefficients, are more suitable for data embedding to reduce the shift distortion. To efficiently utilize zero-valued coefficients for high-capacity embedding while controlling the file size increment, this paper introduces a novel JPEG RDH method based on ternary matrix embedding, where ternary syndrome trellis codes (STC) is employed on selected zero-valued coefficients to minimize the expansion embedding distortion, and other non-zero-valued coefficients are shifted for reversibility. Furthermore, a novel DCT coefficients measurement strategy is proposed for coefficient selection to further reduce the shift distortion. Extensive experimental validations demonstrate the superiority of the proposed method in various evaluation criteria. Notably, the proposed method achieves more than twice the embedding capacity of some state-of-the-art methods at the same PSNR while maintaining file size increment within acceptable bounds.
In this paper, we propose a camera authentication scheme to enhance access security for front-end devices in surveillance networks. The scheme leverages the Photo-Response Non-Uniformity (PRNU) pattern noise of camera sensors and combines traditional encryption techniques to strengthen system security. During registration, the camera captures images to extract PRNU, generates a compressed device fingerprint stored on the server as a root key. For authentication, the server sends a challenge sequence randomly generated from the root key to the front-end, which captures a new image to generate a root key approximation for response. To prevent attackers from extracting device fingerprints from public images, it incorporates anonymization, proposing a DWT-based PRNU anonymization algorithm. This improves PSNR by 8.08 dB and SSIM by 0.08 on average compared to previous methods. Security Analysis and Experimental results show high authentication accuracy and security, effectively resisting replay and man-in-the-middle attacks, providing a robust solution for surveillance network devices.
AI-generated content (AIGC) has become increasingly difficult to distinguish from real images, creating new challenges for media authentication. Existing detectors often rely on either convolutional networks, which focus on local patterns but lack global reasoning, or Transformers, which capture long-range context but suffer from high computational cost. Recent state space models such as Mamba provide linear time processing, yet their causal structure leads to long-range dependency decay, making them less effective for detecting forgery clues that appear across distant image regions. In this work, we propose Multi-scale Linear Local Attention (MLLA), a unified framework for AIGC detection that combines local artifact modeling with efficient global context reasoning. Our design integrates an artifact-aware tokenization (AAT) with a Linear Local Attention (LLA) block that merges depthwise convolutions, linear attention, and rotary positional embedding to overcome the limitations of both Transformers and causal state space models. By stacking LLA blocks in a multi-scale encoder, the network learns fine-grained features in shallow layers and broader semantic clues in deeper layers. Experiments on a wide range of GAN and diffusion datasets show that MLLA achieves state-of-the-art performance and strong generalization to unseen generators. The results confirm that combining local priors with efficient non-causal global modeling is a simple yet powerful direction for AIGC detection.
Transferable adversarial examples (AEs) have attracted considerable attention due to their ability to expose vulnerabilities in black-box deep neural networks (DNNs). However, achieving superior transferability for targeted attacks remains a challenge. In this article, inspired by the observation that AEs with smaller intra-class distances and larger inter-class distances tend to exhibit higher transferability, we propose a novel targeted attack based on Feature Contrastive Optimization (FCO). This attack enhances adversarial transferability by minimizing intra-class distances and maximizing inter-class distances. Specifically, we first define positive samples (belonging to the target class) and negative samples (belonging to non-target classes) that correspond to targeted AEs. Subsequently, leveraging these defined positive and negative samples, we propose two metrics—Intra-class Compactness (IC) and Inter-class Separability (IS)—to construct a novel Feature Contrastive (FC) loss. By integrating this plug-and-play FC loss into standard adversarial objectives, the generated AEs are encouraged to better align with the target class distribution while diverging from those of non-target classes. Extensive experiments on the ImageNet-compatible dataset demonstrate that our approach consistently improves targeted transferability across a broad range of DNN architectures.
External disturbances such as wind gusts and track irregularities can affect the suspension performance of high-speed maglev trains. In addition, in long-term commercial operations, multisource coupled factors such as track settlement, time-varying structural parameters, loosening of mechanical components, and degradation of controller performance can severely threaten the operational reliability, safety, and passenger comfort of the train. Therefore, monitoring abnormal states of the suspension system becomes crucial. This study first classifies potential abnormal states of the suspension system into three categories: low-frequency transient anomalies, high-frequency transient anomalies, and long-term anomalies. Different anomaly detection strategies are designed for each category to achieve comprehensive and accurate monitoring of abnormal states: the isolation forest algorithm is used for real-time monitoring of suspension data to quickly and accurately identify transient abnormal states without relying on pretraining or data labels; an improved long short-term memory (LSTM) model integrating the complete ensemble empirical mode decomposition with adaptive noise (CEEMDAN) method is developed to eliminate noise interference, reduce modeling complexity, and improve time-series prediction accuracy. Long-term abnormal states are further identified offline using steady-state metrics of the suspension system. Simulations and experimental verification demonstrate that the proposed anomaly detection method can accurately identify various abnormal states of the suspension system.
Recent advances in Deepfake generation have enabled forged audio to be synthesized alongside manipulated videos, significantly enhancing realism and posing new challenges for Deepfake detection. However, due to the inherent cross-modal gap between audio and visual information, existing multimodal methods often fail to effectively exploit complementary information across modalities. Moreover, both unimodal and multimodal representations commonly contain redundant and noisy information, which weakens discriminative cues and degrades detection performance. To address these issues, we propose Audio-Visual Information Refinement (AVIR), a robust multimodal Deepfake detection framework that explicitly reduces redundancy and noise. AVIR integrates superpixel modeling and the information bottleneck (IB) principle to learn compact and task-relevant representations. Specifically, a Superpixels-based Defects Mining Transformer (SDMT) is designed to cluster local forgery artifacts into semantically coherent regions, thereby reducing redundant and noisy visual information. To further refine both unimodal and multimodal features, an Artifacts-Emphasis Information Bottleneck (AEIB) module is proposed, which learns minimal sufficient representations by maximizing mutual information with detection labels while constraining irrelevant information from the input. In addition, AEIB is applied to multimodal representations to effectively mitigate the cross-modal gap and enhance discriminative fusion. Extensive experiments on multiple public benchmarks demonstrate that AVIR outperforms state-of-the-art unimodal and multimodal Deepfake detection methods in cross-dataset generalization scenarios.
The development of the internet has greatly facilitated the transmission of images over social networks, while also triggering serious copyright issues. Deep robust watermarking serves as a crucial technique for image copyright protection. However, the image distortions caused by Social Network Transmission Operations (SNTOs) make existing deep robust watermarking methods fragile in real-world social network scenarios. To address this, we propose a Curriculum Learning-based Deep Robust Watermarking method, called CL-DRW, to generate watermarks that can be resilient to SNTOs. Specifically, we develop a watermarking model constructed with an invertible neural network and present a multi-stage training framework based on curriculum learning to train it effectively. We incrementally introduce noise attacks based on their disruptive impact on the watermark, from weak to strong, thereby enabling our model to build robustness against SNTOs gradually. Additionally, we design an SNTOs simulation noise layer, which is built upon a transformer-based deep network and incorporates differentiable JPEG, to simulate the closed-box distortions caused by SNTOs. Extensive experiments indicate that our proposed CL-DRW outperforms state-of-the-art deep watermarking methods in terms of robustness against real-world social network transmission operations. Source code is available at https://github.com/yingshuai-zhao/CL-DRW
Deepfake technology poses a serious threat to society by synthesizing a victims facial features and attributes to carry out deception. Traditional active defense methods against deepfake attacks are typically designed for specific models, and protected images often lose their anti-forgery capability after compression or reconstruction, severely limiting their practical applicability. This paper proposes a Meta-Learning-based active defense Scheme against deep facial forgery attacks (MLPDS), which effectively safeguards facial images against diverse deepfake attacks in real-world scenarios. Our approach adopts a general paradigminjecting noise into the original image to construct a cross-model defense algorithm against deepfake attacks. Specifically, by leveraging a meta-learning strategy, we integrate perturbations generated by multiple deepfake models, enabling robust protection against a variety of forgery models. Furthermore, to maintain the high fidelity of the images, we propose a symmetric gradient quantization strategy based on the arctan function to minimize the perceptual discrepancy between the perturbed and original images. Finally, an end-to-end optimization network is employed to generate universal perturbations tailored to specific images, supported by a pixel-level error metric that constrains deviations from the original content. Since no retraining is required to protect newly encountered images, this approach significantly improves the efficiency and practicality of real-time anti-deepfake defense. Experiments show that the proposed MLPDS algorithm can effectively resist attacks from multiple forgery models, outperforming state-of-the-art defense methods and significantly reducing image distortion with an average PSNR gain of approximately 7 dB, which fully meets the practical desire for efficient and reliable deepfake defense.
Adversarial transferability enables adversarial examples to attack unseen models. State-of-the-art input transformation attacks (ITAs) enhance transferability through ensemble optimization over multiple transformed inputs. However, existing ITAs mainly focus on designing more effective transformation techniques while largely overlooking the ensemble optimization strategy. As a result, most ITAs continue to rely on a loss-level ensemble mechanism, leaving substantial room for improving transferability through better ensemble optimization. To this end, a novel Scaled Logit Ensemble (SLE) optimization framework is proposed to enhance the adversarial transferability of existing ITAs. Specifically, SLE first replaces the loss-level ensemble mechanism with a logit-level ensemble strategy that has been shown to yield superior transferability in model ensemble attacks. Furthermore, to alleviate scale variations among logits generated from different transformed inputs, a Self-Standard Deviation Scaling method is proposed to adaptively calibrate the logits before aggregation. Extensive experiments demonstrate that SLE consistently improves the transferability of multiple ITAs across diverse model architectures, suggesting that ensemble optimization constitutes a promising direction for advancing ITAs.
While person re-identification (ReID) is widely deployed, its security against sophisticated attacks remains a critical concern. Existing attacks based on pixel-level perturbations or backdoors provide only a partial view of vulnerabilities and are often impractical in realistic open-set scenarios. To overcome these limitations and ultimately strengthen robustness, we propose the Diffusion-based Semantic Camouflage Attack (DSCA), a framework that exposes vulnerabilities in the high-level semantic space. Rather than manipulating pixels, DSCA instantiates a conditional diffusion generator that subtly edits latent semantic attributes (for example, clothing color and texture) while preserving visual coherence, thereby impersonating a specified target identity at inference. This generative formulation enables operation in a zero-query, black-box setting without access to or feedback from the victim model. In the offline setting, the network is trained on a diverse set of surrogate ReID models with different backbones, including CNN and Transformer architectures, to encourage cross-model transferability. In the online setting, DSCA directly produces a camouflage image that deceives the victim system into matching the attacker with a specified target identity, achieving a successful attack without any model interaction. Extensive experiments on major ReID benchmarks validate the approach, showing high attack success rates (over 95% in our setting), strong perceptual fidelity, and evasion of advanced defenses. By exposing security gaps at the semantic level, DSCA provides a practical diagnostic tool to inform defense objectives and guide the development of more robust ReID systems.
Reversible data hiding (RDH) for JPEG images remains relatively underexplored, with key challenges lying in coefficient selection and modification strategies. Existing methods select coefficients for embedding through block-wise or frequency-band-based operations, resulting in coarse-grained decisions that constrain embedding performance. In this paper, a novel RDH scheme for JPEG images based on gap-driven histograms generation with coefficient-wise selection is proposed. First, a multi-metric weighted complexity and coefficient-wise selection approach is proposed, integrating four local feature criteria to assess each coefficient individually, enabling more precise per-coefficient selection. Then, a gap-driven adaptive multi-histogram generation strategy is introduced, leveraging gap pairs to minimize shifting distortion by segmenting histograms via bisection and avoiding modifications to high-magnitude coefficients. Experimental results confirm that the proposed method achieves improved visual quality and more efficient file size control compared to existing state-of-the-art approaches.
A digital calibration method with reduced LUT complexity is presented to suppress code-domain nonlinear errors caused by CDAC mismatch in high-resolution SAR ADCs. The method exploits the MSB-dominant and locally correlated behavior of mismatchinduced errors, so that interval-wise error estimation and correction can be performed using the most significant bits of the output code for LUT addressing. Unlike conventional MSB capacitor-weight calibration, the proposed method does not directly estimate individual capacitor errors; instead, it statistically compensates the code-domain error over MSB-indexed intervals. A segment-wise statistical training scheme is introduced for stable coefficient updating. The proposed method is evaluated by pre-layout circuitlevel simulations of a 16-bit SAR ADC designed in a 180 nm CMOS process. Under mismatch-dominant conditions, the SNDR, ENOB, and SFDR are improved from 76.24 dB, 12.37 bit, and 84.52 dB to 95.71 dB, 15.61 bit, and 107.41 dB, respectively. In addition, the required LUT size is reduced from 65536 entries to 4096 entries, corresponding to a storage reduction of approximately 93.75%. These results indicate that the proposed method can reduce LUT storage overhead while maintaining effective calibration accuracy under mismatch-dominant pre-layout simulation conditions.
Quantization-table-modification (QTM) is widely utilized in current studies of reversible data hiding (RDH) for JPEG images. However, the performance of existing QTM-based methods is far from optimal due to the uniform quantization step division and imperfect distortion modeling. In this letter, by incorporating Syndrome-Trellis-Code (STC) into QTM, a novel hybrid JPEG images RDH method is proposed. Firstly, instead of the uniformly dividing strategy conducted in previous works, by adaptively dividing the quantization steps, a hybrid embedding mechanism combining binary and ternary embedding is proposed. Then, the corresponding capacity-distortion model is established, by which the spatial domain distortion is estimated. Finally, based on the derived capacity-distortion model, for performance optimization, STC is utilized to minimize the cover modification. In this way, JPEG images RDH can be effectively conducted so that the visual quality of the marked image is well maintained. Experimental results demonstrate that the proposed method significantly outperforms some state-of-the-art works in terms of visual quality.
Reversible data hiding (RDH) combining multiple histograms modification (MHM) and matrix embedding (ME) has demonstrated competitive capacity-distortion performance. However, the exponential complexity of multi-dimensional parameter optimization restricts existing methods to a single pair of expansion bins. Consequently, the redundancy in smooth regions is underutilized, necessitating the modification of textured regions and resulting in severe distortion at high capacities. To address this issue, a framework with multiple expansion bin pairs combined with ME and MHM is proposed. First, a general capacity-distortion model is established, and the relationship of embedding rates across multiple expansion bin pairs is derived via Lagrange optimization to reduce parameter dimensionality. Furthermore, a depth-first search algorithm constrained by image complexity priors is designed, enabling rapid global optimization and the adaptive allocation of expansion bin pairs. Experimental results demonstrate that the proposed method effectively suppresses invalid modifications at high capacities, yielding superior PSNR compared to some recent works.
Current deep watermarking frameworks typically consist of an encoder, a noise layer, and a decoder (E-N-D), in which jointly optimize the encoder and decoder over a large training set. However, this learned global embedding strategy compromises across diverse images, leaving the embedding potential of individual images under-exploited and introducing an inherent amortization gap. To address this issue, a novel paradigm termed Instance-Aware Encoder Adaptation (IAEA) is proposed in this letter. Built upon the global model trained in the first stage, IAEA freezes the decoder and fine-tunes only the encoder for each cover image and the to-be-embedded watermark in the second stage. This transforms the global optimization into an instance-level refinement under practical decoding constraints, effectively narrowing the amortization gap. The proposed IAEA can exploit the embedding potential of individual images and can be integrated with existing E-N-D methods for performance enhancement, and its effectiveness is experimentally verified.
In reversible data hiding (RDH) based on prediction-error-expansion (PEE), the embedding distortion is lower for bit “0” than for bit “1”. To favor bit “0”, a binary matrix embedding (BME)-based data transformation strategy is introduced. In this strategy, the transformed data replaces the original, and embedding is performed via two-dimensional (2D) mappings. However, the mappings are restricted to at most two modification directions by BME, limiting flexibility and increasing distortion. Moreover, the embedded data distribution is altered by data transformation, rendering traditional mapping selection models inapplicable. To address the above problems, a multivariate matrix embedding (MME)-based 2D PEE framework is proposed in this paper. First, a piecewise data transformation strategy is introduced to enable embedding via generalized mappings with three modification directions. Then, the corresponding capacity-distortion model is established to optimize mapping selection. Finally, adaptive costs are refined based on mapping patterns to minimize the embedding distortion. Experimental results demonstrate that the proposed method achieves superior marked image quality over some state-of-the-art methods.
To address the difficulty in balancing invisibility and robustness against geometric and topological attacks in 3D mesh watermarking, this paper proposes a robust watermarking algorithm based on PCA pose calibration and spherical coordinate mean QIM. First, Principal Component Analysis (PCA) is adopted to calculate the principal axes of the mesh, and statistical skewness detection is combined to eliminate the sign ambiguity of the principal axis directions, thereby constructing a robust coordinate system invariant to rotation, translation, and vertex reordering. In the spherical coordinate space, mesh vertices are partitioned into multiple sectors solely based on the azimuth angle, and redundant grouping is performed on the sectors combined with key-driven pseudo-random permutation. Subsequently, the statistical mean of the vertex polar angles within each group of sectors is calculated as a stable feature, and the Quantization Index Modulation strategy is adopted to modify the mean to embed watermark information. The extraction process reuses the same coordinate system construction and sector partitioning strategies, and majority voting decisions are performed on multiple sectors within the group. Experimental results demonstrate that, the proposed method possesses strong robustness against attacks such as affine transformations (RST), Gaussian noise, Laplacian smoothing, mesh simplification, as well as 3D compression and quantization, meanwhile guarantees high visual quality.