To address the growing demand for digital image security, this paper introduces AMG, a U-Net-based robust watermarking framework designed for screen-shooting resilience and watermark imperceptibility. The framework fuses channel and spatial attention to adaptively enhance key region features during encoding and decoding, achieving precise watermark embedding with visual fidelity. Guided by a Triple Mask of depth, edge, and gradient components, a specialized loss function strategically embeds the watermark into structurally-rich yet visually non-salient regions. Experiments show AMG surpasses state-of-the-art methods in extraction accuracy, robustness, and imperceptibility, offering a systematic solution to balance watermark security and visual quality.
To defend against 3D adversarial face presentation attacks (PAs), this paper proposes an initiative defense scheme to proactively prevent face images from being misused by malicious attackers for crafting adversarial presentation attack instruments (PAIs). The proposed scheme is motivated by theoretically investigating the impostor attack principle behind 3D adversarial PAs, and its core idea is to interfere with the facial biometric fitting process of such attacks by injecting proactive defense components into face images beforehand. To learn the defense components, a hierarchical loss function is constructed based on attack objective disruption and fitting direction disruption, and it is able to well supervise the generation of the defense components and enables them to disturb the ultimate identity impersonation goals and the intermediate optimization orientations of 3D adversarial PAs. To analyze the performance of the proposed scheme, experiments are carried out with four different PAIs crafted by the state-of-the-art 3D adversarial PAs. The results indicate that it can achieve high defense success rates for the 3D adversarial PAIs based on eyes, eyes and nose, respirator, and natural style changes while preserving reasonable facial fidelity, and it exhibits generalization ability for unseen face recognition models, JPEG compression, resizing, and potential anti-defense approaches to a certain extent. Furthermore, compared with the existing proactive defense methods of face images, it obtains better performance when countering the 3D adversarial face PAs.
In recent years, video steganography technique based on H.265/HEVC has received widespread attention. Typically, video steganography selects various syntax elements during the encoding process as carriers, and utilizing Prediction Unit (PU) as carrier is currently one of the most significant research directions. However, due to the limited number of PU types, such algorithms often suffer from insufficient capacity and visual quality. To alleviate the aforementioned issues, this paper proposes an H.265/HEVC video steganography algorithm that utilizes polygon encoding and Improved Deep Learnable Similarity Network Filter (IDLSNF). Firstly, we design a new polygon encoding rule, which maps different integers into several polygons. Secondly, we propose a novel steganography method based on polygon encoding and PU partition mode. This method selects the PUs of 8 & times;8 and 16 & times;16 coding unit in P-frames as carriers and hides the secret message by modifying the partition mode of two adjacent PUs. Due to the ability of polygon encoding to represent more information within a small range, it increases capacity with low steganographic distortion. Thirdly, we further propose a filter by improving DLSN, which enhances the visual quality of the entire stego video by processing I-frames. Extensive experimental results show that the video steganography algorithm proposed in this paper achieves higher capacity and superior visual quality compared to current State-of-the-Art methods. Meanwhile, our algorithm can also obtain good BRI and anti-steganalysis performance. This method has promising application prospects in the field of video covert communication.
Existing adversarial attacks for face anti-spoofing predominantly assume that the model parameters are fully known, and often overlook the transferability of adversarial examples across different models and domains. Furthermore, they typically target relatively homogeneous architectures, relying primarily on basic deep learning models for anti-spoofing. To address these limitations, this paper proposes an illumination-based input transformation method for generating adversarial attacks. A liveness ablation module is introduced to suppress liveness-related cues in the input image prior to attack generation, thereby enhancing the adversarial strength of the crafted examples. Additionally, a random illumination transformation strategy is employed to increase domain divergence by altering illumination factors, which enriches the diversity of input samples and boosts the transferability of adversarial examples across different models and settings. Extensive experiments conducted on two public datasets demonstrate that the proposed method outperforms existing approaches in terms of physical-world transferability. Moreover, the liveness ablation module can be integrated with other attack strategies to furture improve their adversarial effectiveness.
To address the limitations of current 3D mesh watermarking in robustness and imperceptibility, this paper proposes a deep watermarking based on a geometric-weighted aggregation mechanism. The message encoder and decoder networks are first improved to enable the effective embedding of 16-bit binary watermark information. An attack simulation module is then introduced to enhance the decoder's robustness against various distortions. Additionally, an adversarial discriminator is incorporated to guide the encoder in optimizing the embedding strategy, thereby minimizing geometric distortion. Furthermore, a cross-resolution strategy is developed to enable training on low-resolution meshes and perform watermark embedding and extraction on high-resolution meshes. Experimental results demonstrate that it outperforms the existing mainstream approaches in terms of extraction accuracy, geometric fidelity, and imperceptibility.
Existing 3D mesh watermarking methods often struggle to achieve high embedding capacity while remaining robust against both geometric and topological attacks. To address this problem, this paper proposes a blind watermarking method based on spherical-coordinate statistical features and frequency-domain quantization index modulation (QIM). Specifically, it first constructs a spherical coordinate system from the centroid of the 3D model and use sthe statistical distribution of radial distances to form a feature domain that is naturally invariant to rigid transformations, namely rotation, scaling, and translation (RST). To reduce the dependence on mesh topology, it further introduces a stable binning strategy based on the numerical ordering of radial distances. This strategy removes the sensitivity of watermark extraction to the original vertex indexing order and therefore provides robustness against vertex-reordering attacks. On this basis, trimmed means within the bins are used to generate a statistical feature sequence, which is then transformed by the discrete cosine transform (DCT) for energy compaction. QIM is finally applied to embed high-dimensional binary watermark information into the mid-to-high frequency DCT coefficients. Experimental results show that the proposed method can reliably embed a 256-bit watermark while maintaining good visual quality. It also remains robust against a wide range of attacks, including rotation, scaling, translation, additive noise, mesh smoothing, mesh simplification, and vertex reordering.
Face recognition models are vulnerable to spoofing of adversarial patches in the physical world. Attackers can enable face recognition models to make false identity judgments by simply pasting a sticker with a special pattern on the face. However, existing attacks lack the ability to transfer to black-box models, and the improvement of transferability is mainly focused on adversarial perturbations based on the p-norm. To further improve the attack performance and transferability, a highly transferable face recognition adversarial patches generation method named as AdvDiffusion is proposed. It first determines the region for adversarial patches generation based on facial gradient maps, and then an image is reconstructed to generate an adversarial patch by adding noise and denoising it with a pre-trained diffusion model. In the denoising, an adversarial loss is used to fine-tune the model and control the image to generate an adversarial patch with spoofing capability. Experiments and analysis show that the adversarial patches generated by the proposed mehtod have good adversarial attack capability on black-box face recognition models in both digital and physical domains, and also have better robustness under the changes of a complex physical environment compared with some state-of-the-art methods. It has great potential application for black-box attacks in the physical domain.
To reduce embedding distortion during the verification of 3D models and enhance the precision of tamper detection, a semi-fragile reversible watermarking is proposed for 3D models based on quantization interval division modulation. Initially, 3D models are transformed from Cartesian to spherical coordinates. Subsequently, watermarks derived from vertex and one-ring neighborhood surface information are generated in two parts to improve tamper detection precision. The watermark embedding leverages quantization interval division modulation to reduce embedding distortion. Experimental results and analysis reveal that this approach exhibits reduced embedding distortion and enhanced tamper detection precision compared with some existing schemes, and it demonstrates potential of application in 3D model integrity verification.
To address the difficulty in balancing invisibility and robustness against geometric and topological attacks in 3D mesh watermarking, this paper proposes a robust watermarking algorithm based on PCA pose calibration and spherical coordinate mean QIM. First, Principal Component Analysis (PCA) is adopted to calculate the principal axes of the mesh, and statistical skewness detection is combined to eliminate the sign ambiguity of the principal axis directions, thereby constructing a robust coordinate system invariant to rotation, translation, and vertex reordering. In the spherical coordinate space, mesh vertices are partitioned into multiple sectors solely based on the azimuth angle, and redundant grouping is performed on the sectors combined with key-driven pseudo-random permutation. Subsequently, the statistical mean of the vertex polar angles within each group of sectors is calculated as a stable feature, and the Quantization Index Modulation strategy is adopted to modify the mean to embed watermark information. The extraction process reuses the same coordinate system construction and sector partitioning strategies, and majority voting decisions are performed on multiple sectors within the group. Experimental results demonstrate that, the proposed method possesses strong robustness against attacks such as affine transformations (RST), Gaussian noise, Laplacian smoothing, mesh simplification, as well as 3D compression and quantization, meanwhile guarantees high visual quality.
This paper proposes an innovative bearing fault diagnosis scheme that integrates an enhanced complete ensemble empirical mode decomposition with adaptive noise technique and a one-dimensional convolutional neural network for accurate fault feature extraction and identification in antifriction bearings. Focusing on acoustic emission signals from bearings, the proposed method first separates the raw acoustic signal into multiple intrinsic modality components, and then computes the Pearson correlation coefficients among all of them and the raw signal, thereafter reconstructs the signal based on the correlation coefficients. The reconstructed signal serves as input for training the one-dimensional convolutional neural network, which achieves an exceptional fault classification accuracy of 99.8 % on the test dataset. Comprehensive experimental results validate the proposed method's superior performance in detecting various bearing defects, demonstrating significant improvements in diagnostic precision compared to conventional approaches.
Face morphing attacks have emerged as a significant security threat, compromising the reliability of facial recognition systems. Despite extensive research on morphing detection, limited attention has been given to restoring accomplice face images, which is critical for forensic applications. This study aims to address this gap by proposing a novel face de-morphing (FD) method based on identity feature transfer for restoring accomplice face images. The method encodes facial attribute and identity features separately and employs cross-attention mechanisms to extract identity features from morphed faces relative to reference images. This process isolates and enhances the accomplice's identity features. Additionally, inverse linear interpolation is applied to transfer identity features to attribute features, further refining the restoration process. The enhanced identity features are then integrated with the StyleGAN generator to reconstruct high-quality accomplice facial images. Experimental evaluations on two morphed face datasets demonstrate the effectiveness of the proposed approach, improving the average restoration accuracy by at least 5% compared with other methods. These findings highlight the potential of this approach for advancing forensic and security applications.
With the rapid development of Deepfake technology, social security is facing great challenges. Although numerous Deepfake detection algorithms based on traditional CNN frameworks perform well on specific datasets, they still suffer from overfitting due to an over-reliance on localized artifact information. This limitation leads to degraded detection performance across diverse datasets. To address this issue, this study proposes a dual-branch fusion network called LGDF-Net. LGDF-Net uses a dual-branch structure to process the local artifact features and global texture features generated by Deepfake separately, preserving their unique characteristics. Specifically, the local compression branch utilizes a specially designed local compression module (LCM) that allows the network to focus more accurately on key regions of localized artifacts in Deepfake faces. The global expansion branch enhances the analysis of the global facial context through a global expansion module (GEM), which captures image context information and subtle texture features more comprehensively. Additionally, the proposed multi-scale feature extraction module (MSFE) delves into image features at various scales, enriching the extraction of detailed information. Finally, the multi-level feature fusion strategy (MLFF) improves the integration of local and global features through multiple layers, enabling the network to learn the intrinsic connections between these two types of features. A series of experimental validations demonstrate that the proposed scheme outperforms many existing detection networks in terms of accuracy and generalization ability.
To address the challenges faced by existing Deepfake detection methods in effectively highlighting and distinguishing forgery details, a dual-branch network is proposed, comprising a high-frequency feature enhancement branch and a high-frequency feature suppression branch. By integrating high-frequency features into the inputs of both branches, the separation of frequency-domain features from spatial ones is avoided during training, thereby allowing a more comprehensive information set to be retained. The architecture is structured into three hierarchical stages—initial, intermediate, and advanced—to progressively enable feature extraction and refinement. Subtle forgery traces that might be overlooked by conventional methods are captured through differential features, which are obtained by subtracting one branch’s features from the other. Additionally, intermediate texture difference attention fusion modules and advanced texture difference attention fusion modules are introduced to strategically aggregate multi-stage features. The superior performance of the proposed model is demonstrated through extensive experiments, showing significant advancements in Deepfake detection tasks.
With the rapid development of the metaverse, massive amounts of 3D data are created and outsourced in the cloud, and ciphertext policy attribute-based encryption (CP-ABE) is widely used in fine-grained access control to achieve secure outsourced data sharing. However, the prominent security risk is due to the fact that authorized data users may later become traitors and illegally redistribute the 3D models to the public. To protect the rights of the creator, a traitor tracing and access control method for encrypted 3D models is proposed using CP-ABE and fair watermark to meet the security needs in the metaverse. First, a commutative watermark/encryption method based on the orthogonal operation domain is designed, and the 3D model is encrypted by CP-ABE. Then, a fair watermark protocol protects the rights of the parties. Finally, the blockchain acts as a trusted third party and records the authentication information for traitor tracing. The experimental results demonstrate the feasibility and safety of the proposed method.
To address the risk of image piracy caused by screen photography, this paper proposes an end-to-end robust watermarking scheme designed to resist screen-shooting attacks. It comprises an encoder, a noise layer, and a decoder. Specifically, the encoder and decoder are equipped with the DMCB structure, which combines dilated and standard convolutions to effectively enlarge the receptive field and enable the extraction of richer image features. Moreover, it selects optimal watermark embedding regions to ensure high imperceptibility while maintaining reliable extractability after screen-shooting. To further optimize the training process, a dynamic learning rate adjustment strategy is introduced to adaptively modify the learning rate based on a predefined schedule. This accelerates convergence, avoids local minima, and improves both the stability and accuracy of watermark extraction. Experimental results demonstrate its strong robustness under various shooting distances and angles, and the visual quality of images is preserved.
Face morphing attacks pose a significant threat to modern facial recognition systems by fusing two facial images into a single morphed image. This can deceive biometric systems and lead to inaccurate identifications. Despite various detection methods developed to counter these attacks, restoring the original facial image of the accomplice from the morphed image-known as face de-morphing-remains a substantial challenge. In this paper, we propose a StyleGAN-based face de-morphing network to recover the facial images of the accomplice. Our method utilizes the pre-trained StyleGAN model to encode facial images into the semantic latent space, applies a specially designed lightweight identity feature separation network to obtain high-quality semantic latent encodings of the accomplice, and then employs another pre-trained StyleGAN to generate high-quality restored images. Experimental results demonstrate that our approach significantly improves restoration accuracy compared to existing facial de-morphing methods while maintaining an efficient and lightweight identity separation network.
The currently used face recognition systems have been proven to be seriously threatened by face morphing attacks. To counter this challenge, several detection methods have been proposed, with some focusing on restoring the accomplice's facial image. However, these methods are not yet mature enough to achieve high-quality restoration results. In this paper, we propose FSIFD-GAN, a face de-morphing method based on frequency-spatial dual-feature attention separation mechanism. In this method, frequency information is introduced into the face de-morphing research for the first time, and the task of separating accomplice's identity features from morphed images is transferred to frequency domain and spatial domain. FSIFD-GAN uses the attention mechanism to realize the separation operation of frequency and spatial features in two domains respectively, and then uses the frequency-spatial cross-attention operation to realize the interaction of frequency and spatial information, so as to obtain comprehensive accomplice's identity information. Our experiments demonstrate that FSIFD-GAN can effectively restore the accomplice's facial image. This can help to identify the accomplice's identity in criminal investigations and judicial evidence.
Abstract Morphing attacks (MAs) pose a substantial security threat to the Automatic Border Control (ABC) system. While a few morphing attack detection (MAD) methods have been proposed, the face morphing accomplice's facial restoration has not received sufficient attention. Due to the inability to foresee the morphing factor used for a particular morphed image, selecting the appropriate de‐morphing factor becomes a challenging problem in the restoration of the accomplice's facial image. If the morphing factor cannot be chosen reasonably, achieving the desired restoration effect is difficult. Therefore, this paper presents an adaptive de‐morphing factor framework (ADFF) architecture for restoring the accomplice's facial image. By exploiting the morphed images stored in the electronic passport system and the real‐time captured criminal's images, ADFF can effectively restore the accomplice's facial image. Experimental results and analysis show that ADFF can significantly reduce the security threats of MAs on ABC.
Most face anti-spoofing methods address the generalization problem by extracting domain-invariant representations from multiple source domains or unlabelled target data. However, their deployment in real-world applications is unfeasible when data is insufficient or unavailable due to the collection costs and privacy concerns. This work investigates a more practical yet challenging scenario: single-source domain generalization based face anti-spoofing, where only one source domain is available during training and evaluated on multiple unseen target domains. To tackle this problem, a Causality-inspired Single-source Domain Generalization method (CSDG) is developed, which focuses on learning causal spoofing representations from the causality perspective. Specifically, a causal diagram is constructed to estimate the fundamental properties of ideal causal spoofing representations: remain invariant to shifts of domain-related confounders and causally sufficient for the detection category. To satisfy the above properties, the Causal Learning Module (CLM) maximizes the correlation of representations before and after intervention and minimizes the correlation with negative distributions. The intervention is achieved by arbitrarily performing spectrum mixup and structure destruction on source data within the Causal Intervention Model (CIM). Extensive experiments on four benchmark datasets validate the effectiveness of the proposed method.