In digital watermarking schemes, imperceptibility, robustness, and capacity are three fundamental yet conflicting performance metrics. Recent advancements in statistical model-based approaches have attracted growing attention due to their potential to balance these competing requirements. Existing methods are limited in their ability to simultaneously ensure sufficient capacity while enhancing both robustness and imperceptibility. To address this limitation, we propose a hybrid domain statistical image watermarking (HDSIW) scheme by leveraging vector finite Student's-t mixture (FSM) based hidden Markov tree (HMT) modeling of non-subsampled Shearlet transform (NSST) domain fast generic polar complex exponential transform (FGPCET) magnitudes. The NSST-FGPCET magnitudes are offered as a novel watermark embedding domain. A novel edge-based adaptive embedding localization method is proposed. The vector FSM-HMT is constructed using the marginal statistical features and dependencies of the NSST-FGPCET magnitudes. The decoder is derived based on the maximum likelihood criterion and the vector FSM-HMT. In addition, we propose a metric to effectively measure the balance between imperceptibility and robustness. Extensive experiments have shown that the proposed HDSIW outperforms state-of-the-art methods in terms of imperceptibility and robustness. The HDSIW can accommodate sufficient watermark capacity with favorable imperceptibility and robustness.
Synthetic Aperture Radar (SAR) automatic target recognition (ATR) is an important task in military reconnaissance and maritime monitoring. However, spatial-domain models struggle to preserve the geometric configuration of sparse and discontinuous SAR scattering centers. Frequency-domain analysis can reveal scattering variations that are less explicit in spatial features, but existing spatial-frequency aggregation often introduces frequency-domain responses without explicitly assessing their relevance to target-related scattering patterns, which may bring structure-irrelevant cues into the final representation. To address this challenge, we propose the Topology-Spectrum Guided Enhancement Network (TSGE-Net). TSGE-Net consists of two key components, namely the Structure-Aware Module (SAM) and the Cross-domain Feature Fusion Module (CFFM). Specifically, SAM organizes spatially discontinuous scattering responses into topology-aware representations. CFFM then performs cross-domain fusion by selectively incorporating frequency-domain cues that are consistent with the learned scattering organization. Extensive experiments on four public SAR datasets and a mixed-category dataset show that TSGE-Net consistently outperforms competitive baselines and maintains robust performance across diverse SAR classification settings. Code is available at https://github.com/LiJiarui1104/TSGE-Net.git.
Digital image watermarking technology has become a critical approach for safeguarding information security. The key performance metrics for evaluating watermarking algorithms include watermark capacity, robustness, and imperceptibility; however, these factors are often in tension with one another. Achieving an optimal balance among these factors remains a central challenge in watermarking algorithm research. This study presents an image watermarking algorithm founded on statistical principles, utilizing the vector bivariate generalized linear exponential (BGLE) distribution as its core model. To begin, we employ the undecimated dual-tree complex wavelet transform (UDTCWT) in conjunction with the improved polar harmonic Fourier moments (IPHFMs) transform to extract the UDTCWT-IPHFMs magnitude domain, which exhibits high resistance to attacks. The vector BGLE distribution is then used to model the magnitude domain of UDTCWT-IPHFMs, capturing the dependencies across different scales, directions, and subbands effectively. We apply the maximal evidential likelihood estimates (MELEs) method to estimate the model parameters. Leveraging the vector BGLE model, we construct a novel watermark decoder. To further enhance the security of the embedded information, we introduce an innovative dual watermarking technique based on the statistical modeling algorithm. The security is reinforced through mutual authentication between the authentication watermark and the target watermark. Experimental evaluations indicate that the proposed method successfully balances robustness, imperceptibility, and watermark capacity. Compared with existing state-of-the-art techniques, the performance of the constructed decoder in watermark extraction is significantly improved.
Zero-watermarking provides a non-intrusive solution for copyright protection and is particularly suitable for fidelity-sensitive images. However, existing moment-based zero-watermarking methods still suffer from high computational cost, limited numerical accuracy, and insufficient robustness and discriminability, especially for color images. To address these issues, this paper proposes an efficient color image zero-watermarking framework based on Quaternion Fractional Generalized Pseudo-Jacobi-Fourier Moments (QFGPJFMs). First, a fast and accurate computation strategy is developed for generalized pseudo-Jacobi-Fourier moments (GPJFMs) by integrating recursive computation, polar pixel tiling, and Gaussian quadrature integration. Second, fractional-order parameters are introduced to construct FGPJFMs, which provide more flexible basis control and improved representational adaptability. Third, the resulting descriptors are extended to the quaternion domain to jointly model RGB channels and preserve inter-channel correlations. Finally, QFGPJFMs are combined with mixed loworder moment features (MLMF) and asymmetric mapping to build a robust and discriminative zerowatermarking framework for color-image copyright protection. Experimental results show that the proposed method achieves an accuracy rate of 1.0000 under most common signal-processing and geometric attacks on the USC-SIPI dataset, maintains 0.9963 under upper-left cropping with a ratio of 1/8, and yields a false positive rate of 0 in the discriminability test. Additional experiments further verify its security, efficiency, and watermark capacity.
Imperceptibility and robustness are two fundamental yet competing requirements in digital image watermarking, especially under frequency-domain degradations such as JPEG compression. To address this challenge, this paper proposes Wav-UNet, a deep watermarking framework that integrates full-subband wavelet transform (FSWT), inverse FSWT, and deformable convolution into a unified encoder-decoder architecture. Unlike conventional U-Net-based watermarking models that rely on standard downsampling and upsampling, Wav-UNet propagates multi-scale spatial-frequency information throughout both embedding and extraction. Deformable convolutions further adapt the receptive fields to local image structures, improving robustness against localized and geometric distortions. In addition, a perceptual subband-weighting strategy and a local-global strength-control mechanism are introduced to regulate embedding intensity according to human-visual-system sensitivity and regional distortion consistency. The training objective is formulated as a robustness-oriented objective dominated by watermark recovery, with residual regularization and frequency consistency serving as auxiliary constraints. Experimental results show that the proposed method achieves high visual quality, with 42.73 dB PSNR and 0.9836 SSIM, while maintaining strong extraction accuracy under various attacks, including JPEG compression, noise, filtering, cropping, and dropout. Additional JPEG sensitivity experiments further evaluate the robustness boundary under very low quality factors such as QF=10. These results demonstrate that Wav-UNet provides a favorable trade-off among imperceptibility, robustness, and computational efficiency.
Digital image watermarking is an important part of information hiding technology, which can effectively protect the copyright of information. Robustness, imperceptibility and watermark capacity are important factors to evaluate the quality of digital image watermarking algorithms, but they are mutually constrained and have inherently contradictory relationships. In recent years, in order to solve this problem, it has been found that the image watermarking algorithm based on statistical models can achieve an effective balance between the three. This paper presents an image watermarking method based on vector bivariate conditional Weibull distribution (BCWD), undecimated dual tree complex Wavelet transform (UDTCWT), and fast accurate Chebyshev-Fourier moments (FACHFMs) transform domain. When embedding watermark, the watermark is embedded into the UDTCWT-FACHFMs magnitudes with strong robustness by the multiplicative embedding method. In the stage of watermark extraction, since vector BCWD can accurately capture the statistical distribution of the image and the dependence between scales and directions, we use it to model the UDTCWT-FACHFMs magnitudes. The parameters of statistical model are obtained by probability weighted moment based on power density method (PWMBP). Local optimal (LO) rule is often used in watermark detection schemes. Therefore, we combine the LO rule with vector BCWD to propose a new statistical image detector. A large number of experiments prove that the image watermarking scheme based on vector BCWD model designed in this paper can balance robustness, imperceptibility and watermarking capacity. According to the experimental results, after embedding 1024 bits watermark information, the obtained PSNR values are all greater than 42 dB, and the designed detector can still maintain good detection performance under various attacks.
Local binary pattern (LBP) and its variants are effective for general textures but become less robust under grayscale inversion. To better balance discriminability and robustness, we propose a grayscale-inversion and rotation invariant image descriptor, called the Non-local Binary Derivative Pattern (NBDP). For robustness, NBDP encodes the magnitude response in the Gaussian derivative domain to achieve invariance to grayscale inversion. For discriminability, NBDP uses two encoding methods, non-local-to-neighborhood encoding (NBDP_N) and non-local-to-center encoding (NBDP_C), to capture local and non-local interactions. These methods are complementary and rotation invariant. Experiments on three benchmark databases (i.e. Outex, CUReT and KTH-TIPS) demonstrate that NBDP outperforms state-of-the-art methods under linear or nonlinear grayscale-inversion changes.
Stereo images have recently gained considerable attention due to their immersive nature, highlighting an urgent need for robust copyright protection mechanisms. However, most existing zero-watermarking algorithms are tailored for 2D images and do not adequately meet the unique requirements of stereo images. Moreover, current methods for zero-watermarking stereo images often fail to accurately represent and maintain the critical relationship between the left and right views, thereby limiting their effectiveness. To overcome these limitations, this paper proposes an innovative zero-watermarking method specifically designed for stereo images, which leverages an accurate ternary polar linear canonical transform (ATPLCT). We first introduced a new computational technique called the accurate polar linear canonical transform (APLCT) to address the numerical integration problems inherent in the polar linear canonical transform (PLCT). Next, we extend the APLCT using ternary number theory to develop the ATPLCT, which is specifically optimized for capturing stereo image characteristics. Finally, we propose a stereo image zero-watermarking strategy that integrates the ATPLCT with an asymmetric tent map. Comparative experiments and analyses show that our proposed method offers improved performance and greater robustness compared to existing approaches.
Balancing the relationship between imperceptibility, robustness, and watermarking capacity is a difficult challenge for digital image watermarking algorithms to overcome. In this paper, we design a well-performing statistical blind watermarking scheme to protect the copyright of digital images. First, the bidimensional empirical mode decomposition (BEMD) is combined with the Schur decomposition (SD) to construct the robust digital watermarking carrier. Then, the positionally stabilized BEMD-SD domain coefficients are selected to hide the information with the help of entropy threshold edge detection technique. Next, the simplified bivariate generalized t (SBGt) distribution is used for accurate modeling to account for the strong correlation between the coefficients, and the model parameters are computed by the inverse harmonic Newton maximum likelihood estimation (CHN-MLE) method. Finally, the digital watermark decoder is constructed to extract watermark information by combining the maximum likelihood criterion. Experimental results on a large number of test images show that the proposed blind watermarking decoder outperforms most of the state-of-the-art statistical methods and deep learning methods recently proposed in the literature. Our proposed method has excellent imperceptibility and robustness in accommodating watermarks of the same capacity.
Sorted-based LBP variants have been validated as effective grayscale inverse image classification methods. However, most of these methods encode the order of sampling points at the same scale and thus suffer from two problems: 1) Ignoring inter-scale correlation leads to descriptors that are not resistant to real scene changes. 2) The inherent flaws of sorted encoding cause descriptors to discriminate complex texture structures, showing low discriminability. To address these problems, we design the new scale-structure model and region encoding to realize a more robust and discriminative representation called Local Radial Grouped Invariant Order Pattern (LRGIOP). LRGIOP can effectively distinguish texture details in real scenes while resisting various complex imaging conditions. Experiments on several image databases show that the LRGIOP descriptor achieves state-of-the-art classification results under linear or even nonlinear grayscale-inversion transformations.
Robustness, imperceptibility, and watermark capacity are three indispensable and contradictory properties for any image watermarking systems. It is a challenging work to achieve the balance among the three important properties. In this paper, by using the bivariate Beta mixture model (Bivariate BMM), we present a statistical image watermark scheme in nonsubsampled Contourlet transform (NSCT)-fractional order Jacobi Fourier moments (FJFMs) amplitude hybrid domain. The whole watermarking algorithm includes two parts: watermark embedding and detection. NSCT is firstly performed on host image to obtain the frequency subbands, and the NSCT subbands are divided into non overlapping blocks. Then, the significant NSCT domain blocks are selected using local binary patterns (LBP). Meanwhile, for each selected NSCT coefficient block, FJFMs are calculated to obtain the NSCT-FJFMs amplitude. Finally, watermark signals are inserted into the amplitude hybrid domain of NSCT-FJFMs. In order to detect accurately watermark signal, the statistical characteristics of NSCT-FJFMs magnitudes are analyzed in detail. Then, NSCT-FJFMs magnitudes are described statistically by Bivariate BMM, which can simultaneously capture the marginal distribution and strong dependencies of NSCT-FJFMs magnitudes. Also, Bivariate BMM parameters are estimated accurately by the rough-enhanced-Bayes mixture estimation & expectation-maximization (REBMIX&EM) approach. Finally, a statistical watermark detector based on the locally optimum (LO) decision rule and Bivariate BMM is developed in NSCT-FJFMs magnitude hybrid domain. Also, the closed-form expressions on the LO statistic is derived and the receiver operating characteristic (ROC) about our detector is analyzed. Extensive experimental results show the superiority of the proposed image watermark detector over some state-of-the-art statistical watermarking methods.
Copy-move forgery, where an image region is copied into another part of the same image, is one of the most common and easy-to-implement image tampering techniques. Keypoint (also called feature point)-based detection methods exhibit remarkable performance in terms of computational cost and robustness. However, these methods have the following limitations to varying degrees: 1) Failure to extract keypoints from small or smooth regions; 2) Lack of robust and discriminative descriptors for keypoints; 3) Low accuracy and excessive computational cost of keypoints matching; 4) High false negative/positive rate caused by the defects of post processing. To overcome such limitations, we propose an adaptive copy move forgery detection based on new keypoint feature and matching in this paper. Firstly, based on the simple linear iterative clustering (SLIC) and the multi-directional multi-layer double-cross pattern (MDML-DCP), i.e., the segmentation method is adopted, and uniform key points are adaptively extracted from the whole image (including small and smooth regions) by fitting the MDML-DCP value-threshold function of superpixels. Then, due to the strong robustness of moment features to attacks and their stronger descriptive ability compared to traditional invariant moments. Secondly, texture features can get more distinctive features. Next, accurate quaternion fractional pseudo-Jacobi-Fourier moments (AQFPJFM) and gradient local ternary patterns (GLTP) describe the key points to obtain robust and discriminative features. Then, the ITQ-PTH algorithm is introduced for keypoint matching, which is more accurate than traditional locality-sensitive hashing algorithms and can improve the matching accuracy. Finally, the false negative rate and false positive rate are reduced by reliable post-processing methods. Experimental results show that the proposed method achieves an F-score of 96.61
In the field of digital watermarking, three fundamental and interdependent requirements must be satisfied: robustness, invisibility, and payload capacity. Achieving an optimal trade-off among the three requirements remains a significant challenge. This paper introduces a novel statistical image watermarking method that operates in the nonsubsampled contourlet transform (NSCT)-pseudo Jacobi Fourier moments (PJFMs) magnitude domain, wherein a probability density function (PDF) derived from the Weibull-Burr impounded bivariate distribution (WBIBD) is employed. The proposed statistical watermarking framework consists of two main components: embedding and detection. During the embedding phase, the original image is first decomposed using NSCT, followed by the segmentation of high-frequency subbands into non-overlapping blocks. PJFMs is then computed for NSCT coefficient blocks and digital watermarks are embedded into robust NSCT-PJFMs magnitudes. In the detection phase, the robust local NSCT-PJFMs magnitudes are first modeled using the WBIBD, which accurately captures both the marginal distributions and the strong dependencies of these magnitudes simultaneously. The parameters of WBIBD are computed efficiently by trimmed L-moments estimation. Finally, a watermark detector is then constructed by integrating the WBIBD model with the locally most powerful (LMP) test. Furthermore, closed-form expressions for the watermark detector are derived using the WBIBD framework. Experimental results demonstrate that the probability of detection and AUROC value (area under receiver operating characteristic curve) of the proposed watermarking algorithm are higher than those of other detectors. The results indicate superior detection performance, and it achieves a more effective balance between imperceptibility, robustness, and payload.
Local binary pattern (LBP) and many variants adopt the difference information to encode all the neighborhood binarization, demonstrating excellent ability to distinguish images. However, these methods are very sensitive to reverse gray changes. How to improve the robustness of reverse gray changes and maintain excellent classification ability has become a problem to be solved. In this paper, an image captioning method called multiscale image captioning method based on wavelet transform is proposed (WMLBP). This method uses stationary wavelet transform to remove the redundant information of the image and extract features from the frequency domain. This paper also proposes an innovative coding scheme, which designs three descriptors, and uses multi-scale hierarchical thresholds and complementary information to enhance the ability of image feature description. Finally, the image features were described by combining all the descriptors and using the multi-scale and multi-resolution cross-scale joint representation. Tests across multiple databases have demonstrated that the preprocessing of this method can significantly enhance classification capabilities. Compared to other algorithms, it has a higher accuracy rate and faster computational speed. Experiments demonstrate that the method exhibits strong robustness against grayscale inversion changes (both linear and nonlinear) and image rotation, while also demonstrating excellent classification performance. It holds significant importance for addressing the challenges encountered in image classification. Our code is available at: https://github.com/Dawei-W/WMLBP/tree/master.
Local binary pattern (LBP) and its variants have demonstrated excellent distinguishability in facing different challenges. However, most of these LBP methods use scalar thresholds to encode all neighborhood binarization. There are thus two main problems: 1) they are highly prone to encode two neighborhoods with differences in texture structure as the same LBP code. 2) They cannot describe the interactions between the neighborhood and the information in the region where the neighborhood is located. Given this, this paper proposes a Generalized Multiscale Hierarchical Threshold (GMHT) framework, which can effectively capture texture information at different scales and regions of an image, thus solving the first problem. For the second problem, we propose the Regional Gradient Pattern (RGP), which is a 3-D joint and cascade of the gradient operator (G), the extremum operator (E), the variance operator (V) and the center operator (C). Respectively, the four operators encode different texture information of the local and the region to describe the local-region interactions. Benefiting from the calculation method, they are invariant to the grayscale-inversion. Experiments on four texture databases (Outex, KTH-TIPS, CUReT, USPtex1) show that this descriptor achieves state-of-the-art classification results in the presence of linear and even non-linear grayscale-inversion transformations. Our code is available at: https: //github.com/xuyanqi971202/RGP_Demo.git.
Index-to-palm interaction plays a crucial role in Mixed Reality(MR) interactions. However, achieving a satisfactory inter-hand interaction experience is challenging with existing vision-based hand tracking technologies, especially in scenarios where only a single camera is available. Therefore, we introduce Palmpad, a novel sensing method utilizing a single RGB camera to detect the touch of an index finger on the opposite palm. Our exploration reveals that the incorporation of optical flow techniques to extract motion information between consecutive frames for the index finger and palm leads to a significant improvement in touch status determination. By doing so, our CNN model achieves 97.0% recognition accuracy and a 96.1% F1 score. In usability evaluation, we compare Palmpad with Quest’s inherent hand gesture algorithms. Palmpad not only delivers superior accuracy 95.3% but also reduces operational demands and significantly improves users’ willingness and confidence. Palmpad aims to enhance accurate touch detection for lightweight MR devices.
Image watermarking technology poses a significant challenge, requiring a delicate balance between robustness, imperceptibility, and capacity. To solve the balancing problem, a statistical learning based blind image watermarking technique is proposed in this paper. The idea of the proposed method to solve the balancing problem is to enhance imperceptibility and robustness while accommodating sufficient watermarks. Firstly, the local low-order pseudo-Zernike moments (PZM) magnitude in the discrete non-separable Shearlet transform (DNST) domain is constructed as the embedding domain. To ensure stability in embedding positions, we propose an edge embedding strategy. Secondly, by analyzing the statistical characteristics and multi-correlations of DNST-PZM magnitudes, a multi-correlation vector model based on the skew student's-t mixture (SSM) and hidden Markov tree (HMT) is designed. Finally, leveraging the vector SSM-HMT model and the maximum likelihood (ML) criterion, we derive the closed-form decoder expression. Comparative analysis with state-of-the-art methods demonstrates the superior imperceptibility and robustness of our proposed approach in accommodating the same capacity watermark.
Robustness, imperceptibility, and watermark capacity are three indispensable and contradictory properties for any image watermarking systems. It is a challenging work to achieve the balance among the three important properties. In this paper, by using bivariate Birnbaum–Saunders (BRBS) distribution model, we present a statistical image watermark scheme in nonsubsampled shearlet transform (NSST)-pseudo Zernike moments (PZMs) magnitude hybrid domain. The whole watermarking algorithm includes two parts: watermark embedding and extraction. NSST is firstly performed on host image to obtain the frequency subbands, and the NSST subbands are divided into non overlapping blocks. Then, the significant high-entropy NSST domain blocks are selected. Meanwhile, for each selected NSST coefficient block, PZMs are calculated to obtain the NSST-PZMs amplitude. Finally, watermark signals are inserted into the amplitude hybrid domain of NSST-PZMs. In order to decode accurately watermark signal, the statistical characteristics of NSST-PZMs magnitudes are analyzed in detail. Then, NSST-PZMs magnitudes are described statistically by BRBS distribution, which can simultaneously capture the marginal distribution and strong dependencies of NSST-PZMs magnitudes. Also, BRBS statistical model parameters are estimated accurately by modified closed-form maximum likelihood estimator (MML). Finally, a statistical watermark decoder based on BRBS distribution and maximum likelihood (ML) decision rule is developed in NSST-PZMS magnitude hybrid domain. Extensive experimental results show the superiority of the proposed image watermark decoder over some state-of-the-art statistical watermarking methods and deep learning approaches.
Invisibility, robustness and payload are three indispensable and contradictory properties for any image watermarking systems. To achieve the tradeoff among the three requirements, statistical watermarking approaches have received increasing attention in recent years. But, most existing schemes often bear a number of drawbacks, in particular: (1) They mainly utilize transform coefficients, which are always fragile to some attacks, especially global geometric transforms, for watermark inserting and statistical modeling; (2) the adopted statistical models always cannot capture accurately both marginal distribution and various strong dependencies between coefficients; and (3) the used parameter estimation usually has high time complexity and poor computational accuracy. This has motivated us to introduce in this paper a novel statistical image watermarking in singular value decomposition (SVD)-undecimated discrete wavelet transform (UDWT) difference domain using vector Alpha Skew Gaussian (VB-ASG) distribution. We begin with a detailed study on the robustness and statistical characteristics of local SVD-UDWT difference coefficients of natural images. This study reveals the excellent robustness, highly non-Gaussian marginal statistics and strong dependencies of local SVD-UDWT difference coefficients. We also find that conditioned on their generalized neighborhoods, the local SVD-UDWT difference coefficients can be approximately modeled as vector Alpha Skew Gaussian (VB-ASG) variables. Meanwhile, model parameters can be estimated effectively by using approximate maximum likelihood estimation (AMLE) approach. Based on these findings, we model local SVD-UDWT difference coefficients using VB-ASG model that can capture marginal statistics and strong dependencies. Finally, we develop a new statistical image watermark decoder using the VB-ASG model and maximum likelihood (ML) decision rule. Our experimental evaluation results validate that our image watermarking leads to performance improvements comparable to several state-of-the-art statistical watermarking methods and some approaches based on convolutional neural networks. In particular, under the condition of the same watermarking capacity, the imperceptibility and robustness of this method show certain superiority compared with other algorithms.
With the rapid development of image editing software, forged images have become a serious social problem because of their great destructiveness. Copy-move is one of the most commonly used types of forgery. The keypoint-based copy-move forgery detection (CMFD) techniques identify forged regions by extracting image keypoints and using local visual features. These methods show remarkable detection capability in some areas, such as memory requirements and computational cost. After years of research, the challenges of the current keypoint-based methods are as follows: 1) The number or distribution of extracted keypoints is not satisfactory, especially in smoothed and textured areas. 2) Many local visual features have poor robustness to geometric attacks and post-processing disturbances, which fundamentally limits the performance of algorithms. 3) The existing post-processing algorithms can't effectively filter out false matched pairs and accurately locate the tampered areas. To overcome these problems, an accurate and robust image copy-move forgery detection method using adaptive keypoints and hybrid features is proposed. Firstly, an adaptive keypoint extraction method based on the simple linear iterative clustering (SLIC) and the K-multiple-means (KMM) is proposed, which can extract the dense and uniform keypoints in the whole image. Then, we combine the transform domain features based on the Fast Quaternion Generic Polar Complex Exponential Transform (FQGPCET) and the texture features based on the Gray-level co-occurrence matrix (GLCM) to obtain the robust hybrid features. The hybrid features have outstanding descriptive power. Then the feature matching is performed by the double-bit quantized locally sensitive hash (DBQ-LSH). Finally, a high-precision post-processing algorithm includes two-step filtering and two-step clustering is proposed. The experimental results demonstrate that the overall performance of the proposed algorithm is superior to that of other solutions for detecting copy-move forgery images.