PurposeThis work aims to investigate and improve adversarial patch attacks for semantic segmentation, a task increasingly deployed in security-critical applications. Existing attacks often overlook pixel-level uncertainty and spatial variation, resulting in inefficient optimization and limited effectiveness. The purpose of this study is to design an uncertainty-aware attack framework that better identifies and exploits structurally vulnerable regions in segmentation models.Design/methodology/approachWe propose a two-stage uncertainty-aware adversarial patch attack framework. The first stage computes pixel-wise entropy to identify locally uncertain regions. The second stage applies a confidence-based inter-pixel weighting strategy that prioritizes vulnerable pixels by comparing their confidence to a global statistical threshold. These components are unified into a dynamic loss reweighting mechanism. Experiments are conducted on Cityscapes and BDD100 K using ICNet, DDRNet, and SegFormer.FindingsExperimental results show that the proposed method outperforms existing patch-based attacks such as SSAP. By effectively targeting uncertain and structurally vulnerable regions, our method achieves stronger degradation of segmentation performance, with mIoU reduced to as low as 8%. The results demonstrate both high attack effectiveness and strong cross-dataset and cross-model generalization.Originality/valueThis work is the first to incorporate pixel-level uncertainty into adversarial patch optimization for semantic segmentation. Unlike prior patch-based attacks that treat all pixels uniformly, our method explicitly models local entropy and confidence-driven spatial variation, enabling more targeted and effective perturbation. The proposed dynamic loss reweighting framework provides a novel perspective on exploiting structural vulnerabilities in dense prediction tasks. This approach offers both theoretical and practical value for understanding segmentation robustness and designing stronger uncertainty-guided attacks.
Deep Neural Networks (DNNs) have made significant advancements in computer vision, widely applied to various tasks. However, these models remain vulnerable to adversarial attacks. This study aims to reveal the adversarial threats faced by visible-infrared detection models in real-world scenarios and proposes a unified adversarial patch method, i.e., a single patch design effective across both visible and infrared modalities, based on a genetic algorithm. This method enables selective or balanced attacks on visible and infrared detectors, providing an in-depth analysis of the security of models in practical applications. Experimental results show that the method effectively reduces the detection model's accuracy and demonstrates attack effects in simulated real-world scenarios. By optimizing the shape features of adversarial patches using a genetic algorithm and adaptively adjusting the attack strength across modalities via weight coefficients, the proposed method enhances the flexibility and robustness of cross-modal adversarial attacks. Additionally, the method employs the Expectation Transformation (EOT) strategy, showing strong robustness under different viewpoints. Extensive experiments validate the method's effectiveness, with an attack success rate (ASR) exceeding 89%. This study provides a theoretical basis for improving the robustness and security of models and offers valuable insights for safety-critical applications such as intelligent surveillance.
Deep neural network (DNN)-based detectors are highly vulnerable to physical adversarial camouflage, especially under multi-viewpoint settings, posing serious threats to vehicle detection systems. Existing differentiable-rendering-based attacks often suffer from two major limitations: (1) existing methods fail to capture complex environmental features during the rendering process, thereby undermining the stability of physical attacks; (2) current attention-based attack strategies fail to ensure consistent attacks across multi-scale targets, thereby constraining the transferability of adversarial camouflage. To address these limitations, we propose an adversarial camouflage generation method that concurrently integrates transferability, robustness, and generality. First, we develop an environmental-feature renderer that fuses depth maps to accurately synthesize multi-environment imagery, markedly enhancing attack robustness under complex environmental perturbations. Second, we introduce an attention-guided attack optimization strategy that enforces consistency across targets of varying scales, thereby substantially improving attack transferability. Third, we develop a Metaheuristic-based Texture Improvement Strategy (MTIS) that hybridizes the Whale Optimization Algorithm (WOA) and Moth-Flame Optimization (MFO) to escape local optima in the complex, non-convex search space by optimizing within a reduced-dimensional latent space. Unlike prior methods reliant on specific vehicle textures, our framework produces universal, model-agnostic patterns. Extensive experiments on six mainstream object detectors (e.g., Yolov5, Faster R-CNN) demonstrate that our approach achieves an average attack success rate (ASR) of 70.5% under ideal conditions and maintains 68.7% ASR in adverse environments (e.g., rain, fog, snow, night), outperforming existing state of the art physical attacks by 9-10%, and highlighting superior transferability and robustness.
Purpose With the development of wireless networks and artificial intelligence technology, unmanned aerial vehicle (UAV) clusters are widely used in various fields and cluster intelligence attacks are more harmful. However, most methods defending against UAV clusters produce consumption of non-reusable resources. To address this problem, a tethered UAV is adopted to perform active defense against adversary UAV clusters in this article, which can reduce the consumption of nonreusable resources. Design/methodology/approach Using tethered UAV to enter the opponent’s UAV cluster and analyze the flow of packets in adversary UAV cluster to find and occupy the central node. The tethered UAV can acquire and analyze key packets by deploying a grayhole attack at the location of the central node, after which the packets are selectively tampered with and discarded to cripple the opposing UAV cluster. Findings Comparing packet loss rate and delay with a normal network and the network that suffered from grayhole attack, it can be seen that the proposed scheme makes the tethered UAV close to the normal nodes in the UAV cluster and difficult to be detected. In addition, the tethered UAV is able to capture more packets compared to the other two networks, and the average deviation of the tethered UAV in capturing packets is around 5% in repeated experiments. Originality/value This article proposes an active defense method assisted by tethered UAV, which can minimize the consumption of nonreusable resources. The tethered UAV is converged to ordinary nodes of the opponent’s cluster, in which it is not easily detected. It provides a new direction for point-air defense technology.
Artificial intelligence (AI) image generation powered by large language model (large language modelss (LLMs)) enables highly realistic image synthesis and manipulation, posing significant security risks to Internet of Things (IoT) systems, particularly in identity authentication and data integrity. Although multidomain synthetic image detection has advanced, how spatial and frequency domain features affect decision making is still an open question, causing models to emphasize less critical areas and fall into local optimality. Through multidomain empirical analysis, we reveal the common contrastive differences in image textures and further demonstrate that frequency analysis helps capture the spectral differences in images. Building on this, we propose the collaborative spatial and frequency detector (CSFD). First, the image is decomposed into strong and weak texture regions in the spatial domain. Second, it aggregates different components in the frequency domain, using weighted channel attention to enhance spatial reasoning. Finally, the texture regions are combined to discriminate synthetic images. Experimental results demonstrate that incorporating channel attention based on frequency information improves the detection of synthetic images with spectral defects. On a comprehensive AI-generated image detection benchmark, the proposed method improves accuracy by 2.61% over current methods.
Detectors have been extensively utilized in various scenarios such as autonomous driving and video surveillance. Nonetheless, recent studies have revealed that these detectors are vulnerable to adversarial attacks, particularly adversarial patch attacks. Adversarial patches are specifically crafted to disrupt deep learning models by disturbing image regions, thereby misleading the deep learning models when added to into normal images. Traditional adversarial patches often lack semantics, posing challenges in maintaining concealment in physical world scenarios. To tackle this issue, this paper proposes a Prompt-based Natural Adversarial Patch generation method, which creates patches controllable by textual descriptions to ensure flexibility in application. This approach leverages the latest text-to-image generation model—Latent Diffusion Model (LDM) to produce adversarial patches. We optimize the attack performance of the patches by updating the latent variables of LDM through a combined loss function. Experimental results indicate that our method can generate more natural, semantically rich adversarial patches, achieving effective attacks on various detectors.
Purpose Current multi-source image fusion methods frequently overlook the issue of detailed features when employing deep learning technology, resulting in inadequate target feature information. In real-world mission scenarios, such as military information acquisition or medical image enhancement, the prominence of target feature information is of paramount importance. To address these challenges, this paper introduces a novel infrared-visible light fusion model. Design/methodology/approach Leveraging the foundational architecture of the traditional DenseFuse model, this paper optimizes the backbone network structure and incorporates a Unique Feature Encoder (UFE) to meticulously extract the distinctive features inherent in the two images. Furthermore, it integrates the Convolutional Block Attention Module (CBAM) and the Squeeze and Excitation Network (SE) to enhance and replace the original spatial and channel attention mechanisms. Findings Compared to other methods such as IFCNN, NestFuse, DenseFuse, etc., the values of entropy, standard deviation, and mutual information index of the method presented in this paper can reach 6.9985, 82.6652, and 13.6022, respectively, which are significantly improved compared with other methods. Originality/value This paper presents a UFEFusion framework that synergizes with the CBAM attention mechanism to markedly augment the extraction of detailed features relative to other methods. Moreover, the framework adeptly extracts and amplifies unique features from disparate images, thereby elevating the overall feature representation capability.
With the rapid development of deep learning research and applications, security issues in artificial intelligence are becoming increasingly prominent. Deep neural network models are frequently under attack, especially from backdoor attacks, posing significant threats to the security of these models. Currently, a key focus of backdoor attack research is on designing covert triggers to achieve hidden objectives. However, we find that while these covert triggers achieve good results in digital domains, they are susceptible to environmental changes in the real physical world, such as Gaussian blur. To address this issue, we propose a universal, semantic-based, visible trigger (USV Trigger) in this paper, which can be applied to different classification models while maintaining effectiveness and stealthiness. The backdoor attacks designed in this paper are not only applicable to digital domains but also demonstrate good attack effectiveness in the physical world.
In recent years, generative artificial intelligence has been developing rapidly. In the image domain, image generation models based on deep learning have made remarkable achievements. Early frameworks for image generation models were dominated by generative adversarial networks (GANs) and variational autoencoders (VAEs). Nowadays, large-scale generative models based on diffusion models have become mainstream, and the quality of their generated images is significantly improved. We will review the research and development of image generation models and delve into the significant progress made in the field in recent years. Initially, we revisit the development of traditional image generation models like GANs and VAEs, emphasizing their contributions and challenges. We also introduce diffusion models, which have received much attention in the field of image generation due to their unique generative process and excellent generative performance. Subsequently, we emphasized the large vision models with SAM as the focal point. We also pay special attention to large-scale generative models like Stable Diffusion, which have demonstrated unprecedented capabilities in high-quality image generation tasks. Additionally, we explore target models and respective fine-tuning methods for domain-oriented image generation tasks, predicts future directions in image generation, and proposes potential research focuses and challenges.
This paper tackle the challenges associated with low recognition accuracy and the detection of occlusions when identifying long-range and diminutive targets (such as UAVs). We introduce a sophisticated detection framework named UAV-YOLOv5, which amalgamates the strengths of Swin Transformer V2 and YOLOv5. Firstly, we introduce Focal-EIOU, a refinement of the K-means algorithm tailored to generate anchor boxes better suited for the current dataset, thereby improving detection performance. Second, the convolutional and pooling layers in the network with step size greater than 1 are replaced to prevent information loss during feature extraction. Then, the Swin Transformer V2 module is introduced in the Neck to improve the accuracy of the model, and the BiFormer module is introduced to improve the ability of the model to acquire global and local feature information at the same time. In addition, BiFPN is introduced to replace the original FPN structure so that the network can acquire richer semantic information and fuse features across scales more effectively. Lastly, a small target detection head is appended to the existing architecture, augmenting the model’s proficiency in detecting smaller targets with heightened precision. Furthermore, various experiments are conducted on the comprehensive dataset to verify the effectiveness of UAV-YOLOv5, achieving an average accuracy of 87
Abstract To prevent malicious activities and automated programmes from infiltrating and attacking websites or systems, a secure circular text-based CAPTCHA is designed based on a multi-secret visual cryptography. In this paper, multiple circular secret images are randomly generated by the server- side are encrypted into two circular share images, one of the share images is saved while the other is distributed to the user. Random characters are selected from each secret image to dynamically generate a circular CAPTCHA, with enhancing the authentication function for legitimate users and providing more effective resistance to phishing attacks. From the evaluation and recognition of the quality of the CAPTCHA, the circular text-based CAPTCHA has achieved a balance between usability and security at a certain extent.
Visual cryptography scheme is a method of encrypting secret image into n noiselike shares. The secret image can be reconstructed by stacking adequate shares. In the past two decades, many schemes have been proposed to realize the cheating prevention visual cryptography scheme (CPVCS). Significantly, Ren et al. [ 9 ] first introduced the idea of CPVCS with the help of Latin square. Inspired by their work, in this article, a new reliable scheme is proposed. More precisely, to facilitate the certification process, we embed meaningful characters into the randomly chosen authentication patterns in each divided blocks. Furthermore, we fix the security vulnerability in the stacked results of share S g and verification Ver g , where 1≤ g ≤ n . Since the improved scheme encrypts the secret image by utilizing random grids, the generated shares have no pixel expansion. Finally, theoretical analysis and experimental results are conducted to evaluate the efficiency and security of the proposed scheme.
Visual cryptography scheme (VCS) is an image secret sharing method which exploits the spatial responds characteristics of Human visual system (HVS). Applications of traditional OR and XOR based VCSs are seriously limited by the implementation carrier in practice. In this paper, we proposed a new kind of VCS which is implemented on modern display terminal of high refresh rates. Our approach exploits the temporal responds characteristics of HVS that light signals are temporal integrated into a single steady continuous one if the frequency exceeds critical fusion frequency (CFF). Furthermore, basing on the proposed VCS, we implement an information security display technology that can prevent unauthorized photography. Only authorized viewers can recover the secret information with the help of synchronized glass. While unauthorized viewers with naked eye or camera get nothing about the secret information. Experimental results show the effectiveness of proposed temporal integration based VCS and information security display technology.
Information security is arousing wide attention of the world. Apart from the traditional attack on transmission process, threats specific to display terminals should not be ignored as well. In this paper, we investigate the color mixture model of a well-known psychophysical phenomenon: “persistence of vision”, and propose a new display technology which presents secret and cheating information at the same time. Unauthorized viewers only see the cheating information from display, while authorized viewers can obtain the secret information with the help of synchronized glasses. Experimental results show the rationality and effectiveness of the technology.
In order to preventing cheating attack, partition authentication visual cryptography scheme (PAVCS) was proposed, in which each authentication pattern can be revealed on the stacked result of the corresponding partitions of any two shares. Without needing extra share, a PAVCS not only accomplishes sharing a secret image, but also provides mutual authentication of every two shares. In this paper, we propose two types of collusive attacks to (k, n)-PAVCS. The aim of the first type of attack is to destroy the integrity of the secret image revealed by victims, while the second type of attack belongs to cheating attack. Experimental results and theoretical analysis show that the two types of collusive attacks are successful attacks. Furthermore, we put forward the suggestion to devise partition authentication visual cryptography scheme for resisting the proposed two types of collusive attacks.
In the past decade, the researchers paid more attention to the cheating problem in visual cryptography (VC) so that many cheating prevention visual cryptography schemes (CPVCS) have been proposed. In this paper, the authors propose a novel method, which first makes use of Latin square to prevent cheating in VC. Latin squares are utilised to guide the choosing of authentication regions in different rows and columns of each divided block of the shares, which ensures that the choosing of authentication regions is both random and uniform. Without pixel expansion, the new method provides random regions authentication in each divided block of all shares. What is important is that the proposed method is applicable to both (k, n)-deterministic visual cryptography scheme ((k, n)-DVCS) and (k, n)-probabilistic visual cryptography scheme ((k, n)-PVCS). Experimental results and properties analysis are given to show the effectiveness of the proposed method.
Tagged visual cryptography scheme (TaVCS) is a new type of visual cryptography scheme, in which an additional tag image is revealed visually by folding up each share. A TaVCS not only carries augmented information in each share, but also provides user-friendly interface to identify each share. In this paper, we present a novel method to construct (k, n)-TaVCS. It can adjust visual quality of both the reconstructed secret image and the recovered tag image flexibly. Meanwhile, the proposed method provides better visual quality of both the reconstructed secret image and the recovered tag image under certain condition. Experimental results and theoretical analysis demonstrate the effectiveness of the proposed method.
A visual cryptography scheme (VCS) encodes a secret image into several share images, such that stacking sufficient number of shares will reveal the secret while insufficient number of shares provide no information about the secret. The beauty of VCS is that the secret can be decoded by human eyes without needing any cryptography knowledge nor any computation. Variance is first introduced by Hou et al. in 2005 to evaluate the visual quality of size invariant VCS. Liu et al. in 2012 thoroughly verified this idea and significantly improved the visual quality of previous size invariant VCSs. In this paper, we first point out the security defect of Hou et al.’s multi-pixel encoding method (MPEM) that if the secret image has simple contours, each single share will reveal the content of that secret image. Then we use variance to explain the above security defect.
In visual cryptography schemes (VCS), we often denote the set of all parties by P = { 1 , 2 , ⋯ , n } . Arumugam et al. proposed a ( k , n ) -VCS with one essential party recently, in which only subset S of parties satisfying S ⊆ P and | S | ≥ k and 1 ∈ S can recover the secret. In this paper, we extend Arumugam et al.’s idea and propose a ( k , n ) -VCS with t essential parties, say ( k , n , t ) -VCS for brevity, in which only subset S of parties satisfying S ⊆ P and | S | ≥ k and { 1 , 2 , … , t } ∈ S can recover the secret. Furthermore, some bounds for the optimal pixel expansion and optimal relative contrast of ( k , n , t ) -VCS are derived.
A (k, n) threshold secret image sharing scheme, abbreviated as (k, n)-TSISS, splits a secret image into n shadow images in such a way that any k shadow images can be used to reconstruct the secret image exactly. In 2002, for (k, n)-TSISS, Thien and Lin reduced the size of each shadow image to 1 k of the original secret image. Their main technique is by adopting all coefficients of a (k − 1)-degree polynomial to embed the secret pixels. This benefit of small shadow size has drawn many researcher’s attention and their technique has been extensively used in the following studies. In this paper, we first show that this technique is neither information theoretic secure nor computational secure. Furthermore, we point out the security defect of previous (k, n)-TSISSs for sharing textual images, and then fix up this security defect by adding an AES encryption process. At last, we prove that this new (k, n)-TSISS is computational secure.