As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates a fundamental crisis in artificial intelligence security. Existing defense architectures heavily rely on empirical semantic guardrails and probabilistic large model adjudicators, mechanisms that fail to provide deterministic security lower bounds when facing complex semantic symbol decoupling attacks. To overcome this empirical semantic guardrail dilemma, this paper proposes a new security paradigm for agents based on the fundamental limitations of logical reasoning. Based on this paradigm, we further introduce an executable Proof-Constrained Action (ePCA) framework with a neural symbolic isolation architecture. This framework abandons semantic trust in natural language, forcing agents to losslessly formalize their intentions into first-order logical mathematical constraints before performing physical operations. Empirical evaluations of macroscopic and microscopic two-dimensional dynamic adversarial systems demonstrate that our formal verification mechanism achieves zero attack success rate and zero false positive rate across the evaluated scenarios, with extremely low computational latency. This research provides a conditional formal foundation under explicit system assumptions and an engineering paradigm for constructing the underlying defense foundation for future intelligent systems.
Entanglement routing selects a path to establish entanglement connections between two arbitrary nodes in quan-tum networks, which plays an important role in quantum com-munication. In quantum networks, quantum decoherence and limited network performance make it challenging to distribute entangled pairs. Many entanglement routing schemes have been proposed to solve this issue but most of them are in a centralized and synchronized manner. However, they may be infeasible in large-scale quantum networks. Therefore, in this paper, we propose a distributed and asynchronous entanglement routing scheme called DFER in which quantum nodes manage requests autonomously. The major challenge is quantum nodes have little knowledge about entangled pairs, which hinders the ability to establish fidelity guaranteed entanglement connections. To ad dress this challenge, we develop DLFR algorithm which estimates the fidelity of end-to-end entanglement connections based on link-level fidelity and calculates link-level fidelity requirement based on the characteristic of purification. Among nodes which meet link-level fidelity requirement, we design DFPS path se lection algorithm to select next hop with the highest expected throughput to distribute entangled pairs. Numerous simulation results demonstrate that DFER can efficiently distribute fidelity guaranteed entangled pairs with high throughput.
Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-speech, however, high-quality systems are still commonly built through an intermediate acoustic representation before waveform synthesis. In this work, we present BareWave, a fully waveform-native framework for direct text-to-wave generation in flow-matching TTS. We consider this setting to raise three training challenges: raw-waveform modeling lacks a strong pretrained representational scaffold, different stages of training benefit from different noise schedules, and data-space perceptual objectives do not automatically share the temporal structure of the velocity-space flow objective. As a result, direct waveform training is hard to optimize efficiently, hard to push toward a strong final operating point with a fixed recipe, and hard to integrate effective perceptual refinement. Guided by this view, we develop a direct text-to-wave training framework that combines training-time representation alignment, staged noise scheduling, and velocity-aware perceptual alignment (VAPA), while preserving a single waveform-native inference path without pretrained components at test time. Experiments on zero-shot voice cloning show that strong intelligibility, speaker similarity, and naturalness can be achieved under a fully waveform-native inference path, supporting waveform-native flow-matching TTS as a practical direction. Project page with audio demos is available at https://barewave.github.io/.
Face swapping has become a prominent research area in computer vision and image processing due to rapid technological advancements. The metric of measuring the quality in most face swapping methods relies on several distances between the manipulated images and the source image, or the target image, i.e., there are suitable known reference face images. Therefore, there is still a gap in accurately assessing the quality of face interchange in reference-free scenarios. In this study, we present a novel no-reference image quality assessment (NR-IQA) method specifically designed for face swapping, addressing this issue by constructing a comprehensive large-scale dataset, implementing a method for ranking image quality based on multiple facial attributes, and incorporating a Siamese network based on interpretable qualitative comparisons. Our model demonstrates the state-of-the-art performance in the quality assessment of swapped faces, providing coarse- and fine-grained. Enhanced by this metric, an improved face-swapping model achieved a more advanced level with respect to expressions and poses. Extensive experiments confirm the superiority of our method over existing general no-reference image quality assessment metrics and the latest metric of facial image quality assessment, making it well suited for evaluating face swapping images in real-world scenarios.
Transformer models have advanced single image super-resolution (SISR), but most emphasize enlarging the receptive field while underusing high-frequency cues across spatial and channel dimensions. To mitigate this self-attention(SA)’s low-pass bias, We propose FANSR (Frequency Adaptive Network for Efficient Image Super Resolution), a lightweight framework that considering receptive-field expansion with explicit high-frequency modeling. FANSR introduces two frequency-based token mixers that transform features into frequency representations along spatial and channel axes, efficiently isolating high-frequency components and injecting them as sparse priors into the subsequent Self-Attention block. We further introduce a Frequency Enhancement Loss (FE-Loss) that adaptively prioritizes high-frequency regions, enabled by the model’s explicit distinction of low- and high-frequency content. Extensive experiments show that FANSR achieves a superior trade-off between image quality and latency compared to state-of-the-art lightweight methods. Code and pretrained models will be released.
Digital watermarking constitutes a fundamental technical pillar for addressing the visual trust crisis engendered by artificial intelligence-generated content. Existing digital watermarking survey literature is predominantly confined to a static classification perspective, rendering it inadequate for elucidating the continuously escalating dynamic adversarial co-evolution between watermarking defenses and attack methodologies. To this end, grounded in the perspective of security game theory, a “prevention-tracing-resistance” ternary analytical framework was proposed oriented toward the full lifecycle of artificial intelligence-generated content. The core mechanisms of watermarking technology were systematically examined across three principal dimensions, interdicting unauthorized data mining to achieve source-level governance of the content generation process, copyright provenance tracing and identity authentication for high-dimensional synthesized forgery content, and resisting deep adversarial erasure attacks to ensure the robust persistence of embedded watermarks. Finally, the critical challenges pertaining to insufficient cross-modal generalization capability and the continuous escalation of adversarial attack models were analyzed, and future evolutionary trajectories of proactive defense watermarking were prospected. It aims to provide a systematic reference framework for both theoretical inquiry and engineering practice in next-generation artificial intelligence-generated content proactive defense systems, thereby fostering the secure and sustainable development of a trustworthy artificial intelligence-generated content ecosystem.
Digital watermarking embeds imperceptible information into images to support copyright protection and source tracing. While recent deep learning-based watermarking methods achieve strong robustness under various distortions, they typically assume fixed input resolutions. However, images often appear in diverse sizes and aspect ratios, and existing extensions such as residual scaling or block-wise embedding offer only partial solutions, leading to degraded robustness, reduced visual quality. To address these challenges, we propose VaRiA (Variable-Resolution Image Adaptive robust watermarking), a Transformer-based framework that models pixel relationships consistently across resolutions, while a dual-stage training strategy first ensures robustness on fixed resolutions and then adapts to variable resolutions. Experiments demonstrate that VaRiA achieves nearly 100% watermark extraction accuracy under diverse distortions and resolutions, while maintaining high image fidelity (44 dB PSNR), confirming its robustness and practical value.
Image steganography is an essential technique for concealing information by embedding secret information within images to make it undetectable. In recent years, with the rapid development and popularization of text-to-image generation models, many generated images have been disseminated through the Internet, thus making generated images ideal covers for steganography. Given that the distribution of generated images is more easily modeled than natural images, steganographic methods based on generated images exhibit higher security. Nevertheless, these methods typically require white-box access to the generative model, while contemporary popular generative models are black-box models. We observed that slight modifications in the input parameters of black-box image generative models result in subtle differences between generated images, offering new camouflage advantages for image steganography. Based on this observation, we propose an image steganography method based on the fluctuation of generative models. This approach leverages the fluctuation of image generative models, disguising stego images to appear as if they were generated by the parameter fluctuations of the generative model. Experimental results show that our proposed method outperforms baseline methods when facing steganalysis attacks, significantly enhancing steganographic security without compromising image quality.
Clothes-invariant feature extraction is critical to the clothes-changing person re-identification (CC-ReID). It can provide discriminative identity features and eliminate the negative effects caused by the confounder--clothing changes. But we argue that there exists a strong spurious correlation between clothes and human identity, that restricts the common likelihood-based ReID method P(Y|X) to extract clothes-irrelevant features. In this paper, we propose a new Causal Clothes-Invariant Learning (CCIL) method to achieve clothes-invariant feature learning by modeling causal intervention P(Y|do(X)). This new causality-based model is inherently invariant to the confounder in the causal view, which can achieve the clothes-invariant features and avoid the barrier faced by the likelihood-based methods. Extensive experiments on three CC-ReID benchmarks, including PRCC, LTCC, and VC-Clothes, demonstrate the effectiveness of our approach, which achieves a new state of the art.
Recent advancements in video generation technologies have been significant, resulting in their widespread application across multiple domains. However, concerns have been mounting over the potential misuse of generated content. Tracing the origin of generated videos has become crucial to mitigate potential misuse and identify responsible parties. Existing video attribution methods require additional operations or the training of source attribution models, which may degrade video quality or necessitate large amounts of training samples. To address these challenges, we define for the first time the "few-shot training-free generated video attribution" task and propose SWIFT, which is tightly integrated with the temporal characteristics of the video. By leveraging the "Pixel Frames(many) to Latent Frame(one)" temporal mapping within each video chunk, SWIFT applies a fixed-length sliding window to perform two distinct reconstructions: normal and corrupted. The variation in the losses between two reconstructions is then used as an attribution signal. We conducted an extensive evaluation of five state-of-the-art (SOTA) video generation models. Experimental results show that SWIFT achieves over 90% average attribution accuracy with merely 20 video samples across all models and even enables zero-shot attribution for HunyuanVideo, EasyAnimate, and Wan2.2. Our source code is available at https://github.com/wangchao0708/SWIFT.
Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no longer reside in any single input, but instead emerge from how otherwise benign or weakly harmful conditions interact across modalities and time. Existing safety benchmarks, however, remain largely prompt-centric or tied to fixed conditioning interfaces, leaving such compositional risks difficult to study systematically. To bridge this gap, we introduce Multi2AV-Safety, the first safety benchmark, to the best of our knowledge, to cover all 11 non-singleton T/I/A/V conditioning configurations for audio-video generation, comprising 11,024 attack instances. Evaluation on Multi2AV-Safety reveals systematic weaknesses in representative multimodal safety guards across attack mechanisms and harm-evidence structures. Our evaluation reveals two complementary failure modes: harmful semantics can emerge from the combination of individually benign inputs, while explicit harmful cues can become harder to detect when mixed with benign multimodal context. Together, these results identify compositional risk perception as a central capability gap in safeguarding multimodal-conditioned audio-video generation: current safety guards fail to reliably integrate safety evidence across modalities and time, even when all conditioning inputs are observable. The dataset will be publicly released in October 2026.
Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deployment criterion. The evidence requirements are strongly asymmetric. One attack that obtains harmful help from the deployed service suffices to establish that such help remains, and such attacks appear repeatedly in the coded record. Establishing that little remains cannot follow from the safeguard's own numbers alone; it also requires evidence about what the surrounding system still allows after the safeguard performs its local function. Such evidence is supported or derived in only a small minority of the depth-coded claims, and one such claim bounds its scoped residual. A better local score is therefore not, by itself, a stronger claim about the deployment. Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer.
With the development of deep learning, high-value and high-cost models have become valuable assets, and related intellectual property protection technologies have become a hot topic. However, existing model watermarking work in black-box scenarios originates mainly from training-based backdoor methods, which probably degrade primary task performance. To address this, we propose a branch backdoor-based model watermarking protocol named BranchWM to protect the intellectual property of the model. This protocol adopts a construction based on a message authentication scheme as the branch indicator, following a comparative analysis with other secure cryptographic primitives. We prove the lossless performance of the protocol by reduction. In addition, we analyze potential threats to the protocol and present a secure and feasible watermarking instantiation for language models. We further conduct empirical evaluations of the instantiated BranchWM, demonstrating its effectiveness and security for ownership verification.
Recent advances in generative models and technological innovations have significantly addressed the fundamental challenges of character image animation. However, existing approaches predominantly focus on character animation from a single reference image, substantially limiting their applicability in scenarios such as multiple character interaction animation. To fill this gap, this paper introduces MultiAnimate, a comprehensive framework that enables concurrent animation of multiple characters within a shared environment while preserving both identity consistency and spatial relationships. The framework achieves these objectives through multiple well-designed mechanisms. First, we incorporate an identity-specific reference net that enables appearance extraction from multiple reference images, distinguishing MultiAnimate from existing approaches constrained to single reference inputs. Second, we implement an identity-aware pose encoder to address the character-pose binding challenge, wherein an attention mechanism enables the network to accurately differentiate and process multiple pose sequences during generation. Third, we introduce an interaction guider module that enhances the framework's capability to handle complex inter-character interactions by leveraging character-specific mask information, serving as an optional component that refines the pose sequences. Extensive experiments and ablation analyses demonstrate our framework's superiority in multiple character animation, particularly in scenarios involving complex motion sequences.
Composition is a cornerstone of visual aesthetics, influencing the appeal of an image. While its principles operate independently of specific content, in practice, composition is often coupled with semantics. As a result, existing methods often enhance composition either through implicit learning or by semantics-based layout control, rather than explicitly modeling composition itself. To address this gap, we introduce Composer, a framework rooted in aesthetic theory, designed to model composition in a semantic-agnostic manner. First, it supports composition transfer by extracting key composition-aware representations from a reference image and leveraging a tailored conditional guidance module to control composition based on pre-trained diffusion models. Second, when users specify only text themes without a composition reference, Composer supports theme-driven composition retrieval by leveraging the in-context learning capabilities of Large Vision-Language Models (LVLMs), achieving explicit composition planning. To enhance composition in a reference-free mode, we conduct text-to-composition fine-tuning on the trained control module to enable implicit composition planning. Furthermore, we curated a high-quality dataset comprising 2 million image-text pairs using state-of-the-art generative models to support model training. Experimental results demonstrate that Composer significantly enhances aesthetic quality in text-to-image tasks and facilitates personalized composition control and transfer, offering users precision and flexibility in the creative process.
Online social networks (OSNs) offer an abundant and freely available source of images, providing fertile ground for steganographic communication. However, the mandatory lossy operations applied by these platforms—primarily JPEG recompression— make robustness a pressing challenge. Existing robust steganographic methods focus on improving the embedding process, but inevitably compromise security. In this paper, we break this trade-off by proposing, for the first time, a robust cover screening method that enables successful message extraction after JPEG recompression, even when combined with non-robust steganographic methods. To ensure that the screened covers are compatible with arbitrary steganographic settings—including distortion functions, coding schemes, and messages—we introduce Robustness-Minimizing Modification (RMM), which simulates the worst-case impact of steganographic modifications on cover robustness. Images that remain unchanged under JPEG recompression after RMM are screened as robust covers. Our experiments reveal that such robust covers exist widely in both natural and generated images. Therefore, recent advances in generative modeling enable cost-effective and scalable expansion of candidate covers, addressing potential limitations of the screening method in practice. Our experiments also demonstrate that these screened covers can achieve 100% message extraction even with non-robust steganography at high embedding rates, while maintaining security comparable to other covers.
The remarkable proficiency of diffusion models raise concerns about malicious applications in digital image manipulation, which poses significant challenges in distinguishing synthetic counterparts. In response to this challenge, we define a new benchmark, Diff-IML dataset, tailored for real-world diffusion-based image manipulations. It leverages image manipulation techniques based on diffusion models, encompassing four distinct types: remove, fill, replace and outpaint, to simulate real-world image editing scenarios. To tackle the above task, we propose a novel method for uncovering intrinsic manipulation cues, by leveraging self-correlations within low-level representations to compensate the spatial information loss in high-level localization decoder. To further enhance the completeness of the forgery features, an auxiliary classifier is designed to predict mask proportions of manipulation regions. Upon the Diff-IML dataset, extensive experimental results demonstrate our approach achieves superior performance to prior manipulation localization methods.
The astonishing proficiency and unprecedented level of realism of diffusion models in creating and manipulating images have undoubtedly drawn concerns.Many methods have been proposed to detect generated images. Typically, they usually take RGB images as input, and use backbones like ResNet, CLIP visual encoder to extract features. Even though these backbones are capable to detect fake images, they are mainly designed to extract the high-level semantic information, rather than inherently designed for fake image detection. To this end, in this paper, we want to optimize the embedding space tailored for detecting fake images via representation learning. We notice that Neighboring Pixel Relationships (NPR) is capable to capture the intrinsic forgery clues, which means that NPR may be a good input to perform representation learning that aims at learning the embedding space tailored for detecting fake images.Therefore, we leverage features from both RGB modality and NPR modality to perform two proposed representation learning methods, Cross-Modal Contrastive Learning (CMCL) and Cross-Modal Mutual Distillation (CMMD), in order to learn the forgery-aware embedding space. The CMCL boosts the discrimination of features between real and fake images, while the CMMD simultaneously transfers the learned knowledge between two modalities, being able to learn compact features within the intra-class. CMCL and CMMD work collaboratively so that each modality learns a more comprehensive forgery-aware representation to distinguish real and fake images.Extensive experiments on GenImage, DRCT-2M, and Co-Spy-Bench datasets show that our method achieves state-of-the-art results.
Backdoor attacks pose a critical threat to the security and reliability of deep neural networks (DNNs), enabling malicious triggers to manipulate model predictions while maintaining normal performance on clean inputs. Addressing this challenge requires robust defense mechanisms that eliminate backdoors without compromising model utility. This paper introduces PVDI (Preserving Vital and Disrupting Irrelevant attentions), a novel blind purification method designed to neutralize backdoors while retaining the primary prediction capabilities of both backdoored and clean models. PVDI leverages vital and irrelevant attentions during fine-tuning: vital attention is preserved to maintain the model's core functionality, while irrelevant attention is disrupted to neutralize backdoor behavior. Extensive evaluations across diverse datasets and attack scenarios demonstrate PVDI's superior purification performance, achieving significant reductions in attack success rates while preserving the utility of backdoored models and minimizing adverse impact on clean models. PVDI outperforms existing state-of-the-art defenses and sets a new benchmark for backdoor defense. This work represents a significant step forward in combating backdoor attacks.