With the continuous advancement of speech synthesis algorithms, synthetic speech has become increasingly realistic, posing unprecedented challenges to speaker recognition (verification) systems and information security. To address the escalating security risks caused by high-fidelity synthetic speech, this paper proposes a synthetic speech detection model named Group Delay and Latent Attention enhanced XLS-R+Conformer (GDLA-XLS-R+Conformer), based on the previously successful XLS-R+Conformer baseline. The model incorporates group delay to supplement phase features, thereby enhancing the ability to detect phase anomalies in synthetic speech. Furthermore, it employs a Symmetrical Multi-head Latent Cross Attention (SMLCA) and a Learnable Gating Fusion (LGF) to achieve tight integration of waveform and phase information. Experimental results show that the proposed GDLA-XLS-R+Conformer outperforms existing methods on multiple datasets, including ASVspoof 2019 logical access (LA), ASVspoof 2021 LA, and ASVspoof 2021 deepfake (DF), demonstrating its effectiveness and robustness across diverse scenarios.
Existing face forgery detection (FFD) models often generalize poorly across datasets and manipulation types because their features entangle genuine content and method-specific artifacts with forgery cues. We propose Dual Feature Progressive Disentanglement (DFPD), a unified framework that learns robust, domain-invariant forgery representations by disentangling both local spatial regions and common forgery semantics. A Local Feature Disentanglement (LFD) module exploits pixel-wise entropy to build entropy-guided spatial masks, emphasizing transferable regions while suppressing ambiguous backgrounds and low-confidence patches. In parallel, a Common Forgery Disentanglement (CFD) module couples adaptive balanced normalization with adversarial learning to attenuate domain-specific style variations and promote cross-domain alignment. To further prevent residual content leakage, we adopt an orthogonal feature projection that explicitly removes content components from the learned forgery embedding. These components are trained under a Progressive Feature Disentanglement Strategy (PFDS), which applies fine-to-coarse disentanglement constraints from shallow to deep network stages to stabilize optimization. Extensive cross-dataset and cross-manipulation evaluations, including robustness tests under common image degradations and benchmarks with emerging diffusion-based generative forgeries, show that DFPD consistently outperforms state-of-the-art detectors on unseen domains and attacks. Ablation studies further validate the complementary contributions of LFD, CFD, and PFDS within the proposed framework.
Collaborative perception significantly enhances the situational awareness of autonomous vehicles but simultaneously introduces severe vulnerabilities to adversarial attacks from malicious agents. Prevailing defense paradigms are hindered by a fundamental trade-off: Hypothesize-and-Verify methods offer generalization but are prohibitively slow due to their multi-pass decoding, while Feature-Vector Classification is efficient but suffers from poor generalization due to lossy feature vectorization. This paper addresses this critical trade-off by introducing Function-Space Adjudication (FSA), a novel design principle that adjudicates the trustworthiness of features directly in the function space, thus preserving high-fidelity information. As an innovative implementation of this paradigm, we propose FNO-Guard, a framework that leverages a Fourier Neural Operator to learn a direct mapping from a pair of full feature functions to a maliciousness score. The proposed approach uniquely synthesizes the single-pass efficiency of Feature-Vector Classification with the holistic, information-rich analysis of consensus-based approaches. By operating on complete feature maps, FNO-Guard learns to identify fundamental functional inconsistencies—the hallmark of any adversarial manipulation, rather than overfitting to specific attack signatures, thus achieving strong generalization. Extensive experiments on the simulated V2X-Sim benchmark and the real-world DAIR-V2X dataset demonstrate that FNO-Guard achieves strong defense performance against diverse attacks while maintaining real-time efficiency. In particular, it outperforms existing defenses under standard attack settings and achieves 62 FPS, providing a nearly 3 × speedup over Hypothesize-and-Verify methods. Comparisons across diverse architectures confirm that full-map preservation improves defense, with the FNO operator achieving optimal performance. This empirically validates FNO-Guard and the FSA principle for secure collaborative perception.
With the increasing application of electrical network frequency (ENF) in forensic audio and video analysis, ENF signal detection has emerged as a critical technology. However, high-pass filtering operations commonly employed in modern communication scenarios, while effectively removing infrasound to enhance communication quality at reduced costs, result in a substantial loss of fundamental frequency information, thereby degrading the performance of existing detection methods. To tackle this issue, this paper introduces Multi-HCNet, an innovative deep learning model specifically tailored for ENF signal detection in high-pass filtered environments. Specifically, the model incorporates an array of high-order harmonic filters (AFB), which compensates for the loss of fundamental frequency by capturing high-order harmonic components. Additionally, a grouped multi-channel adaptive attention mechanism (GMCAA) is proposed to precisely distinguish between multiple frequency signals, demonstrating particular effectiveness in differentiating between 50 Hz and 60 Hz fundamental frequency signals. Furthermore, a sine activation function (SAF) is utilized to better align with the periodic nature of ENF signals, enhancing the model’s capacity to capture periodic oscillations. Experimental results indicate that after hyperparameter optimization, Multi-HCNet exhibits superior performance across various experimental conditions. Compared to existing approaches, this study not only significantly improves the detection accuracy of ENF signals in complex environments, achieving a peak accuracy of 98.84%, but also maintains an average detection accuracy exceeding 80% under high-pass filtering conditions. These findings demonstrate that even in scenarios where fundamental frequency information is lost, the model remains capable of effectively detecting ENF signals, offering a novel solution for ENF signal detection under extreme conditions of fundamental frequency absence. Moreover, this study successfully distinguishes between 50 Hz and 60 Hz fundamental frequency signals, providing robust support for the practical deployment and extension of ENF signal applications.
In recent years, with the rapid development of deep forgery technology, its verisimilitude is increasing day by day, and its social impact is becoming more and more serious. However, although a variety of face deep forgery video detection algorithms have been proposed and have shown certain detection capabilities on open source data sets, in the face of increasingly sophisticated deep forgery technology, the differences between genuine and fake videos are gradually difficult to be captured by the naked eye, and existing detection methods generally have problems such as low cross-compressibility detection and poor robustness. Therefore, in order to improve the detection accuracy and model robustness, a deep forged video detection method named SFDT is proposed. In this scheme, the framework structure of fusion of air frequency domain is adopted. Firstly, feature extraction is enhanced by improving MVIT in the airspace and using ASFF adaptive module; secondly, frequency domain features are extracted by dynamic filter in the frequency domain; then FECAM is used to reduce the loss caused by frequency domain information transmission; finally, multi-mode fusion module is used for feature fusion. This air-frequency fusion detection scheme can not only improve the cross-compression detection performance of the model, but also effectively deal with various interference in the transmission of video, and improve the robustness of the model.
X-ray image-based prohibited item detection plays a crucial role in modern public security systems. Despite significant advancements in deep learning, challenges such as feature extraction, object occlusion, and model complexity remain. Although recent efforts have utilized larger-scale CNNs or ViT-based architectures to enhance accuracy, these approaches incur substantial trade-offs, including prohibitive computational overhead and practical deployment limitations. To address these issues, we propose Xray-YOLO-Mamba, a lightweight model that integrates the YOLO and Mamba architectures. Key innovations include the CResVSS block, which enhances receptive fields and feature representation; the SDConv downsampling block, which minimizes information loss during feature transformation; and the Dysample upsampling block, which improves resolution recovery during reconstruction. Experimental results demonstrate that the proposed model achieves superior performance across three datasets, exhibiting robust performance and excellent generalization ability. Specifically, our model attains mAP50-95 of 74.6% (CLCXray), 43.9% (OPIXray), and 73.9% (SIXray), while demonstrating lightweight efficiency with 4.3 M parameters and 10.3 GFLOPs. The architecture achieves real-time performance at 95.2 FPS on the GPUs. In summary, Xray-YOLO-Mamba strikes a favorable balance between precision and computational efficiency, demonstrating significant advantages.
The electric network frequency (ENF), often referred to as the industrial heartbeat, plays a crucial role in the power system. In recent years, it has found applications in multimedia evidence identification for court proceedings and audio–visual temporal source identification. This paper introduces an ENF region classification model named UniTS-SinSpec within the UniTS framework. The model integrates the sinusoidal activation function and spectral attention mechanism while also redesigning the model framework. Training is conducted using a public dataset on the open science framework (OSF) platform, with final experimental results demonstrating that, after parameter optimization, the UniTS-SinSpec model achieves an average validation accuracy of 97.47%, surpassing current state-of-the-art and baseline models. Accurate classification can significantly aid in ENF temporal source identification. Future research will focus on expanding dataset coverage and diversity to verify the model’s generality and robustness across different regions, time spans, and data sources. Additionally, it aims to explore the extensive application potential of ENF region classification in preventing crimes such as telecommunications fraud, terrorism, and child pornography.
Existing face recognition methods cannot effectively eliminate the influence of corrupted features caused by occlusion. As the features flow deeper, the corrupted features get entangled with the effective features used for identity classification, which affects the recognition results. To address the problem, this paper designs an occluded face recognition method based on segmentation and multi-stage mask learning strategy. The model consists of three components: occlusion detection and segmentation, feature extraction, and mask learning unit. The proposed method only needs one end-to-end process to learn feature masks and deep occlusion-robust features without relying on additional occlusion detectors. The mask learning units take different sizes of occlusion segmentation representations and facial features of different stages as input, generate corresponding feature masks for different stages of feature extraction, and effectively eliminate the influence of corrupted features caused by occlusion at each stage of feature extraction through mask operations. Finally, a feature pyramid is constructed to fuse features of different stages for identity classification. Experimental results show that the proposed method can effectively improve the accuracy of occluded face recognition. The accuracy on the occluded LFW dataset and the real masked datasets MFR2 and Mask_whn reach 98.77%, 96.70% and 81.53%, respectively, which has an accuracy improvement of 2.04, 0.48 and 4.44 percentage points compared with the existing mainstream methods.
With the proliferation of urban video surveillance systems, the abundance of surveillance video data has emerged as a pivotal asset for enhancing public safety. Within these video archives, the identification of abnormal human actions carries profound implications for security incidents. Nevertheless, existing surveillance systems primarily rely on conventional algorithms, leading to both missed incidents and false alarms. To address the challenge of automating multi-object surveillance video analysis, this study introduces a comprehensive method for the detection and recognition of multi-object abnormal actions. This study comprises a two-stage framework: the coarse detection stage employs an enhanced YOWOv2E model for spatio-temporal action detection, while the precise detection stage utilizes a two-stream network for precise action classification. In parallel, this paper presents the PSA-Dataset to address the current limitations in the field of abnormal action detection. Experimental results, collected from both public datasets and a self-built dataset, illustrate the effectiveness of the proposed method in identifying a wide spectrum of abnormal actions. This work offers valuable insights for automating the analysis of human actions in videos pertaining to public security.
针对目前基于深度学习的恶意代码分类算法易出现灾难性遗忘导致分类准确率不高、收敛过慢的问题,提出基于隐式随机梯度下降的恶意代码分类算法.与现有算法不同,该算法构造内外循环网络结构来协同学习最优网络模型以提高恶意代码分类准确率.在内循环优化阶段,通过优化带有偏好正则项的损失函数迫使内循环网络沿外循环网络方向更新权重从而避免内循环网络遗忘过去学到的知识.在外循环优化阶段,通过求解近似外循环网络梯度并利用隐式随机梯度下降优化外循环网络权重,使得外循环网络能够更快更稳定地收敛.在三个恶意代码数据集上的实验结果表明,该算法有效避免了灾难性遗忘,使用较少训练轮数取得了最高的分类准确率,显著提升了恶意代码分类的稳定性和鲁棒性.
Rapid development of deepfake technology led to the spread of forged audios and videos across network platforms, presenting risks for numerous countries, societies, and individuals, and posing a serious threat to cyberspace security. To address the problem of insufficient extraction of spatial features and the fact that temporal features are not considered in the deepfake video detection, we propose a detection method based on improved CapsNet and temporal-spatial features (iCapsNet-TSF). First, the dynamic routing algorithm of CapsNet is improved using weight initialization and updating. Then, the optical flow algorithm is used to extract interframe temporal features of the videos to form a dataset of temporal-spatial features. Finally, the iCapsNet model is employed to fully learn the temporal-spatial features of facial videos, and the results are fused. Experimental results show that the detection accuracy of iCapsNet-TSF reaches 94.07%, 98.83%, and 98.50% on the Celeb-DF, FaceSwap, and Deepfakes datasets, respectively, displaying a better performance than most existing mainstream algorithms. The iCapsNet-TSF method combines the capsule network and the optical flow algorithm, providing a novel strategy for the deepfake detection, which is of great significance to the prevention of deepfake attacks and the preservation of cyberspace security.
In recent years, voice deepfake technology has developed rapidly, but current detection methods have the problems of insufficient detection generalization and insufficient feature extraction for unknown attacks. This paper presents a forged speech detection method (HuRawNet2_modified) based on a self-supervised pre-trained model (HuBERT) to improve detection (and address the above problems). A combination of impulsive signal-dependent additive noise and additive white Gaussian noise was adopted for data boosting and augmentation, and the HuBERT model was fine-tuned on different language databases. On this basis, the size of the extracted feature maps was modified independently by the α-feature map scaling (α-FMS) method, with a modified end-to-end method using the RawNet2 model as the backbone structure. The results showed that the HuBERT model could extract features more comprehensively and accurately. The best evaluation indicators were an equal error rate (EER) of 2.89% and a minimum tandem detection cost function (min t-DCF) of 0.2182 on the database of the ASVspoof2021 LA challenge, which verified the effectiveness of the detection method proposed in this paper. Compared with the baseline systems in databases of the ASVspoof 2021 LA challenge and the FMFCC-A, the values of EER and min t-DCF decreased. The results also showed that the self-supervised pre-trained model with fine-tuning can extract acoustic features across languages. And the detection can be slightly improved when the languages of the pre-trained database, and the fine-tuned and tested database are the same.
With the rapid development of deepfake technology, the authenticity of various types of fake synthetic content is increasing rapidly, which brings potential security threats to people's daily life and social stability. Currently, most algorithms define deepfake detection as a binary classification problem, i.e., global features are first extracted using a backbone network and then fed into a binary classifier to discriminate true or false. However, the differences between real and fake samples are often subtle and local, and such global feature-based detection algorithms are not optimal in efficiency and accuracy. To this end, to enhance the extraction of forgery details in deep forgery samples, we propose a multi-branch deepfake detection algorithm based on fine-grained features from the perspective of fine-grained classification. First, to address the critical problem in locating discriminative feature regions in fine-grained classification tasks, we investigate a method for locating multiple different discriminative regions and design a lightweight feature localization module to obtain crucial feature representations by augmenting the most significant parts of the feature map. Second, using information complementation, we introduce a correlation-guided fusion module to enhance the discriminative feature information of different branches. Finally, we use the global attention module in the multi-branch model to improve the cross-dimensional interaction of spatial domain and channel domain information and increase the weights of crucial feature regions and feature channels. We conduct sufficient ablation experiments and comparative experiments. The experimental results show that the algorithm outperforms the detection accuracy and effectiveness on the FaceForensics++ and Celeb-DF-v2 datasets compared with the representative detection algorithms in recent years, which can achieve better detection results.
随着科学技术的迅速发展,基于深度学习生成的合成语音给语音认证系统和网络空间安全带来了新的挑战.针对现有检测模型准确率较低和语音特征挖掘不够充分的问题,提出了一种基于Involution算子和交叉注意力机制改进的合成语音检测方法.前端将语音数据提取线性频率倒谱系数(LFCC)特征和恒定Q变换(CQT)谱图特征,两个特征分别输入到后端的双分支网络中.后端网络使用ResNet18作为主干网络先进行浅层的特征学习,并将Involution算子嵌入主干网络,扩大特征图像学习区域,增强在空间范围内学习到的频谱图像特征信息.同时在训练分支之后引入cross-attention交叉注意力机制,使LFCC特征和CQT谱图特征构建交互的全局信息,强化模型对特征的深层挖掘.所提模型在ASVspoof2019 LA测试集上取得了 0.84%的等错误率和0.026的最小归一化串联检测代价函数的实验结果,展现了优于主流的检测模型.结果表明,改进的模型能够有效融合不同的频谱特征,提高模型的特征学习能力,从而强化模型的检测能力.
网站指纹识别技术通过分析流量特征判断用户访问的网站站点,能够有效监管TOR匿名网络的用户行为.现有的识别方法通常需要大规模的数据样本以获得高的识别准确率,且普遍存在概念漂移问题.针对以上问题,本文提出一种基于残差和协作对抗网络(Residual network and Collaborative and Adversarial Network,Res-CAN)的网站指纹识别模型.该模型使用残差网络(Residual network)作为特征提取器以减少网络的优化难度.同时,将协作对抗网络(Collaborative and Adversarial Network,CAN)应用于网站指纹识别问题,使得特征提取器同时学习领域相关和领域无关特征,实现源域与目标域的特征空间对齐.实验结果表明,本文提出的方法在小样本环境下网站指纹识别准确率达到91.2%,优于现有的利用对抗领域自适应网络(Domain-Adversarial Neural Networks,DANN)迁移学习方法,且抗概念漂移能力较高.
Tor anonymous traffic identification technology provides a mechanism to combat illegal and criminal activities in the dark network using Tor anonymous communication tools.However,some challenges exist,such as data collection difficulties,unbalanced datasets,and the poor ability of the Tor analysis model to detect and adapt to conceptual drift.First,the collected original Tor PCAP traffic is segmented,denoised,and processed into byte sequences.Then,one-dimensional sequences are transformed into visual grayscale images and input to an improved multi-size Deep Convolution Generate Adversarial Network(DCGAN) to generate Tor traffic samples for data balancing. Finally,a Stacked Denoising Auto-Encoder(SDAE) is used for sequence dimensionality reduction,and the extracted features are input to an Online Sequential Extreme Learning Machine(OS-ELM) to realize the online flow recognition of Tor traffic.The experimental results show that the improved DCGAN can be used to improve the quality of data sets and improve the model recognition rate by about 2.8 percentage points. The accuracy of the traffic analysis model combined with OS-ELM and SDAE can reach 95.7%,and the recognition efficiency is greatly improved compared with traditional Convolutional Neural Network(CNN) and Long Short-Term Memory(LSTM) network models.
随着物联网的大规模使用,其安全问题也日益严峻,如何在资源有限的物联网环境中准确实时检测网络攻击是亟需解决的关键问题.基于流量特征的入侵检测系统是物联网安全的一种解决方案,但该方案存在流量特征数量繁多、不利于训练快速轻量的检测模型的问题.针对该问题,文章提出一种基于特征选择的物联网轻量级入侵检测方法相关性系数和方差膨胀因子的特征选择方法.该方法在流粒度下对流量特征进行选择,通过机器学习算法对正常流量和恶意流量进行分类.实验结果表明,该方法能在有限的资源下快速有效地识别网络攻击行为,综合精确度与召回率达到99.4%.
Aiming at the problems in the field of Android malicious family detection,such as insufficient code visualization method construction information,large classification effect affected by the number of data sets and low classification accuracy,an Android malicious family classification method based on multi feature file synthetic image and Xception improved model is proposed.Fir-stly,three feature files corresponding to RGB multi-channel are selected to synthesize color images.Then,the improved Xception model introduces the focal loss function to alleviate the negative impact caused by the uneven distribution of samples.Finally,the attention mechanism is integrated into the improved model to extract the image features of malicious code from different dimensions,which improves the classification effect of the model.Experimental results show that the malicious code images synthesized by the proposed method contain richer features,have higher accuracy than the mainstream malicious family classification methods,and have better classification effect for unbalanced data sets.
Face forgery detection is drawing ever-increasing attention in the academic community owing to security concerns. Despite the considerable progress in existing methods, we note that: Previous works overlooked fine-grain forgery cues with high transferability. Such cues positively impact the model’s accuracy and generalizability. Moreover, single-modality often causes overfitting of the model, and Red-Green-Blue (RGB) modal-only is not conducive to extracting the more detailed forgery traces. We propose a novel framework for fine-grain forgery cues mining with fusion modality to cope with these issues. First, we propose two functional modules to reveal and locate the deeper forged features. Our method locates deeper forgery cues through a dual-modality progressive fusion module and a noise adaptive enhancement module, which can excavate the association between dual-modal space and channels and enhance the learning of subtle noise features. A sensitive patch branch is introduced on this foundation to enhance the mining of subtle forgery traces under fusion modality. The experimental results demonstrate that our proposed framework can desirably explore the differences between authentic and forged images with supervised learning. Comprehensive evaluations of several mainstream datasets show that our method outperforms the state-of-the-art detection methods with remarkable detection ability and generalizability.
The accuracy of existing face recognition models cannot improve due to the influence of masks and other occlusion factors.The current mainstream research methods integrate and apply the occluded and unoccluded scenes to multiple scenes after separate training.Aiming at the limitation of occluded face recognition model, this paper proposed an improved face feature rectification network(FFR-Net) model.This model could be used for face recognition with or without occlusion, and be applied to mask and glasses occlusion recognition scenes.FFR-Net proposed a face feature rectification module.In order to make full use of the feature information of the unocclusion area, the spatial branch of the module introduced involution operator to expand the image information interaction area and enhance the face feature information in the spatial range.The channel branch introduced coordinate attention to capture cross channel information to enhance the feature representation, which was conducive for the model to locate and identify the target area more accurately.Using Meta-ACON as a new dynamic activation function, it improved model generalization and calculation accuracy by dynamically adjusting the degree of linearity or nonlinearity.Finally, this paper trained the improved FFR-Net on the CASIA-Webface processed face dataset with or without mask occlusion.The accuracy of the test results on the LFW processed face dataset with or without mask occlusion and Meglass dataset are 82.50% and 89.75% respectively, which is superior to the existing algorithm, and verifies the effectiveness of the proposed method.