To enhance the multi-scale object perception and representation capability of the YOLOv8n model in complex autonomous driving scenarios, this paper proposes a lightweight object detection algorithm named SDF-YOLO, based on a multi-scale dynamic feature fusion mechanism. On top of the original YOLOv8n structure, multiple architectural optimization strategies are introduced to strengthen its multi-scale feature fusion ability. First, an additional detection head is added to process shallow small-scale features, thereby improving fine-grained object detection capability. Second, four Sample, Parameters-Free Attention-Adjacent Context Coordination Modules (Sim-ACCoM) are embedded into the backbone and neck networks. By combining convolutions with varying dilation rates, Sim-ACCoM effectively enhances contextual awareness and cross-scale feature integration during feature extraction. Moreover, a Dynamic Interpolation Fusion (DIF) module is also proposed prior to the concatenation operation of the neck and following the C2f module which allows complementary representations of features and also improves richness of the representation. Lastly, parameters are pruned using the Group-Slim algorithm to enhance the efficiency of the parameters and lightweight properties. Experiments using publicly available datasets and a manually annotated Car dataset of 11 object categories show the SDF-YOLO gets an improvement of 1.7%, 3.4%, 3.3% and 3.7% in Precision, Recall, mAP50 and mAP50-95 parameters respectively, and the parameters are reduced by 1.2 M number than the original YOLOv8n. These findings prove that SDF-YOLO is an efficient method of the detector and structural efficiency. Experimental findings confirm SDFYOLO is effective to enhance detection strength as well as precision in challenging conditions due to scale differences, which shows its practicality in autonomous driving perception tasks.
Salient object detection plays a pivotal role in the interpretation of remote sensing imagery. However, this task remains challenging due to the inherent complexities of remote sensing scenes, which include cluttered backgrounds, substantial scale variations, indistinct object boundaries, and environmental noise. To improve localization accuracy in these complex scenarios, we propose a novel edge-guided stratified fusion network (ESFNet) for accurate remote sensing salient object extraction. The network achieves stratified fusion of fine-grained edge cues and rich semantic information through its hierarchical architecture, enabling precise salient object localization. Specifically, we design an EMFM to explicitly incorporate contour semantics, thereby enhancing boundary delineation while effectively suppressing background clutter and noise interference. Moreover, a stratified attention module aggregates contextual information across multiple semantic levels using a hierarchical attention mechanism. This dual-path architecture facilitates synergistic interaction between edge precision and semantic coherence, significantly improving feature discriminability and spatial consistency. Extensive experiments on two benchmark datasets demonstrate that our ESFNet achieves state-of-the-art performance, exhibiting remarkable superiority in detecting salient objects with sharp boundaries and consistent semantics across diverse remote sensing scenarios.
Transfer-based black-box attacks are an important tool for evaluating deployed vision models, yet adversarial examples generated from Vision Transformer (ViT) surrogates often exhibit limited cross-architecture transferability. Existing momentum-based attacks are effective for convolutional neural network (CNN) surrogates, but they can accumulate stale directions and overfit the surrogate when the source model is a ViT. This paper presents Ada-MGNS, a ViToriented transferable attack that combines adaptive momentum with deep attention guidance. The adaptive component measures the directional discrepancy between the current guided gradient and the accumulated trajectory, and then attenuates stale momentum when the search direction becomes unstable. The guidance component fuses the classification gradient with an auxiliary gradient extracted from the last transformer block’s attention responses, encouraging perturbations to disturb both output decisions and semantic aggregation. Experiments on ImageNet with four ViT surrogates, thirteen standard black-box targets, and five defense models show that Ada-MGNS consistently improves attack success rates over representative ViT-specific baselines, remains compatible with DI/TI transformations and effective against adversarially trained and purification-based defenses.
Due to the challenges of synthetic aperture radar (SAR) data acquisition and the high cost of manual annotation, utilizing labeled optical images to learn from unlabeled SAR data has received great attention. Cross-modal domain adaptation (DA) from optical to SAR imagery presents a particularly difficult problem because of the inherent modality gap between these two imaging paradigms. To address the problem of unsupervised ship detection in SAR images, we conduct domain adaptation experiments from ship images in the DIOR dataset to the SSDD dataset. However, traditional domain adaptation methods are insufficient to address the significant modality differences between optical and SAR images. Although the unconstrained feature alignment strategy is effective between domains with small differences, it inadvertently expels SAR features from the supervised recognition space, ultimately reducing detection performance. To mitigate this issue, we propose a new framework, DAIR, which integrates an innovative inversion regularization module (IRM) and a task-correlation enhancement (TCE) strategy to improve domain adaptation. Specifically, IRM acts as a feature-space regularizer to counteract deviation caused by aggressive alignment, while TCE explicitly models task interdependency to alleviate the effects of task independence. Evaluated on the DIOR and SSDD datasets, our method achieved improvements of 5% and 8% in AP50 and APm, respectively, over the baseline method DA Faster R-CNN. In few-shot scenarios, our method attained gains of 14.1%, 10.6%, and 15.4% in AP50 compared to the state-of-the-art method under 3-shot, 5-shot, and 10-shot settings, which demonstrates stronger generalization ability with limited data. Source code and models are available at https://github.com/whatbb/DAIR/tree/main
Radio interference sources mounted on high-speed mobile platforms present a considerable challenge to accurate localization and tracking due to their wide-ranging and dynamic characteristics. This article tackles the intricate problem of tracking erratically maneuvering interference sources and introduces a dynamic tracking scheme for the spatial decomposition of motion states. Firstly, the proposed scheme decomposes the complex three-dimensional motion into three independent one-dimensional linear movements. This facilitates the concurrent application of interacting multiple model matching filters across all three dimensions, enabling precise characterization of the complex motion states. Secondly, an enhancement is made to the update process of the noise covariance matrix within the real-time noise estimator. Specifically, a time-varying exponential forgetting factor is incorporated to ensure the robustness of the adaptive cubature Kalman filter algorithm when combined with the noise estimator. Simulation results show that the proposed tracking scheme outperforms traditional methods in terms of both the speed of tracking error convergence and the fluctuations reduction. The improved accuracy and reliability of the proposed scheme position it as a promising approach for addressing the challenges associated with tracking dynamic radio interference sources on high-speed mobile platforms.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Reconfigurable intelligent surface (RIS) technologies can effectively extend the coverage of MIMO systems with improved signal-to-interference-and-noise ratio (SINR). However, the channel estimation of RIS-assisted MIMO systems is very challenging, though crucially important. First, the passive reflective nature of RIS makes it hard to obtain channel state information (CSI) of the transmitter-to-RIS and RIS-to-receiver links, respectively. Second, limited by the number of transmit antennas, reflective elements typically need to be divided into groups, and RIS needs to turn on/off element groups successively to obtain the full CSI, causing extra overhead and difficulty in control. To tackle these problems, we first propose a Two-fold-SVD based Separated Reflecting Channel Estimation (2-SVD-SRCE) scheme, which based on pilot blocks can effectively extract transmitter-to-RIS and RIS-to-receiver links’ CSI, respectively. Then, we propose a novel pilot reconstruction design at the transmitter and develop a Pilot Reconstruction based Separated Reflecting Channel Estimation (PR-SRCE) scheme, which can overcome the limitation that the number of reflective elements in each group cannot exceed transmit antennas. Analyses and simulations show that our proposed scheme can significantly simplify the RIS control procedures, while achieving good performances in both complexity and channel estimating accuracy.
The metaverse, as an envisioned paradigm of the future internet, aims to establish an immersive and multidimensional virtual space in which global users can interact with one another, as in the real world. With the rapid development of emerging technologies—such as digital twins (DT), blockchain, and artificial intelligence (AI)—the diverse potential application scenarios of the metaverse have attracted a great deal of research attention and have created a prosperous market. The demand for ubiquitous communications, pervasive sensing, ultra-low latency computing, and distributed storage has consequently surged, due to the massive heterogeneous devices and data in the metaverse. In order to achieve the metaverse, it is essential to establish an infrastructure system that integrates communications, sensing, computing, and storage technologies. Information about the physical world can be obtained by pervasive sensing, computing resources can be scheduled in a reasonable manner, quick data access can be achieved through the coordination of centralized and distributed storage, and, as the bridge, mobile communications systems connect communications, sensing, computing, and storage in a new system, which is the integration of communications, sensing, computing, and storage (I-CSCS). Following this trend, this paper discusses the requirements of the metaverse for spectrum resources, ultra-reliable transmission, seamless coverage, and security protection in wireless mobile communications systems, and analyzes the fundamental supporting role of the sixth-generation mobile communications system (6G) in the metaverse. Then, we explore the functions and roles of the integrated sensing and communications technologies (ISAC), as well as the integration of communications, computing, and storage technologies for the metaverse. Finally, we summarize the research directions and challenges of I-CSCS in the metaverse.
OTFS is recognized as a promising modulation technology to enhance and extend coverage for high-mobility communications. In this paper, we propose a Modified-Constellation assisted Orthogonal Time-Frequency-Space with Index Modulation (MCOTFS-IM) scheme to improve BER performance over doubly selective fading channels. The proposed scheme optimizes the positions of QAM constellation points that are close to the origin, such that the Euclidean distance for detecting active or inactive status of a grid point in Delay-Doppler domain can be enlarged. This adjustment offers a more accurate detection rate, thus improving the overall BER across index bits and information bits for index modulation. We further apply the Jacobi preconditioning Conjugate gradient (Jac-CG) method to achieve performance of MMSE detection with relatively low complexity. Simulation results demonstrate that the proposed MCOTFS-IM outperforms OTFS-IM in terms of BER under various system setups.
As drones become increasingly prevalent in human life, they also raise security concerns such as unauthorized access and control, as well as collisions and interference with manned aircraft. Therefore, ensuring the ability to accurately detect and identify between different drones holds significant implications for coverage extension. Assisted by machine learning, radio frequency (RF) detection can recognize the type and flight mode of drones based on the sampled drone signals. In this paper, we first utilize Short-Time Fourier Transform (STFT) to extract two-dimensional features from the raw signals, which contain both time-domain and frequency-domain information. Then, we employ a Convolutional Neural Network (CNN) built with ResNet structure to achieve multi-class classifications. Our experimental results show that the proposed ResNet-STFT can achieve higher accuracy and faster convergence on the extended dataset. Additionally, it exhibits balanced performance compared to other baselines on the raw dataset.
Physical layer security (PLS) of wireless communications is becoming increasingly significant with the advent of 6G. The reconciliation of channel-based secret key generation (SKG) methods is necessary to avoid inconsistencies after key generation, resulting in increased communication overhead, leakage of key information, and reduced security performance. This paper proposes a fuzzy secret key generation (FSKG) method that does not require the legitimate user to perform a consistency check after fuzzy extraction of the phase, which is used directly to encrypt the constellation rotation of the key. The proposed scheme allows users to code-decode the key bits to reduce the transmission error probability and improve the key consistency, and the simulation results prove that the proposed FSKG scheme outperforms traditional SKG schemes like CQA and CQG. The simulation results also indicate that the best secure key rate can be achieved by applying PSK modulation.
This paper presents an effective method for strengthening the discriminative ability of high-level deep features by enhancing and aggregating discriminative part-level features for the fine-grained vehicle recognition task. In general, the task of visual recognition concentrates more on the visual differences at the object level. However, for fine-grained object recognition, the visual differences between target objects typically exist in local discriminative areas, so it is more concerned about extracting fine-grained features from these part regions. In this context, we propose solving this issue with a novel feature extraction method from two perspectives: the generation of more feature descriptors of part regions through the learning process of deep networks and the aggregation of part-level discriminative features. This approach is designed to improve the backbone networks to generate finer-level part features through a part-level feature enhancement module and to investigate the intrinsic part-level features of the backbone networks with the help of a feature aggregation module. The enhancement module efficiently finds the finer features highly correlated to the part regions. Then the feature aggregation module builds correlations of similar part features through feature grouping and fusion. Moreover, our proposed method does not require additional parts annotations and achieves comparable performance on two widely-used benchmarks for recognizing fine-grained vehicle types. Experimental results and explainable visualizations demonstrate the effectiveness of the proposed method.
In this paper, we propose a novel part-level feature extraction method to enhance the discriminative ability of deep convolutional features for the task of fine-grained vehicle recognition. Generally, the challenges for fine-grained vehicle recognition are mainly caused by the subtle visual differences between part regions of vehicles. Therefore, it is essential to extract discriminative features from part regions. Many existing methods, especially deep convolutional neural networks (D-CNNs), tend to detect the discriminative part regions explicitly or learn the part information implicitly through network restructuring and neglect the abundant part-level information contained in the high-level features generated by CNNs. In light of this, we propose a simple and effective part-level feature extraction method to enhance the representation of part-level features within the global features of target object generated by the backbone networks. The proposed method is built on the deep convolutional layers from which the discriminative part features could be integrated and extracted accordingly. More specifically, a basic feature grouping module is adopted to integrate the feature maps of deep convolutional layers into groups in each of which the related discriminative parts are assembled. The feature grouping process is performed in a multi-stage manner to ensure the integration process. Then a fusion module follows to model the coarse-to-fine relationship of the part features and further ensure the integrity and effectiveness of the part features. We conduct comparison experiments on public datasets, and the results show that the proposed method achieves comparable performance with state-of-the-art algorithms.
The development of 6G communication is now putting higher performance requirements on the mobile satellite system. How to achieve comprehensive coverage by adding satellites to mobile communication system has become a popular topic on integrated satellite-ground network. However, using traditional methods to analyze coverage performance is limited by the topology of the satellite constellation. In this paper we use the stochastic geometry to model the LEO constellation. The stochastic geometry analysis method weakens the influence of constellation topology and provides a new representation of the coverage performance characteristics. This paper considers the models of beam coverage angle and atmospheric attenuation on coverage performance and proposes a new interference simplification method. It provides a new tool for the future optimization analysis of coverage performance of LEO satellite constellation and gives a new general expression for the coverage probability. The simulation shows that the coverage analysis model established in this paper can clearly represent the characteristics of the variation of the coverage probability with different parameters, which is conforms to the changing trends of the actual satellite constellation. For example, it can obtain the number of satellites with optimal coverage probability at different orbital altitudes.
In this paper, we propose a novel light-weight feature integration and fusion method to enhance the discriminative ability of deep convolutional features for the task of fine-grained vehicle recognition. The proposed method is built on the deep convolutional layers from which the discriminative part features could be integrated and fused accordingly. More specifically, a basic feature integration module is adopted to integrate the feature maps of deep convolutional layers into groups in each of which the related discriminative parts are assembled together. Then a fusion module follows to model the coarse-to-fine relationship of the part features and further ensure the integrity and effectiveness of the part features. We conduct comparison experiments on public dataset, and the results show that the proposed method achieves comparable performance with state-of-the-art algorithms.
Ambient backscatter technology is an important technology of the Internet of Things (IoT), which can use ambient radio frequency (RF) sources to communicate between passive devices. However, most studies on the signal detection problem of ambient backscatter communication are based on the fact that tags have two backscatter states, reflective and non-reflective. In this paper, we present a signal detection scheme for ambient backscatter systems which the tag can backscatter in three states. First, we design a maximum a posteriori (MAP) detector to detect the multiple symbols sent by the tag using the continuously received signals. Then we propose a coding scheme that reduces the bit error rate (BER) and improves the system throughput. Furthermore, we derive expressions of detection threshold and BER, and give the system throughput under different coding schemes. At last, experimental simulations are conducted to prove our study.
Security has become a critical aspect of 5G ultra-low-latency communications (URLLC) due to the openness and vulnerabilities of wireless system information. By tampering/jamming pilots or reference signals in initial access, denial of service (DoS) attacks can disturb uplink access authentication and paralyse the normal data processing in URLLC. We in this paper propose a semi-random coding strategy inspired by amplitude amplification in quantum domain to encode and decode pilot information on multidimensional resources, such that the uncertainty of attacks can be eliminated. In particular, each of pilot signals is encoded as a unique non-random cover-free codeword but randomly transmitted across $N$ subcarriers without assuming any prior distribution while pilot decoding is done to identify each pilot by the receiver who though only observes the superposition of these codewords. We find that the key of pilot decoding lies in how to identify and eliminate the codeword from attacker. It can be proved that the identification process is always equivalent to finding a unique solution to black-box model with $N$ distinguishable binary inputs. By implementing quantum algorithm on this model, we can quickly determine the result and also precisely portray the characteristics of computing performance in a semi-random environment. A novel expression of failure probability of this URLLC system with short packet transmission is also derived to characterize the reliability performance. Numerical results show how our proposed scheme can maintain high reliability and low latency under pilot-aware attack even without knowing the prior distribution for the attack.
Fine-grained vehicle categorization has evolved into a significant subject of study due to its importance in the Intelligent Transportation System. A highly accurate and real-time vehicle categorization system will help to support many applications not only in the security aspect but also many walks of life. In this paper, facing the growing importance of this study, we present an image dataset named Frontal-103 to promote the development of the vision-based research on the vehicle, and particularly for the task of fine-grained vehicle categorization. This paper provides a detailed analysis of Frontal-103 in its current state: 1,759 fine-grained vehicle models in 103 vehicle makes and 65,433 web-nature images in total. Apart from the specific viewpoint and vehicle hierarchy, Frontal-103 is superior to the other state-of-the-art vehicle image datasets not only in the scale and diversity but also the accuracy and fine-grained level. We further discuss the peculiar challenges and issues lies in the task of fine-grained vehicle categorization and illustrate the usefulness of our dataset in addressing those problems. We hope Frontal-103 will be beneficial to the vision-based vehicle analysis and contribute to the computer vision community.
In a virtualized environment, it is not difficult to retrieve guest OS information from its hypervisor. However, it is very challenging to retrieve information in the reverse direction, i.e., retrieve the hypervisor information from within a guest OS, which remains an open problem and has not yet been comprehensively studied before. In this paper, we take the initiative and study this reverse information retrieval problem. In particular, we investigate how to determine the host OS kernel version from within a guest OS. We observe that modern commodity hypervisors introduce new features and bug fixes in almost every new release. Thus, by carefully analyzing the seven-year evolution of Linux KVM development (including 3,485 patches), we can identify 19 features and 20 bugs in the hypervisor detectable from within a guest OS. Building on our detection of these features and bugs, we present a novel framework called Hyperprobe that for the first time enables users in a guest OS to automatically detect the underlying host OS kernel version in a few minutes. We implement a prototype of Hyperprobe and evaluate its effectiveness in six real world clouds, including Google Compute Engine (a.k.a. Google Cloud), HP Helion Public Cloud, ElasticHosts, Joyent Cloud, CloudSigma, and VULTR, as well as in a controlled testbed environment, all yielding promising results.
This paper presents a novel feature representation and recognition scheme for vehicle make and model recognition (VMMR) from the frontal image of the vehicle. In general, some domain knowledge is introduced and further exploited to accomplish this task. The modular components in the frontal appearance of vehicles present distinct visual characteristics, and the varying discrimination ability of them made the main discriminant region changes when vehicles compared at inter- or intra-brand level. Inspired by the peculiar vehicle properties, this paper focuses on the localized visual characteristics in the discriminant subregions, and make the representation of the frontal vehicle in a multi-scale spatial manner. This representation scheme encodes local spatial information and component-specific characteristics to the feature descriptors, which enhances the discrimination ability of the feature representation and mitigates the multiplicity problem of VMMR to some extent. Extensive experiments on a large-scale vehicle image dataset have shown the efficiency of the methods in this work.