In disaster scenarios, unmanned aerial vehicle (UAV)-based remote sensing requires low-latency and high-quality image transmission and reconstruction over bandwidth-limited and unstable wireless channels. However, traditional separated communication frameworks are vulnerable to the "cliff effect" and error propagation, and existing deep joint source-channel coding (DeepJSCC) methods, although optimized end-to-end, still suffer from blurred local details, directional texture loss, and a trade-off between pixel fidelity and perceptual naturalness. To overcome these challenges, this paper proposes a disaster recovery-oriented DeepJSCC framework (DRJSCC) that integrates global-local structure perception, multi-directional channel attention, and dual-branch decoding to enhance both semantic robustness and visual quality. Specifically, the global-local structure perception module strengthens structural consistency through a cross-scale "local enhancement-global modeling" pathway, while the multi-directional channel attention mechanism captures orientation-sensitive statistics to restore weak textures such as cracks and shadows. The dual-branch decoder further decouples pixel-level and perceptual objectives to achieve balanced reconstruction. Extensive experiments on xView2, RescueNet, and COCO datasets under AWGN and Rayleigh channels demonstrate that DRJSCC achieves up to 2 dB Peak Signal-to-Noise Ratio (PSNR) improvement and 30 % Learned Perceptual Image Patch Similarity (LPIPS) reduction in the medium-tohigh SNR range while maintaining a 41 FPS inference rate. These results highlight the effectiveness of DRJSCC as a robust and efficient end-to-end paradigm for real-time UAV disaster image transmission and reconstruction.
Continuous Phase Modulation (CPM) signals offer excellent spectral efficiency and constant envelope properties for wireless communications, but traditional detection methods suffer from prohibitive computational complexity. This paper presents CPMNet, a novel deep learning-based detection framework that addresses these limitations through an enhanced residual network architecture incorporating spatial attention mechanisms, multi-scale feature fusion, and bidirectional LSTM networks. CPMNet performs sequence-to-sequence detection without requiring channel estimation or equalization. Experimental results on Advanced Range Telemetry (ARTM) Tier 2 signals show performance varies with modulation complexity: while exhibiting 2-4 dB gaps compared to Maximum Likelihood Sequence Detection (MLSD) in high signal-to-noise ratio (SNR) AWGN channels for lower-order modulations, CPMNet maintains robust performance for high-order modulations where MLSD becomes impractical. In multipath fading channels, CPMNet significantly outperforms MLSD by 3-6 dB across various conditions, demonstrating superior resilience to channel impairments. The framework exhibits excellent generalization with only 1-2 dB degradation in unseen environments. Most critically, CPMNet maintains constant computational complexity regardless of CPM parameters, contrasting sharply with MLSD’s exponential complexity growth, making it particularly advantageous for high-order CPM signals that are computationally prohibitive for traditional methods.
In Computational Ghost Imaging (CGI), the bucket signals play a crucial role, as they capture the encoded information about the object, enabling the reconstruction of images even without traditional detectors. By analyzing the bucket signals, it is possible to classify the target using image-free GI. This paper integrates Long Short-Term Memory (LSTM) networks into CGI, leveraging their gated mechanisms to filter noise, capture sequential features, and extract global object-specific information. The proposed method is evaluated through both simulation and physical experiments. Simulation results show a classification accuracy of 91 % at a sampling rate of 5 %. Additionally, we conducted robustness experiments by introducing Gaussian noise to the input data, under which the LSTM model maintained relatively high accuracy compared to baseline methods. Furthermore, physical experiments validate the feasibility of the approach and demonstrate stable classification performance under real-world conditions, confirming its potential for practical low-sampling, image-free recognition applications.
Automated detection of cervical lesion cell clumps is crucial for cervical cancer screening. However, the dense packing and overlap of cells, caused by adhesion molecules, make detection challenging. To address this issue, we propose the Morphology-Aware Detector (MA-Det). Specifically, by innovatively employing Deformable Convolution, we propose the Dynamic Context Aggregation Module (DCAM) which dynamically captures contextual information, improving the feature representation of lesion cells. To enhance discrimination in high-density regions, we design the Distance-Weighted Interaction Module (DWIM) and Adaptive Morphology-Aware (AMA) loss into the detection head. These components improve spatial awareness by utilizing interactions between cell features and proposals. In particular, our method achieves state-of-the-art performance on the CDetector and CRIC datasets while significantly reducing parameters.
This article comprehensively investigates the scouring mechanism of underwater cement paste (UWP) through parameter calibration, flume erosion testing, numerical simulations, and force chain analysis. Building upon the established ARR constitutive model, concurrent calibration involving flowability and rheological parameter experiments confirms that when fluidity ratio, yield stress, viscosity below 3.5, 7.5, 3.5
To address the high Peak-to-Average Power Ratio (PAPR) issue in Continuous Phase Modulation Orthogonal Frequency Division Multiplexing (CPM-OFDM) systems, which causes power amplifier nonlinear distortion, this paper proposes a novel deep learning network named AF-HDC-CPMNet. Its core is an Attention-enhanced Hybrid Dilated Convolution (AF-HDC) module, which employs coprime dilation rates to capture multi-scale phase dependencies and a branch attention mechanism to dynamically weight features critical for global phase continuity. Additionally, a phase-preserving decoder using nearest-neighbor upsampling is designed to mitigate phase distortion. Simulation results show that the proposed method achieves a significant PAPR reduction to 3.3 dB at a CCDF of $10^{-3}$ while maintaining excellent Bit Error Rate (BER) performance, outperforming traditional and other deep learning-based schemes.
The automated detection of cervical cells plays a vital role in cervical cancer screening. Cancer cells typically exhibit more prominent edge features compared to normal cells. However, during the feature fusion process, blurred edges often lack precise high-frequency information, which hinders the model’s ability to accurately identify cancer cells. To address this challenge, we propose the Hybrid-Domain Feature Pyramid Network (HD-FPN), which enhances the model’s sensitivity to cellular edge information. Specifically, we propose the Frequency-Aware Sampling that preserves critical texture features by leveraging high-frequency decomposition using wavelet transform. Furthermore, we design a multi-domain feature fusion approach to progressively refine edge features and minimize background noise. Our method achieves state-of-the-art performance on the CDetector and CRIC datasets in terms of both detection accuracy and efficiency.
This letter proposes a convolutional neural network-bidirectional long short-term memory (CNN-BiLSTM) architecture for continuous phase modulation (CPM) signal detection by optimising the extraction of time-frequency features and temporal dependencies with reduced complexity. It significantly outperforms the existing maximum likelihood sequence detection (MLSD) and CNN with fully connected layer (CNN-FC) detectors in higher-order modulation and multipath scenarios, achieving a 97.99% parameter reduction compared to CNN-FC. Numerical results confirm its exceptional balance of detection performance and computational efficiency, making it ideal for complex channels and resource-constrained systems.
We propose a novel deep learned video compression technique, named scalable motion estimation (SME), which is designed for video data generated by sensor systems in smart devices. These devices face unique challenges due to limited bandwidth, constrained computational resources, and the complexity of real-world scenarios. Existing learned video compression methods primarily rely on optical flow for motion estimation; however, optical flow often suffers from inaccuracies and displacement errors, resulting in imprecise motion vectors, and the increasing complexity of models further limits their deployment in resource-constrained environments. To address these issues, the SME framework introduces a hierarchical motion vector enhancement strategy that progressively optimizes motion information in video captured by sensors. In addition, we propose the temporal context reinforcement (TCR) module, which partitions high-dimensional contextual information into three components. Compared to direct context encoding methods, the TCR module optimizes bit allocation, achieving a more balanced distribution of high- and low-frequency information. Extensive experiments on benchmark datasets demonstrate that the SME framework substantially enhances the efficiency and performance of video compression systems on RGB sensor-captured videos, making it particularly suitable for video applications.
Semantic communication as a key proposal in 6 G communication, aims to transmit information at the semantic level. Semantic communication system can perceive the deep semantic of transmitted information to enhance the efficiency of information transmission. Meanwhile, low-altitude economy dominated by unmanned aerial vehicles is emerging. Using unmanned aerial vehicle laser radar to acquire point cloud data for scene construction has become a popular research direction. In this study, we propose a method that can transform point-cloud data into rasterized image and reconstruct it on the receiver based on joint source-channel coding. It can achieve the rapid end-to-end reconstruction of point-cloud rasterized images.
Learned video compression (LVC) methods typically align spatial-temporal transformation features with optical flow to perform motion compensation. However, existing LVC methods typically rely on a single reference feature to provide local detail features. This ignores global structural information, resulting in limited capabilities when dealing with fast motion or occlusion scenarios. In addition, with the encoding of P-frames, errors accumulate, leading to a degradation in reconstruction quality. To address these issues, we propose the Neighbor-Aware Feature-Driven Motion Compensation (NAFD-MC) that utilizes the spatial-temporal correlations of the neighboring features to explore the global structural information and the local detail information. Furthermore, we introduce the Synergy Filtering Module (SFM) to enhance inter-frame consistency and alleviate the error accumulation. Experimental results demonstrate that our method outperforms the H.266/VVC reference software VTM-13.2 in public benchmark datasets.
In the past few years, learned video compression has received increasing attention. However, most current learned methods rely on the bilinear warping operations in motion compensation, which is equivalent to a low-pass filtering operation and brings the distortion of reconstruction. To address this problem, we propose a spatial-temporal motion compensation for learned video compression (STMC-LVC). STMC-LVC uses a spatial-temporal motion compensation network (STMC-Net) to fully consider the spatial-temporal correlation between successive frames, performing accurate motion compensation. Specifically, STMC-Net mainly consists of the initial feature prediction module and the fusion module. The initial feature prediction module predicts reference features based on DCN to provide sufficient information for motion compensation. The fusion module obtains temporal attention information by calculating the similarity between features, performs spatial attention operations, and finally performs spatial-temporal fusion, further improving the robustness of motion compensation. In addition, STMC-LVC uses a conditional coding framework. We use a concatenation of the current feature and predicted feature as context to explore spatial-temporal correlations in feature space. Experimental results show that our method effectively improves video compression. Specifically, our model achieves an average of 33.86% and 58.70% bitrate savings than x265 (veryslow) on PSNR and MS-SSIM, respectively.
In the era of sixth generation mobile networks (6G), industrial big data is rapidly generated due to the increasing data-driven applications in the Industrial Internet of Things (IIoT). Effectively processing such data, for example, knowledge learning, on resource-limited IIoT devices becomes a challenge. To this end, we introduce a cloud-edge-end collaboration architecture, in which computing, communication, and storage resources are flexibly coordinated to alleviate the issue of resource constraints. To achieve better performance in hyper-connected experience, real-time communication, and sustainable computing, we construct a novel architecture combining digital twin (DT)-IIoT with edge networks. In addition, considering the energy consumption and delay issues in distributed learning, we propose a deep reinforcement learning-based method called deep deterministic policy gradient with double actors and double critics (D4PG) to manage the multi-dimensional resources, that is, CPU cycles, DT models, and communication bandwidths, enhancing the exploration ability and improving the inaccurate value estimation of agents in continuous action spaces. In addition, we introduce a synchronization threshold for distributed learning framework to avoid the synchronization latency caused by stragglers. Extensive experimental results prove that the proposed architecture can efficiently conduct knowledge learning, and the intelligent scheme can also improve system efficiency by managing multi-dimensional resources. This article proposes a multi-dimensional resource management method for digital twin-enabled Industrial Internet of Things with edge networks to achieve hyper-connected experience, real-time communication, and sustainable computing. The proposed method, named D4PG, improves the traditional deep deterministic policy gradient with double actors and double critics to enhance the exploration ability and the inaccurate value estimation of each agent. Extensive experiments in a distributed test-bed are conducted to evaluate the proposed architecture, and the numeric results show that the proposed methods outperform the others. image
The unique nature of the high frequencies in the terahertz band (0.1-10 terahertz) makes it impractical to use lower frequency channel models. Therefore, there is an urgent need to develop new channel models to realize the potential of terahertz communications, and a variety of channel measurement methods and measurement environments can accelerate development efforts. Since Intelligent Reflective Surface (IRS) is a potential solution to improve the cost-effectiveness and energy efficiency of sixth generation (6G) wireless communication systems. Therefore, this paper presents a method to implement IRS functionality in a 3D ray-tracing simulator and applies it to a novel indoor T-corridor scenario with no line-of-sight path (NLoS) between the transmitter (Tx) and receiver (Rx). High-resolution 3D measurements are processed at 300 GHz to obtain bidirectional angular delay power spectra as well as delay and angular power distributions, and finally the channel is characterized by path loss. Numerical results show that the incorporation of IRS into wireless communication systems in this scenario can significantly improve the coverage and provide useful insights for optimizing the deployment of IRS in indoor T-corridor scenarios.
Reconfigurable intelligent surfaces (RIS) have emerged as a crucial technology for sixth-generation (6G) communication systems. Mastering the modeling and analysis methods of RIS is essential for its effective implementation. In this paper, a channel reconstruction model is proposed based on ray tracing (RT) simulations in an indoor corridor environment at 300 GHz. It can reconstruct the extracted multipath components (MPCs) to enhance their accuracy and alignment with real-world measurement environments. By analyzing the power delay profile (PDP), power angular profile (PAP), and path loss (PL) before and after integrating RIS, the accuracy and effectiveness of the proposed method are verified. It has also confirmed that RIS significantly improves the transmission efficiency and extends the signal coverage of communication systems.
In cognitive wireless networks, secondary users gain access to the spectrum by identifying and occupying primary users’ frequency resources. However, Primary User Emulation Attack (PUEA) has become a common method of attack in cognitive wireless networks. In a PUEA, malicious users try to deceive secondary users by mimicking the signals of primary users, thereby preventing them from utilizing idle frequency resources. In this paper, we propose a channel difference-based energy detection method to defend against PUEA, which takes into account the distance differences between the primary user, attacker, and secondary user. In the proposed scheme, the fusion center does not discriminate directly based on the received signal power, but rather based on the difference between the signal power received by each user and the average of all received signal powers. Closed-form expressions for the detection probability and false alarm probability of the proposed scheme are derived, and we analyze the impact of the minimum distance between the simulated primary user attacker and the secondary user, as well as the number of cooperative users, on the detection performance. Numerical results demonstrated the effectiveness of the proposed scheme in defending against PUEA.
Attributional network anomaly detection is increasingly becoming a focal point of academic research due to its applications in critical domains like social networking and financial fraud. This task faces many challenges due to the differences in attributes between anomalous nodes and others, as well as their complex interactions. While existing shallow methods fall short in concurrently addressing both attributes and structures, leading to suboptimal performance, methods based on deep learning have made significant advances in enhancing anomaly detection. However, they still lack sufficient exploitation of anomaly information. In response to the aforementioned issue, this paper introduces an attribute graph node anomaly detection method based on training strategy optimization (ANATSO). The model constructs a secondary view through edge perturbation and samples subgraphs from each view. These subgraphs are then used to initialize networks for subgraph-node and node-node instance pairs. Additionally, a phased training process is implemented, where the outcomes from the initial phase are used to optimize the node input processing in the subsequent phase, ultimately calculating anomaly scores for each node. This research was evaluated on five benchmark datasets, and the results demonstrate that compared to traditional baseline methods, our model significantly improves the precision of anomaly detection.
Accurate urban PM2.5 forecasting serves a crucial function in air pollution warning and human health monitoring. Recently, deep learning techniques have been widely employed for urban PM2.5 forecasting. Unfortunately, two problems exist: (1) Most techniques are focused on training and prediction on a central cloud. As the number of monitoring sites grows and the data explodes, handling a large amount of data on the central cloud can cause tremendous computational pressures and increase the risk of data leakages. (2) Existing methods lack an adaptive layer to capture the varying impacts of different external factors (e.g., weather conditions, temperature, and wind speed). In this paper, a federated deep learning network (FedDeep) is developed for edge-assisted multi-urban PM2.5 forecasting. First, we assign each urban region to an edge cloud server (ECS). An external spatio-temporal network (ESTNet) is then deployed on each ECS. Data from different urban regions are uploaded to the corresponding ECS for training, which avoids processing all the data on the central cloud and effectively alleviates computational pressure and data leakage issues. Second, in ESTNet, we develop a gating fusion layer to adaptively fuse external factors to improve prediction accuracy. Finally, we adopted PM2.5 data collected from air quality monitoring sites in 13 prefecture-level cities, Jiangsu Province for validation. The experimental results proved that FedDeep outperformed the advanced baselines in terms of prediction accuracy and model efficiency.
Background and Objectives:In cervical cell diagnostics, autonomous screening technology constitutes the foundation of automated diagnostic systems. Currently, numerous deep learning-based classification techniques have been successfully implemented in the analysis of cervical cell images, yielding favorable outcomes. Nevertheless, efficient discrimination of cervical cells continues to be challenging due to large intra-class and small inter-class variations. The key to dealing with this problem is to capture localized informative differences from cervical cell images and to represent discriminative features efficiently. Existing methods neglect the importance of global morphological information, resulting in inadequate feature representation capability.Methods:To address this limitation, we propose a novel cervical cell classification model that focuses on purified fusion information. Specifically, we first integrate the detailed texture information and morphological structure features, named cervical pathology information fusion. Second, in order to enhance the discrimination of cervical cell features and address the data redundancy and bias inherent after fusion, we design a cervical purification bottleneck module. This model strikes a balance between leveraging purified features and facilitating high-efficiency discrimination. Furthermore, we intend to unveil a more intricate cervical cell dataset: Cervical Cytopathology Image Dataset (CCID).Results:Extensive experiments on two real-world datasets show that our proposed model outperforms state-of-the-art cervical cell classification models.Conclusions:The results show that our method can well help pathologists to accurately evaluate cervical smears.