
As human exploration of the ocean expands, the demand for continuous, high-quality, and ubiquitous maritime communication is steadily increasing. However, the dynamic nature of the marine environment and resource constraints present significant challenges for traditional heuristic resource allocation methods, complicating the balance between high-quality communication and limited network resources. This results in suboptimal system throughput and an over-reliance on specific problem structures. To address these issues, in this paper, we introduce a joint resource allocation method based on knowledge embedding. The proposed approach includes an action distribution alignment module designed to improve resource utilization by preventing unreasonable action-output combinations. Furthermore, by integrating knowledge embedding with meta-reinforcement learning techniques, a physical guidance loss function is formulated, which effectively reduces the sample size required for model training, thereby enhancing the algorithm's generalization capabilities. Simulation results show that the proposed method achieves an increase in average system throughput of 31.19% compared to the model-agnostic meta-learning proximal policy optimization (MAML-PPO) algorithm and 80.91% compared to the RL2 algorithm, across various channel environments.
Siting low-altitude takeoff and landing platforms (vertiports) is a fundamental challenge for developing urban air mobility (UAM). This study formulates this issue as a variant of the capacitated facility location problem, incorporating flight range and service capacity constraints, and proposes SPID, a deep reinforcement learning (DRL)-based solution framework that models the problem as a Markov decision process. To handle dynamic coverage, the designed DRL framework-based SPID uses a multi-head attention mechanism to capture spatiotemporal patterns, followed by integrating dynamic and static information into a unified input state vector. Afterward, a gated recurrent unit (GRU) is used to generate the query vector, thereby enhancing sequential decision-making. The action network within the DRL network is regulated by a loss function that integrates service distance costs with unmet demand penalties, enabling end-to-end optimization. Subsequent experimental results demonstrate that SPID significantly enhances solution efficiency and robustness compared with traditional methods under flight and capacity constraints. Especially, across the social performance metrics emphasized in this study, SPID outperforms the suboptimal solutions produced by traditional clustering and graph neural network (GNN)-based methods by up to approximately 29%. This improvement comes with an increase in distance-based cost that is kept within 10%. Overall, we demonstrate an efficient, scalable approach for vertiport siting, supporting rapid decision-making in large-scale UAM scenarios.
The rapid advancement of the low-altitude economy (LAE) necessitates a fundamental shift from fragmented systems toward deeply integrated communication, sensing, navigation, and control capabilities. To this end, this paper proposes a low-altitude digital-intelligent network (LADIN) as an overarching architecture, with integrated sensing and communication (ISAC) serving as the core enabling technology that pervasively unifies its three layers. At the heterogeneous infrastructure layer, we detail an ISAC waveform design based on orthogonal frequency division multiplexing, enabling dual-purpose hardware to simultaneously achieve high-speed data transmission and high-precision environmental sensing. Within the intelligent data fusion layer, ISAC’s role expands into a multimodal fusion paradigm, providing the crucial electromagnetic sensing modality. This layer constructs a unified spatiotemporal feature space by introducing pluggable back-projection adapters and spatiotemporal modeling. These adapters systematically integrate heterogeneous data from ISAC, optical cameras, and light detection and ranging (LiDAR) by inverting their respective observation models, thereby overcoming representational disparities and association ambiguities. At the service and management layer, this coherent representation directly drives algorithmic processes and control policies. ISAC resources are virtualized into dynamically allocable assets, enabling closed-loop control that responds to the real-time state of the feature space, such as reconfiguring base station operational modes based on live situational awareness. Validation through multi-frequency collaborative sensing and multimodal fusion use cases demonstrates significant performance gains in tracking robustness, detection of near-zero radar cross-section targets such as balloons, and seamless urban airspace governance, conclusively establishing the transformative potential of a deeply integrated, ISAC-centric approach for future LAE systems.
In recent years, multi-label zero-shot learning (ML-ZSL) has garnered increasing attention because of its wide range of potential applications, such as image annotation, text classification, and bioinformatics. The central challenge in ML-ZSL lies in predicting multiple labels for unseen classes without requiring any labeled training data, which contrasts with conventional supervised learning paradigms. However, existing methods face several significant challenges. These include the substantial semantic gap between different modalities, which impedes effective knowledge transfer, and the intricate and typically complex relationships among multiple labels, making it difficult to model them in a meaningful and accurate manner. To overcome these challenges, we propose a graph-augmented multimodal chain-of-thought (GMCoT) reasoning approach. The proposed method combines the strengths of multimodal large language models with graph-based structures, significantly enhancing the reasoning process involved in multi-label prediction. First, a novel multimodal chain-of-thought reasoning framework is presented which imitates human-like step-by-step reasoning to produce multi-label predictions. Second, a technique is presented for integrating label graphs into the reasoning process. This technique enables the capture of complex semantic relationships among labels, thereby improving the accuracy and consistency of multi-label generation. Comprehensive experiments on benchmark datasets demonstrate that the proposed GMCoT approach outperforms state-of-the-art methods in ML-ZSL.
Multi-aircraft task allocation (MATA) plays a vital role in improving mission efficiency under dynamic conditions. This paper proposes a novel coevolutionary genetic programming (CoGP) framework that automatically designs high-performance reactive heuristics for dynamic MATA problems. Unlike conventional single-tree genetic programming (GP) methods, CoGP jointly develops two interacting populations, i.e., task prioritizing heuristics and aircraft selection heuristics, to explicitly model the coupling between these two interdependent decision phases. A comprehensive terminal set is constructed to represent the dynamic states of aircraft and tasks, whereas a low-level heuristic template translates developed trees into executable allocation strategies. Extensive experiments on public benchmark instances simulating post-disaster emergency delivery demonstrate that CoGP achieves superior performance compared with state-of-the-art GP and heuristic methods, exhibiting strong adaptability, scalability, and real-time responsiveness in complex and dynamic rescue environments.
This study proposes a method for analyzing synchronization in oscillator systems, illustrated by modeling the dynamics of a circuit of two resistively coupled pulse oscillators. The dynamic characteristic of synchronization is the fuzzy entropy (FuzzyEn), which is calculated from a time series composed of the ratios of the number of pulse periods (subharmonic ratio, SHR) at phase-locking intervals. Low and high entropy values indicate strong and weak synchronization between the two oscillators, respectively. The proposed method effectively visualizes synchronized modes of the circuit using entropy maps of synchronization states. In addition, a classification of synchronization states is proposed based on the dependency of FuzzyEn on the embedding vector length of the SHR time series. An extension of this method for analyzing non-pulse (non-spike) signals is demonstrated using the example of phase-phase coupling rhythms of the local field potential of the rat hippocampus. The proposed entropy-statistical approach, using integers and pulse signal forms, is well-suited for signal synchronization analysis and can be implemented on digital mobile platforms.
The market of low-altitude economy has the po-tential to reach trillion dollars in 10 years globally.In China,it serves as a hallmark of national strategic emerging industries,and represents new quality pro-ductive forces.Exploring innovative engineering and technologies for low-altitude economy infrastructure is expected to promote sustainable growth in this sector.
The booming of artificial intelligence (AI) agents has brought about promising business scenarios for sixth-generation (6G) mobile networks, while simultaneously posing significant challenges to network functionalities and infrastructure. These AI agents can be deployed on end devices (e.g., intelligent robots and intelligent cars) or as digital entities (e.g., personal AI assistants). As novel service entities with autonomous decision-making and task execution capabilities, AI agents introduce potential risks of uncontrollable actions and privacy disclosures. AI agents also require new 6G capabilities beyond traditional communication, including multimodality information interaction (e.g., AI models and tokens) and support for service requirements (e.g., computing and sensing of data). In this article, we introduce the concept of AI-agent communication network (ACN), a new paradigm to enable global information interaction and on-demand capability provisioning for single or multiple AI agents. We first introduce the vision and architectural framework of ACN. Then, key technologies and future research directions related to ACN are discussed. Furthermore, we provide potential use cases to elaborate on how ACN can expand the service capabilities of 6G networks.
Robotic navigation in unknown environments is challenging due to the lack of high-definition maps.Building maps in real time requires significant computational resources.Nevertheless,sensor data can provide sufficient environmental context for robots' navigation.This paper presents an interpretable and mapless navigation method using only two-dimensional(2D)light detection and ranging(LiDAR),mimicking human strategies to escape from dead ends.Unlike traditional planners,which depend on global paths or vision-based and learning-based methods,requiring heavy data and hardware,our approach is lightweight and robust,and it requires no prior map.It effectively suppresses oscillations and enables autonomous recovery from local minimum traps.Experiments across diverse environments and routes,including ablation studies and comparisons with existing frameworks,show that the proposed method achieves map-like performance without a map-reducing the average path length by 50.51%when compared to the classical mapless Bug2 algorithm and increasing it by only 17.57%when compared to map-based navigation.
In recent years, physics-informed neural networks (PINNs) have shown remarkable potential in modeling conservative systems of rigid-body dynamics. However, when applied to practical interaction tasks of manipulators (e.g., part assembly and medical operations), existing PINN frameworks lack effective external force modeling mechanisms, resulting in significantly degraded prediction accuracy in dynamic interaction scenarios. Additionally, because industrial robots (including UR5 and UR10e robots) are generally not equipped with joint torque sensors, obtaining precise dynamics training data remains challenging. To address these issues, this study proposes two enhanced PINNs that integrate motor dynamics and external force modeling. First, two data-driven Jacobian matrix estimation methods are introduced to incorporate external forces: one learns the mapping between end-effector velocity and joint velocity to approximate the Jacobian matrix, while the other first learns the system's kinematic behavior and then derives the Jacobian matrix through analytical differentiation of the forward kinematics model. Second, current-to-torque mapping is embedded as physical prior knowledge to establish direct correlations between system motion states and motor currents. Experimental results on two different manipulators demonstrate that both models achieve high-precision torque estimation in complex external force scenarios without requiring joint torque sensors. Compared with state-of-the-art methods, the proposed models improve overall modeling accuracy by 31.12% and 37.07% on average across various complex scenarios, while reducing joint trajectory tracking errors by 40.31% and 51.79%, respectively.
Millimeter-wave (mmWave) communication is the key to increasing the demand for high data rates and low latency resulting from the rapid evolution of wireless communications, especially in the fifth generation (5G) of wireless communication systems and beyond. The mmWave band suffers from high path loss and obstacle blockage, significantly reducing the transmission range. Note that high-directional beams are required to perform well in the mmWave band. Hence, beam alignment is crucial for high-data-rate transmission between the transmitter (Tx) and the receiver (Rx). One of the drawbacks is getting an accurate beam alignment when the transceiver (Tx, Rx, or both) is mobile. Beam tracking plays a considerable role in 5G communications, especially in vehicular communications, due to the repeated change of the transceiver (Tx, Rx, or both) position. This work presents an overview of the different beam-tracking methods used in mmWave communications, focusing on hybrid beamforming techniques. We also compare the various tracking techniques in a recommendation table. This overview suggests that some tracking methods used in the sub-6-GHz band, such as least mean squares (LMS), recursive least squares (RLS), and Kalman filter, are unsuitable for the mmWave band (due to higher frequency and shorter coherence time), and it recommends faster tracking strategies.
Inferring protocol state machines from observable information presents a significant challenge in protocol reverse engineering (PRE), especially when passively collected traffic suffers from message loss, resulting in an incomplete protocol state space. This paper introduces an innovative method for actively inferring protocol state machines using the minimally adequate teacher (MAT) framework. By incorporating session completion and deterministic mutation techniques, this method broadens the range of protocol messages, thereby constructing a more comprehensive input space for the protocol state machine from an incomplete message domain. Additionally, the efficiency of active inference is improved through several optimizations for the LM+ algorithm, including traffic deduplication, the construction of an expanded prefix tree acceptor (EPTA), query optimization based on responses, and random counterexample generation. Experiments on the real-time streaming protocol (RTSP) and simple mail transfer protocol (SMTP), which use Live555 and Exim implementations across multiple versions, demonstrate that this method yields more comprehensive protocol state machines with enhanced execution efficiency. Compared to the LM+ algorithm implemented by AALpy, Act_Infer achieves an average reduction of approximately 40.7% in execution time and significantly reduces the number of connections and interactions by approximately 28.6% and 46.6%, respectively.
Signal-to-noise ratio (SNR) fluctuations significantly affect spectrum sensing performance in wireless communications. Traditional convolutional neural network (CNN) exhibits limited feature extraction capabilities and inefficient feature utilization at low SNR levels, leading to suboptimal spectrum sensing performance. This paper proposes a spectrum sensing method based on a multi-scale feature fusion network (MSFFNet) to address this issue. First, the proposed method employs a multi-scale feature extraction block (MSFEB) to capture multi-scale information from the input data comprehensively. Next, an adaptive feature screening strategy (AFSS) highlights key features while suppressing redundant information. Finally, a multi-level feature fusion mechanism (MLFFM) optimizes and integrates features across scales and levels, enhancing spectrum sensing performance. Simulation results demonstrate that compared to other methods, the proposed approach achieves superior performance in low-SNR communication scenarios. At an SNR of -14 dB, the detection probability Pd reaches 0.936, while the false alarm probability Pfa is only 0.1. Furthermore, this paper constructs a multi-level mixed-SNR dataset to simulate real communication environments and enhance the robustness of spectrum sensing.
This paper addresses the urgent need to detect low, slow, and small (LSS) unmanned aerial vehicles (UAVs) in complex and critical environments, proposing an active low-altitude target detection method based on the cat's eye effect. The detection system incorporates a control module, a laser emission component, a co-optical path panoramic scanning optical mechanism structure, an echo reception component, target detection, and visualization processing to achieve small target detection. The light source is emitted by a near-infrared laser, and the scanning optical path is realized using micro-electro-mechanical system (MEMS) mirrors and servo mechanisms. The echo reception signal is received by an avalanche photodiode (APD) and the target detection module, which captures the reflected signal and distance information. The detection software integrates the local pyramid attention (LPA) module and the field pyramid network (FPN) through the UAV micro lens identification algorithm. It eliminates false alarms by incorporating SKNet21 and uses the APD to collect echo intensity and flight time, thereby reducing the false alarm rate. The results demonstrate the feasibility of the proposed target detection method, which achieves a mean average precision of 0.809 at an intersection over union (IoU) of 0.50, a mean average precision of 0.324 at an IoU of 0.50-0.95, and a throughput of 49.8 Giga floating-point operations per second (GFLOPs), indicating that it can address the current limitations in LSS target detection.
Medical image segmentation is critical for clinical diagnosis, but the scarcity of annotated data limits robust model training, making few-shot learning indispensable. Existing methods often suffer from two issues-performance degradation due to significant inter-class variations in pathological structures, and overreliance on attention mechanisms with high computational complexity (O(n2)), which hinders the efficient modeling of long-range dependencies. In contrast, the state space model (SSM) offers linear complexity (O(n)) and superior efficiency, making it a key solution. To address these challenges, we propose PPFFR (parallel prototype filter and feature refinement) for few-shot medical image segmentation. The proposed framework comprises three key modules. First, we propose the prototype refinement (PR) module to construct refined class subgraphs from encoder-extracted features of both support and query images, which generates support prototypes with minimized inter-class variation. We then propose the parallel prototype filter (PPF) module to suppress background interference and enhance the correlation between support and query prototypes. Finally, we implement the feature refinement (FR) module to further enhance segmentation accuracy and accelerate model convergence with SSM's robust long-range dependency modeling capability, integrated with multi-head attention (MHA) to preserve spatial details. Experimental results on the Abd-MRI dataset demonstrate that FR with MHA outperforms FR alone in segmenting the left kidney, right kidney, liver, and spleen, and in terms of mean accuracy, confirming MHA' s role in improving precision. In extensive experiments conducted on three public datasets under the 1-way 1-shot setting, PPFFR achieves Dice scores of 87.62%, 86.74%, and 79.71% separately, consistently surpassing state-of-the-art few-shot medical image segmentation methods. As the critical component, SSM ensures that PPFFR balances performance with efficiency. Ablation studies validate the effectiveness of the PR, PPF, and FR modules. The results indicate that explicit inter-class variation reduction and SSM-based feature refinement can enhance accuracy without heavy computational overhead. In conclusion, PPFFR effectively enhances inter-class consistency and computational efficiency for few-shot medical image segmentation. This work provides insights for few-shot learning in medical imaging and inspires lightweight architecture designs for clinical deployment.
Avoidance of topic drift and enabling crossing tunnels are two main difficulties in focused crawling. To overcome the problem of topic drift, we design a comprehensive priority evaluation (CPE) method based on the web text, anchor text, and context of hyperlinks, which improves the topic-relevance evaluation of unvisited hyperlinks. Subsequently, we propose an improved Bayesian classifier with weights (BCW), which adds label weights to the feature words of the Bayesian classifier to enhance the accuracy of webpage classification. To cross tunnels through which some topic-relevant webpages can be reached from low-relevance webpages, we construct a content block segmentation (CBS) technology for webpages based on the backtracking method, which segments a webpage into multiple blocks and then judges the relevance of every content block, extracting hyperlinks with high comprehensive relevance. Finally, a BCW-based focused crawling strategy combining the CPE and CBS strategies (BCW_CC) is proposed and experimentally evaluated for focused crawling in two domains: rainstorm disasters and sports. The results demonstrate the effectiveness of the developed BCW_CC method.
Track-to-track association (T2TA), which aims at unifying track batch numbers and reducing track redundancy, serves as a precondition and foundation for track fusion and situation awareness. The current problems of T2TA come mainly from two sources: track data and association methods. Ubiquitous problems include errors and inconsistent update periods in track data, as well as suboptimal association results and dependencies on prior information and assumed motion models for association methods. Focusing on these two aspects, we propose a multiple-hypothesis algorithm for multi-sensor T2TA with an intelligent track score (MH-T2TA). A spatial-temporal registration module is designed based on self-attention and a contrastive learning architecture to eliminate errors and unify the distributions of asynchronous tracks. A multiple-hypothesis algorithm is combined with deep learning to estimate the association score of a pair of tracks without relying on prior information or assumed motion models, and the optimal association pairs can be obtained. With three kinds of loss functions, tracks coming from the same targets become closer, tracks coming from different targets become more distant, and the estimated track scores are very similar to the real ones. Experimental results demonstrate that the proposed MH-T2TA can associate tracks in complex scenarios and outperform other T2TA methods.
Recently, audio-visual speech recognition (AVSR) has attracted increasing attention. However, most existing works simplify the complex challenges in real-world applications and only focus on scenarios with two speakers and perfectly aligned audio-video clips. In this work, we study the effect of speaker number and modal misalignment in the AVSR task, and propose an end-to-end AVSR framework under a more realistic condition. Specifically, we propose a speaker-number-aware mixture-of-experts (SA-MoE) mechanism to explicitly model the characteristic difference in scenarios with different speaker numbers, and a cross-modal realignment (CMR) module for robust handling of asynchronous inputs. We also use the underlying difficulty difference and introduce a new training strategy named challenge-based curriculum learning (CBCL), which forces the model to focus on difficult, challenging data instead of simple data to improve efficiency.
This study addresses the challenges of near-field interference suppression and resource allocation in extremely large-scale multiple-input multiple-output (XL-MIMO) communication systems, particularly under dense-user scenarios. We propose a quality-of-service (QoS)-aware joint user scheduling and power control scheme. Leveraging the spherical wave (SW) characteristics of near field channels, a dual-domain interference suppression strategy is developed by analyzing the spatial correlation of beam focusing vectors in terms of both angular separation and distance constraints. Based on this, a spatial correlation-based scheduling (SCS) algorithm is designed. By integrating this user selection strategy with a dynamic power allocation mechanism, the proposed approach optimizes the sum spectral efficiency while ensuring the user QoS. This framework is further extended to modular XL-MIMO systems. We show how modular deployment can enhance spatial resolution and develop an adapted QoS-aware user scheduling algorithm, called modular SCS (SCS-mod), for this architecture. Simulation results validate that the proposed algorithms significantly outperform existing schemes in terms of sum spectral efficiency and the number of scheduled users, especially under high user density and high transmission power conditions.