Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual question answering and chain-of-thought reasoning data, which emphasizes linguistic reasoning rather than action-grounded planning. As a result, the learned representations capture semantic knowledge but lack spatial dependencies crucial for reliable trajectory prediction. We propose DriveTeach-VLA, a framework that explicitly teaches VLAs what to see and where to look. Driving-aware Vision Distillation (DVD) injects driving-specific perceptual priors into the vision encoder, while 2D Trajectory-Guided Prompts (2D-TGP) provide spatial conditioning aligned with feasible driving trajectories. Together, they form a vision-guided learning pipeline: what to see (DVD pretraining) - where to look (TGP-guided SFT) - how to act (TGP-guided GRPO). DriveTeach-VLA achieves the state-of-the-art performance on NAVSIM and nuScenes. Our code is available at: https://github.com/ShivaTeam/DriveTeach-VLA.
Multiple-input multiple-output (MIMO) beamforming has been widely recognized for their ability to enhance the performance of integrated sensing and communication (ISAC). Fully digital beamforming is prohibitively expensive, while hybrid beamforming is capable of achieving comparable performance at a much lower cost. However, existing partially-connected hybrid beamforming typically employs uniform subarray structures and neglects the impact of mutual coupling on the performance. To overcome these limitations, this paper proposes a partially-connected hybrid beamformer that jointly optimizes sparse array for reducing the mutual coupling effect and subarray structure for performance improvement. The optimal solutions obtained by convex optimization are used for labeling and for constructing the training dataset for the subsequent convolutional neural network (CNN)-based design. A neighborhood clustering strategy is proposed to significantly reduce the number of classes, thereby improving classification accuracy. Experimental results demonstrate that the proposed joint design strategy substantially improves the ISAC performance, even exceeding the full array when considering the mutual coupling effects.
Orthogonal time frequency space (OTFS) modulation offers strong resilience to Doppler effects but suffers from high system latency, limiting its use in low-latency communications. This paper proposes a channel estimation algorithm for low-latency OTFS systems with large delay spreads and fractional Doppler effects. In the delay-time (DT) domain, impulse pilots are placed at equidistant intervals along the first row of the DT grid to eliminate interference from aliased delays, and a threshold detection method estimates the delays. Doppler shifts and path gains are then estimated using the discrete Fourier transform (DFT). Simulation results show that the proposed algorithm achieves near-optimal bit error rate (BER) under sparse delay spreads.
Despite significant advances in generic object detection, a persistent performance gap remains for tiny objects compared to normal-scale objects. We demonstrate that tiny objects are highly sensitive to annotation noise, where optimizing strict localization objectives risks noise overfitting. To address this, we propose Tiny Object Localization with Flows (TOLF), a noise-robust localization framework leveraging normalizing flows for flexible error modeling and uncertainty-guided optimization. Our method captures complex, non-Gaussian prediction distributions through flow-based error modeling, enabling robust learning under noisy supervision. An uncertainty-aware gradient modulation mechanism further suppresses learning from high-uncertainty, noise-prone samples, mitigating overfitting while stabilizing training. Extensive experiments across three datasets validate our approach’s effectiveness. Especially, TOLF boosts the DINO baseline by 1.2% AP on the AI-TOD dataset.
This paper is concerned with the near-field channel estimation (CE) in extremely large-scale multi-input multi-output (XL-MIMO) systems under spatial non-stationarity (SnS). To this aim, we first analyze the channel characteristics in the systems and design a structured SnS-aware mask matrix, which reveals the relationship between the angular-domain block sparsity of the channel and the SnS structure. Inspired by the relationship, we establish a near-field channel model for the XL-MIMO systems under SnS. Second, we propose a model-data hybrid driven approach, termed SBL4CE-Net, to estimate the near-field channel. SBL4CE-Net unfolds the block sparse Bayesian learning (BSBL) algorithm into a multilayer solution framework and designs customized neural networks (NN) based on parameter features of BSBL to learn its hyperparameters. Simulation results demonstrate that SBL4CE-Net achieves good normalized mean square error (NMSE) performance, demonstrating a balance between interpretability and data-driven adaptability for SnS affected near-field channels.
This paper is concerned with autonomous aerial vehicle (AAV) video coding and transmission in scenarios such as aerial search and monitoring. Unlike existing methods of modeling AAV video source coding and channel transmission separately, we investigate the joint source-channel optimization (JSCO) issue for video coding and transmission. Particularly, we design eight-dimensional delay-power-rate-distortion models in terms of source coding and channel transmission and characterize the correlation between video coding and transmission, with which a JSCO problem is formulated. Its objective is to minimize end-to-end distortion and AAV power consumption by optimizing fine-grained parameters related to AAV video coding and transmission. This problem is confirmed to be a challenging sequential-decision and non-convex optimization problem. We therefore decompose it into a family of repeated optimization problems by Lyapunov optimization and design an approximate convex optimization scheme with provable performance guarantees to tackle these problems. Based on the theoretical transformation, we propose a Lyapunov repeated iteration (LyaRI) algorithm. Both objective and subjective experiments are conducted to comprehensively evaluate the performance of LyaRI. The results indicate that, compared with its counterparts, LyaRI achieves better video quality and stability performance, with a 47.74% reduction in the variance of the obtained encoding bit.
High altitude platform (HAP) millimeter wave (mmWave) links are sensitive to wind-induced attitude shaking. Small roll, pitch, or yaw tilts will misalign narrow communication beams and markedly reduce the gain of the array antenna. In this paper, we propose a vision-augmented large language model (VA-LLM) that forecasts short-horizon attitudes and pre-steers communication beams before misalignment occurs. Specifically, we design a five-channel vision module to render multivariate flight data into pseudo-RGB images with reversible instance normalization (RevIN) and backbone-aligned normalization. We design a learned cross-variable attention (CVA) branch to condense intra-timestep channel relations into a compact token set. Concurrently, we explore a time-series-aware prompt-as-prefix (PaP) to inject frequency, amplitude, and top- $K$-based periodic statistics from the input flight-data window into the token set. After fusing vision-augmented, learned, and text tokens, a frozen LLM refines the representation; a lightweight temporal mixer with axis-wise heads directly regresses future roll/pitch/yaw, which are mapped to continuous steering angles for the array antenna. On real flight test data, VA-LLM improves the average signal-to-noise-ratio (SNR)-type array-gain ratio by 6.56 % over the baseline, achieves a 7.31 % gain over the baseline at the 12step horizon, and yields up to 14.10 % gains over the ablations.
This paper proposes a consistency-model-based channel estimation algorithm for multiple-input multiple-output (MIMO) systems. The proposed algorithm employs a consistency model (CM) to learn the angle-domain channel distribution and uses the trained CM as a plug-and-play (PnP) generative prior for MIMO channel estimation. The proposed algorithm alternates between a pilot-observation-based data-consistency update and a CM-prior-based denoising update. In addition, the proposed algorithm adaptively selects the penalty parameter according to residual energy and residual whiteness, and adjusts the CM denoising level according to the observed signal-to-noise ratio (SNR), thereby avoiding the performance degradation caused by fixed inference schedules under varying observation conditions. Simulation results show that the proposed algorithm not only reduces the number of inference steps by 50
Timely detection of anomalies in traffic systems is crucial for mitigating risks and economic losses. Current time series anomaly detection methods often use reconstruction errors to identify anomalies, and their accuracy depends on how well they can reconstruct normal patterns from the original series. However, traffic data typically exhibit high noise, spatiotemporal heterogeneity, and uncertain inter-series correlations, complicating the learning of normal patterns. To address these challenges, we introduce SpectraBayes, which explores the reconstruction of density and volume series in the frequency domain for anomaly detection. First, we transform the series into the frequency domain and apply a low-pass filter to remove noise. Then, we embed periodic information into the frequency-domain representation through phase shifts to enhance the temporal awareness. Additionally, we model the inter-series correlations between density and volume resiliently using cross-spectrum probabilistic modeling. Optimized by maximizing the Evidence Lower Bound (ELBO), SpectraBayes ensures robust reconstruction while avoiding overfitting against the uncertain data. SpectraBayes outperforms 21 existing anomaly detection models on traffic series anomaly detection tasks, achieving mean improvements of 2.71% across three metrics over the second-best model. Furthermore, it is lightweight and maintains robust performance under varying noise levels.
Near-space information networks (NSINs) composed of high-altitude platforms (HAPs) and high- and low-altitude unmanned aerial vehicles (UAVs) are a new regime for providing quick, robust, and cost-efficient sensing and communication services. Precipitated by innovations and breakthroughs in manufacturing, materials, communications, electronics, and control techniques, NSINs have been envisioned as an essential component of the emerging sixth-generation of mobile communication systems. This article reveals some critical issues needing to be tackled in NSINs through conducting experiments and discusses the latest advances in NSINs in the research areas of channel modeling, networking, and transmission from a forward-looking, comparative, and technical evolutionary perspective. In this article, we highlight the characteristics of NSINs and present the promising use cases of NSINs. The impact of airborne platforms' unstable movements on the phase delays of onboard antenna arrays with diverse structures is mathematically analyzed. The recent advances in HAP channel modeling are elaborated on, along with the significant differences between HAP and UAV channel modeling. A comprehensive review of the networking techniques of NSINs in network deployment, handoff management, and network management aspects is provided. Besides, the promising techniques and communication protocols of the physical (PHY) layer, medium access control (MAC) layer, network layer, and transport layer of NSINs for achieving efficient transmission over NSINs are reviewed, and we have conducted experiments with practical NSINs to verify the performance of some techniques. Finally, we outline some open issues and promising directions for NSINs deserved for future study and discuss the corresponding challenges.
The emerging high-altitude platform (HAP) networks are envisioned as critical components in space-air-ground integrated networks. This paper investigates the uplink channel estimation for large-scale reconfigurable intelligent surface (RIS)-aided HAP networks. To overcome the HAP shaking effect and high computational overhead of massive passive arrays,we propose a shaking-aware fast three-stage channel estimation (SA-FTCE) algorithm in the angular domain, tailored for uniform planar arrays (UPAs). SA-FTCE achieves a computationally efficient estimate by progressively pruning the angular channel matrix to a lower dimension by eliminating inactive azimuth and elevation angular regions. Specifically, in Stage 1, we derive the angle-of-arrival (AoA) interval for the RIS-HAP link through the spatial relationship between the AoA variation and HAP attitude shaking, and introduce a shaking-aware AoA search for initial pruning. In Stage 2, a novel Kronecker variational Bayesian inference (Kronecker-VBI) algorithm is proposed for the low-complexity detection of the effective angular region (EFAR) for further pruning. Finally, the channel estimation is efficiently obtained by a VBI based estimator within the drastically reduced angular space. The simulation results show that the proposed SA-FTCE scheme is faster than its counterparts and achieves comparable estimation accuracy.
Small changes in high-altitude platform (HAP) attitude can cause significant deviations in HAP downlink beam directions, thereby severely degrading HAP downlink communication performance. In this paper, we develop a multimodal large language model (LLM) enabled beamforming framework to achieve robust HAP downlink communications. Specifically, we design a vision-language LLM (VL-LLM) that learns from multivariate flight telemetry to forecast short-term HAP attitudes under platform shaking and support delay-aware proactive beam steering. We design an offline forecast-error calibration procedure to obtain upper bounds on forecast errors and improve the reliability of proactive analog beam steering. Based on the attitude forecasts, we proactively update the analog beamformer and propose a QoS-driven beamforming and admission method with a lightweight feasibility-enforcement step to satisfy instantaneous transmit-power and QoS requirements. Simulation results indicate that the designed VL-LLM can accurately capture changes in the HAP attitude and the proposed beamforming method achieves a 22.1% higher user service ratio and a 12.5% higher sum-rate than representative baselines. The measured mean and p99 computational latencies are 36.24 ms and 40.13 ms, respectively, supporting low-latency online implementation.
High altitude platforms (HAPs) are emerging as a key enabler for 6G coverage, yet limited energy must be split between propulsion and communications. Most prior HAP studies ignore propulsion power or rely on surrogates that miss hull-propeller interference, leading to misestimated communication power budgets and degraded beamforming. More importantly, HAP power allocation is intrinsically a multi-system, multidisciplinary problem in which aerodynamics, propulsion-system efficiency, and communication-system performance (quality of service (QoS) and energy efficiency (EE)) are tightly coupled.To address these challenges, this paper designs an interactive generative artificial intelligence (AI)-empowered HAP power allocation agent.By interacting with the AI agent, we develop an accurate propulsion power consumption model that takes into account the aerodynamic interference between the HAP's hull and the propeller. Assisted by the AI agent, we further formulate a HAP beamforming problem to improve user QoS and enhance the EE of the HAP communication system.This paper also proposes a QoS-enhanced energy-efficient (Q3E) beamforming algorithm to solve the formulated problem.Simulation results demonstrate the accuracy of the propulsion-power model and the effectiveness of the Q3E algorithm.
Industrial gas leakage detection is critically important for safety and environmental protection. While infrared imaging enables detection of invisible gases, two challenges remain: existing datasets lack realistic industrial scenarios, and current methods struggle to distinguish gas plumes from background interferences or segment discontinuous gas distributions. This paper introduces a benchmark comprising an Industrial RGB-Thermal Dataset (IRTD) with gas emission and leakage data from laboratory and industrial sites. A VLM-assisted RGBThermal detection framework with a Cross-Attention based Feature Difference (CAFD) module is designed to enhance gasspecific feature differentiation by computing inter-modal feature discrepancies. Evaluations on public datasets and IRTD demonstrate state-of-the-art results.
Drones have become increasingly widely applied in surveillance systems due to their mobility, making aerial video anomaly detection methods more crucial. Anomalies in aerial videos often present as semantic conflicts, such as the presence of unexpected objects or unusual behaviors that do not align with the context. Previous approaches often relied on manually crafted knowledge graphs to detect such conflicts, which suffer from poor scalability. Recently, owing to their sufficient alignment training, multimodal large language models (MLLMs) have emerged as a generalized solution for semantic understanding. However, the direct application of MLLMs does not yield satisfactory anomaly detection performance in aerial videos. First, aerial videos often manifest platform-induced pseudo-motion, which obscures the true motion of objects and exacerbates detection errors. Second, without sufficient labeled data for fine-tuning, generic MLLMs often lack scene-level semantic guidance to reliably distinguish abnormal events that include contextually inappropriate behaviors. To address these challenges, we propose SemAero, an MLLM-based framework to address these challenges by: 1) designing an ego-motion reduction module to enhance model perception on object movement, 2) generating scene-specific prompts adaptively with step-by-step guidance for reasonable output, and 3) refining scores with dual-stream consistent feature for better domain-specific anomaly detection. Evaluated across 8 diverse aerial scenes and 73 sub-datasets, SemAero achieves a 3.09% improvement in AUC-ROC over the second-best model, demonstrating its ability in aerial video anomaly detection.
To address the complexities of spatial non-stationary (SnS) effects and spherical wave propagation in near-field channel estimation (CE) for extremely large-scale multiple-input multiple-output (XL-MIMO) systems, this paper proposes an SnS-aware CE framework based on adaptive subarray partitioning. We first investigate spherical wave propagation and various SnS characteristics and construct an SnS near-field channel model for XL-MIMO systems. Due to the limitations of uniform subarray patterns in capturing SnS, we analyze the adverse effects of the non-ideal array segmentation (over- and under-segmentation) on CE accuracy. To counter these issues, we develop a dynamic hybrid beamforming-assisted power-based subarray segmentation paradigm (DHBF-PSSP), which integrates power measurements with a dynamic hybrid beamforming structure to enable joint subarray partitioning and decoupling. A power-adaptive subarray segmentation (PASS) algorithm leverages the statistical properties of power profiles, while subarray decoupling is achieved via a subarray segmentation-based sampling method (SS-SM) under radio frequency (RF) chain constraints. For subarray CE, we propose a subarray segmentation-based assorted block sparse Bayesian learning algorithm under the multiple measurement vectors framework (SS-ABSBL-MMV). This algorithm exploits angular-domain block sparsity under a discrete Fourier transform (DFT) codebook and inter-subcarrier structured sparsity. Simulation results confirm that the proposed framework outperforms existing methods in CE performance.
This paper is concerned with rate control (RC) for autonomous aerial vehicle (AAV) video encodingto produce a stable video bitrate and high quality videos under specific constraints. To this aim, we theoretically investigate the relationships between encoding parameters and bitrate and distortion, and design an accurate online rate-distortion (R-D) model by employing linear regression and predictive artificial intelligence (AI) methods. Considering the sensitivity of AAV video encoding to latency and power consumption, we further establish the mathematical relationships between encoding parameters and encoding time and power consumption. By integrating the R-D and encoding time and power consumption models, we formulate a multi-timescale optimization problem and propose a novel algorithm to solve it. Specifically, we decompose it into a family of single-timescale problems via an alternating direction method of multipliers (ADMM). In addition, we design an iterative optimization scheme to solve the single-timescale problems with low computational complexity. Extensive experiments are conducted to validate the designed model and algorithm. Experimental results indicate that the designed model has small estimation errors, and the variance of Y-PSNR achieved by the proposed algorithm is not greater than 26.3% of its counterpart.
With the advent of extremely large-scale multipleinput multiple-output (XL-MIMO) systems, wireless channels exhibit challenging characteristics distinct from conventional systems, such as spherical wave effects and spatial non-stationarity (SnS). To mitigate the adverse impact of these characteristics on subsequent communication processes (e.g., channel estimation), this paper extends the near-field channel model for line-of-sight (LoS) XL-MIMO systems to accommodate SnS properties. First, we derive in detail the negative impacts of over-segmentation and under-segmentation under the influence of SnS, emphasizing the importance of precise subarray partitioning. Then, we introduce a power-adaptive subarray segmentation algorithm (PASS). Simulation results validate the accuracy of the proposed algorithm.
Multivariate time series anomaly detection is crucial in sensitive domains such as cybersecurity and grid monitoring, significantly contributing to the reliability and safety of system operation. However, current methods suffer from inadequate utilization of decomposed time series, insufficient mining of contextual dependencies within the time series, and limited robustness against anomalies during training. To address these limitations, we propose the consistency-enhanced normalizing flow (ConFlow) model, which utilizes the consistency of decomposed time series and contextual temporal embedding to enhance the discriminative ability of the flow model. First, to refine the extraction of time series components, we propose a cascade decomposition and mixing module that iteratively decouples the time series. Second, these components are mapped to Gaussian distributions through the context-aware normalizing flow, incorporating both inter- and intra-series information into the density estimation. Third, the density consistency among decomposed time series is measured to reweight the estimation, while highly inconsistent series are viewed as anomalies and masked during training to improve model robustness. Finally, anomalies are detected using reweight density estimation. Experiments on five widely used datasets in the time series anomaly detection field demonstrate the superiority of our method over state-of-the-art (SOTA) approaches.