Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural language instructions and navigate without task-specific training. Prior works have demonstrated the potential of open-source large language models (LLMs) in zero-shot VLN-CE, yet two major limitations remain: (1) difficulty in accurately following instructions, and (2) susceptibility to loops in spatially confined or semantically similar regions. In this work, we introduce ReThinkNav, a framework designed to further advance open-source LLMs in zero-shot VLN-CE. ReThinkNav integrates contextual reasoning for enhanced instruction comprehension and progress estimation, enabling the LLM to accurately infer both the appropriate action and its rationale. In addition, a Loop Detection and Recovery (LDR) module detects loops and adjusts decisions accordingly. Experiments on the R2R-CE benchmark demonstrate excellent zero-shot performance, while real-world validation on the Unitree G1 humanoid robot confirms its practical applicability. The code is available at https://github.com/damonds27/ReThinkNav.
Integrated sensing and communication (ISAC) has emerged as a key enabler for 6G wireless networks, particularly in Industrial Internet of Things (IIoT) scenarios that demand both real-time sensing and ultra-reliable low-latency communication (URLLC) service. However, existing ISAC designs assume infinite blocklength and single-target scenarios, failing to capture the fundamental trade-off under finite blocklength (FBL) constraints required for URLLC. This paper investigates a dual-target MIMO-ISAC system where a base station performs simultaneous sensing of two targets while maintaining URLLC service to a multi-antenna user under strict latency constraints. We formulate the first multi-objective optimization problem that jointly minimizes two Cramer-Rao bounds (CRBs) while maximizing the achievable rate under FBL constraints. We characterize the Pareto boundary across three resource sharing modes: Communication and Sensing Coexistence (CSC), where tasks are scheduled sequentially; Single-Function Integration (SFI), where sensing tasks share resources; and Multi-Function Integration (MFI), where all functions operate concurrently. For CSC, we prove convexity and derive closed-form optimal solutions. For SFI, we develop a block coordinate descent algorithm exploiting partial convexity. For MFI, we design a successive convex approximation algorithm to handle non-convex rate constraints. The global optimum is obtained by adaptive mode selection. Simulation results validate the proposed design and reveal key trade-off between sensing and communication in resource constrained FBL-ISAC systems.
Object goal navigation aims to guide an agent to find a specific target object in an unseen environment using only first-person visual observations. It requires the agent to enhance scene understanding and train a robust navigation policy. To address this, we proposed two complementary techniques, commonsense-guided object graph reasoning (COGR) and policy regularization (PR). Specifically, COGR improves the agent's scene understanding by integrating object relationships, including category proximity and spatial correlation. It extracts co-occurrence embeddings of the target object from a large language model (LLM) as commonsense knowledge to guide object graph reasoning, enabling the agent to reason beyond visual co-occurrence observed in training environments. PR is a knowledge distillation-inspired regularization mechanism, where a commonsense-free model is used to regularize the navigation policy of the commonsense-guided model. We propose PR to mitigate potential performance degradation caused by knowledge bias from the LLM, enabling the training of a more robust navigation policy. Experiments in the AI2-Thor and RoboThor environments demonstrate the effectiveness and efficiency of our proposed method, and real-world deployment further validates its transferability.
High-definition map (HD Map), as a critical component of digital infrastructure, plays a vital role in the advancement of autonomous driving technologies. Existing HD Map models are primarily structured according to the update frequency of map elements, and they insufficiently account for the service objects and task execution entities involved in autonomous driving. Moreover, these models often lack a systematic analysis of the interrelationships between map elements and their functional roles in autonomous driving scenarios. To address these limitations, this paper proposes a novel high-definition map model, termed the Human-Vehicle-Road-Map (HVRM) model, which considers the service objects, operational carriers, and supporting media of autonomous driving from the perspective of task execution. The proposed model integrates three core entities, namely humans, vehicles, and roads, into a unified high-definition map framework, and systematically analyzes the relationships among map elements and their significance in autonomous driving. The effectiveness of the proposed model is validated through simulation experiments. This research contributes to the existing body of HD map studies and offers new insights for the further development of autonomous driving systems.
Over the past decade, the fusion of light detection and ranging (LiDAR) and inertial navigation systems (INS) has become a reliable solution for environmental perception in intelligent mobile platforms. However, LiDAR performance degrades under adverse weather conditions. Millimeter-wave radar offers complementary benefits: it is robust to environmental disturbances and provides direct Doppler velocity measurements. Moreover, most existing fusion frameworks rely on the extended Kalman filter (EKF), which suffers from linearization errors and inconsistency. To address these limitations, we propose and derive a tightly-coupled LiDAR–radar–inertial odometry framework based on the equivariant filter (EqF). The framework leverages both LiDAR and radar to provide complementary constraints on the 9-dimensional navigation state, significantly improving accuracy and robustness in challenging environments. Furthermore, we conduct an in-depth observability analysis using Lie derivative theory, examining both the nonlinear system and its discrete filter system. We show that the EqF preserves the same unobservable directions as the underlying nonlinear system, thereby ensuring consistent state estimation, a property that standard EKF does not possess. In addition, experiments on real-world datasets demonstrate that our method outperforms the state-of-the-art LiDAR–inertial odometry approach and other EKF-based approaches in terms of localization accuracy, robustness and velocity estimation precision.
Unmanned aerial vehicle (UAV)-enabled wireless communication technique has emerged as a promising paradigm in the next generation communication networks. However, due to the security issues arising from the inherently open nature of air to ground (A2G) links, the security of both communication behavior and information faces significant threat. To address this issue, we investigate a UAV-enabled covert and secure communication network against cooperative detection and eavesdropping from multiple illegal eves, which can share the intercepted signals with others. Specifically, from the perspective of secure communication considering channel uncertainty and cooperative eavesdropping, we characterize users' upload secure throughput by an ergodic integral expression, and formulate a max-min throughput optimization problem through joint design of UAV trajectory, user transmit power and segmentation ratio, under mobility, energy consumption and covertness constraints. Then, to address the problem with complicated covertness metric with no closed-form expression under cooperative detection, we derived closed form expressions of the covertness metric, by utilizing the Pinsker inequality along with properties of both Taylor series and the Digamma function. Subsequently, the monotonicity of the derived metric is analyzed and further utilized to simplify the covertness constraint into transmit power constraint equivalently. Afterward, a novel convex approximation method is introduced to construct a convex subproblem for the formulated highly complicated non-convex problem, enabling an efficient algorithm that delivers high-quality solutions for the formulated problem. Finally, the effectiveness and superior performance of proposed design are verified through simulations.
Reliable and continuous vehicle positioning in urban canyons remains a critical challenge for intelligent transportation systems. To address this issue, this article presents a left equivariant error model for strapdown inertial navigation system and global navigation satellite system (SINS/GNSS) integrated navigation in the inertial frame. Unlike the imperfect invariant extended Kalman filter and its variant, the proposed method copes with the biases states in a more natural way. Also, different from the current equivariant filter based SINS/GNSS integrated navigation, the linearized measurement matrices of this method are independent of the system’s position. This method allows the initial alignment stage and integrated navigation stage to be combined into one stage, enabling faster and more accurate navigation results in challenging urban settings. Firstly, we introduce the two-frame group and associated group actions on the state space and input space of the biased INS from the perspective of left equivariant filter. Secondly, we derive the left equivariant error dynamics in inertial frame and the measurement models for both loosely coupled (LC) SINS/GNSS and tightly coupled (TC) SINS/GNSS, respectively. Thirdly, we conduct the simulation and field test to evaluate the practical performance of the proposed method. Experimental results demonstrate that the proposed method achieves superior transient response in urban canyon environments compared to conventional approaches, enabling reliable one-step alignment and continuous positioning for intelligent vehicles operating in complex traffic scenarios.
Open-vocabulary instance segmentation aims to segment objects of arbitrary categories from textual queries, but existing methods typically couple semantic recognition with mask or contour prediction, limiting generalization to unseen objects with diverse shapes and complex boundaries. In this paper, we propose WorldSnake, a contour-based open-vocabulary instance segmentation framework that reformulates the task through semantic–geometric decoupling. Instead of predicting semantic masks directly, we decompose the task into semantic object localization and category-agnostic contour modeling, where an open-vocabulary detector provides instance localization and a contour network focuses on geometric boundary evolution with reduced reliance on category-specific semantic supervision. To improve contour modeling, we design a Hierarchical Progressive Contour Refinement (HPCR) module that iteratively updates contours from coarse to fine using multi-scale features. In addition, we introduce a semantics-agnostic training strategy that learns contour deformation from contour-level supervision, encouraging the learning of transferable boundary representations across object categories. The results suggest that semantic–geometric decoupling and contour-based geometric modeling are effective for open-vocabulary visual perception. Our code will be available at https://github.com/Giansar-Wu/WorldSnake.
Object-goal navigation aims to guide an agent to find a specific target object in an unfamiliar environment based on first-person visual observations. It requires the agent to learn informative visual representations and robust navigation policy. To promote these two components, we propose two complementary techniques, dual object graph (DOG) and bisimulation metric (BM). DOG integrates current and historical object relationships, including category proximity and spatial correlation. It constructs a current object graph (COG) to model real-time object relationships and maintains a historical object graph (HOG) to preserve long-term object relationships. DOG improves visual representation learning. Both DOG and BM aim to improve robust navigation policy, enabling the agent to escape from deadlock states, such as looping or getting stuck. Specifically, BM is a self-supervised reinforcement learning (RL) technique that groups behaviorally similar observations in representation space. In the process, BM learns robust latent representations that capture only the task-relevant information from observations against distractions such as variations in background or viewpoint, thus providing guidance for effective navigation policy. Experiments in the AI2-Thor and RoboThor environment demonstrate that our method significantly improves the effectiveness and efficiency of navigation in unfamiliar environments, and real-world deployment further validates its transferability.
Large language model (LLM)-driven embodied agents rely on environment states to interpret scenes, generate high-level plans, and drive physical execution, making planner-visible state representations a critical security boundary. Existing attacks primarily manipulate user instructions, prompt contexts, model behavior, or perceptual inputs, while paying limited attention to whether environment-state text itself can serve as deceptive task evidence and propagate beyond planning to affect execution outcomes. Because embodied tasks are constrained by entity grounding, action preconditions, spatial relations, and environmental constraints, planning deviation alone does not guarantee adversarial execution. To address this gap, we investigate environment-state text as an independent attack surface and present the first closed-loop Environment State-Text Injection (ESTI) attack for LLM-driven embodied agents. Without modifying the original user instruction, model parameters, or executor, ESTI reformulates an adversarial objective as false state evidence compatible with the current environment and influences planning and execution through object properties, spatial relations, affordances, task-stage rules, and execution feedback. We further develop ESTI-Bench to evaluate attack propagation across the planning-to-execution closed loop and compare ESTI with Vanilla IPI, EIRAD, and BADROBOT across ProgPrompt/VirtualHome, VoxPoser/RLBench, and AI2-THOR/iTHOR. ESTI consistently outperforms existing baselines, improving planning-level and execution-level attack success rates by up to 89.32% and 43.69%, respectively. Further analysis shows that grounding, consistency, and executability jointly determine whether manipulated state evidence can propagate through the embodied closed loop and produce verifiable environmental changes.
Cooperative localization is essential for swarm applications like collaborative exploration and search-and-rescue missions. However, maintaining real-time capability, robustness, and computational efficiency on resource-constrained platforms presents significant challenges. To address these challenges, we propose D-GVIO, a buffer-driven and fully decentralized GNSS-Visual-Inertial Odometry (GVIO) framework that leverages a novel buffering strategy to support efficient and robust distributed state estimation. The proposed framework is characterized by four core mechanisms. Firstly, through covariance segmentation, covariance intersection and buffering strategy, we modularize propagation and update steps in distributed state estimation, significantly reducing computational and communication burdens. Secondly, the left-invariant extended Kalman filter (L-IEKF) is adopted for information fusion, which exhibits superior state estimation performance over the traditional extended Kalman filter (EKF) since its state transition matrix is independent of the system state. Thirdly, a buffer-based re-propagation strategy is employed to handle delayed measurements efficiently and accurately by leveraging the L-IEKF, eliminating the need for costly re-computation. Finally, an adaptive buffer-driven outlier detection method is proposed to dynamically cull GNSS outliers, enhancing robustness in GNSS-challenged environments.
Human perception possesses the remarkable ability to mentally reconstruct the complete structure of occluded objects, which has inspired researchers to pursue amodal instance segmentation for a more comprehensive understanding of the scene. Previous works have shown promising results, but they often capture the contextual dependencies in an unsupervised way, which can lead to undesirable contextual dependencies and unreasonable feature representations. To tackle this problem, we propose a Pixel Affinity-Parsing (PAP) module trained with the Pixel Affinity Loss (PAL). Embedded into CNN, the PAP module can leverage learned contextual priors to guide the network to explicitly distinguish different relationships between pixels, thus capturing the intra-class and inter-class contextual dependencies in a non-local and supervised way. This process helps to yield robust feature representations to prevent the network from misjudging. To demonstrate the effectiveness of the PAP module, we design an effective Pixel Affinity-Parsing Network (PAPNet). Notably, PAPNet also introduces shape priors to guide the amodal mask refinement process, thus preventing implausible shapes in the predicted masks. Consequently, with the dual guidance of contextual and shape priors, PAPNet can reconstruct the full shape of occluded objects accurately and reasonably. Experimental results demonstrate that the proposed PAPNet outperforms existing state-of-the-art methods on multiple amodal datasets. Specifically, on the KINS dataset, PAPNet achieves 37.1% AP, 60.6% AP50 and 39.8% AP75, surpassing C2F-Seg by 0.6%, 2.4% and 2.8%. On the D2SA dataset, PAPNet achieves 71.70% AP, 85.98% AP50 and 77.10% AP75, surpassing PGExp by 0.75% and 0.33% in AP50 and AP75, and being comparable to PGExp in AP. On the COCOA-cls dataset, PAPNet achieves 41.29% AP, 60.95% AP50 and 46.17% AP75, surpassing PGExp by 3.74%, 3.21% and 4.76%. On the CWALT dataset, PAPNet achieves 72.51% AP, 85.02% AP50 and 80.47% AP75, surpassing VRSPNet by 5.38%, 0.07% and 5.35%. The code is available at https://github.com/jiaoZ7688/PAP-Net.
Object-goal navigation aims to guide an agent to find a specific target object in an unfamiliar environment based on first-person visual observations. It requires the agent to learn informative visual representations and robust navigation policy. To promote these two components, we proposed two complementary techniques, context-aware graph inference (CGI) and generative adversarial imitation learning (GAIL). CGI improves visual representation learning by integrating object relationships, including category proximity and spatial correlation. It uses the translation on hyperplane (TransH) method to infer context-aware object relationships under the guidance of various contexts over navigation episodes, including image, action, and memory. Both CGI and GAIL aim to improve robust navigation policy, enabling the agent to escape from deadlock states, such as looping or getting stuck. GAIL is an imitation learning (IL) technique that enables the agent to learn from expert demonstrations. Specifically, we propose GAIL to address the non-discriminative reward problem that exists in object-goal navigation. GAIL designs a dynamic reward function and combines it with environment rewards, thus providing guidance for effective navigation policy. Experiments in the AI2-Thor and RoboThor environments demonstrate that our method significantly improves the effectiveness and efficiency of navigation in unfamiliar environments.
Objectives: We focus on exploring an unsupervised learning-based model which can take advantage of a single image and events to estimate dense and time-continuous optical flow. Methods: We propose a multi-scale optical flow recurrent estimation network, called MREIFlow, which mainly consists of a triplet feature encoder, a feature fusion subnetwork, and a flow iterative decoder. The triplet feature encoder is capable of extracting multi-scale features from a single image and events. The feature fusion subnetwork is designed to integrate spatial-temporal information in features and generate pseudo-features. The flow iterative decoder can perform feature correlation and estimate full-resolution optical flow iteratively in a coarse-to-fine way. We train the network in an unsupervised learning way to avoid using labeled data. In order to provide plentiful and reliable supervisory signals, we design a hierarchical unsupervised learning scheme. In the scheme, the multi-level optimization objectives are designed to supervise network training in different aspects. We also introduce a teacher model to provide auxiliary supervision. Besides, we apply a loss selection strategy to mine promising supervisory signals. Results: Experiments on the MVSEC dataset show that the proposed method achieves remarkable performance under indoor and outdoor circumstances. Compared with previous methods, the proposed method can provide dense and time-continuous optical flow estimation. Conclusion: The proposed network can take advantage of a single image and events to produce dense and time-continuous optical flow estimation, and can be trained through unsupervised learning techniques. However, it requires high computation and memory to guarantee the precision of optical flow. To address this problem, we will explore an efficient and lightweight network architecture in the future.
In this paper, we propose a leg odometry assisted global navigation satellite system (GNSS)/inertial navigation system (INS) integrated navigation system for the quadruped robot. First, we present a left invariant extended Kalman filter for global measurement on Rotating earth in the local world frame. Next, we construct a leg odometry and derive the contact point motion dynamics and leg odometry observation equation for the quadruped robot. Finally, the field experiment shows that the odometry assisted GNSS/INS integrated system can significantly reduce the position error compared to the one without GNSS signal.
Existing RGB image-based object detection methods achieve high accuracy when objects are static or in quasi-static conditions but demonstrate degraded performance with fast-moving objects due to motion blur artifacts. Moreover, state-of-the-art deep learning methods, which rely on RGB images as input, necessitate training and inference on high-performance graphics cards. These cards are not only bulky and power-hungry but also challenging to deploy on compact robotic platforms. Fortunately, the emergence of event cameras, inspired by biological vision, provides a promising solution to these limitations. These cameras offer low latency, minimal motion blur, and non-redundant outputs, making them well suited for dynamic obstacle detection. Building on these advantages, a novel methodology was developed through the fusion of events with depth to address the challenge of dynamic object detection. Initially, an adaptive temporal sampling window was implemented to selectively acquire event data and supplementary information, contingent upon the presence of objects within the visual field. Subsequently, a warping transformation was applied to the event data, effectively eliminating artifacts induced by ego-motion while preserving signals originating from moving objects. Following this preprocessing stage, the transformed event data were converted into an event queue representation, upon which denoising operations were performed. Ultimately, object detection was achieved through the application of image moment analysis to the processed event queue representation. The experimental results show that, compared with the current state-of-the-art methods, the proposed method has improved the detection speed by approximately 20% and the accuracy by approximately 5%. To substantiate real-world applicability, the authors implemented a complete obstacle avoidance pipeline, integrating our detector with planning modules and successfully deploying it on a custom-built quadrotor platform. Field tests confirm reliable avoidance of an obstacle approaching at approximately 8 m/s, thereby validating practical deployment potential.
Visual semantic SLAM integrates geometric measurements with semantic perception, making it widely applicable in autonomous driving and robotics. Semantic-assisted localization and dynamic object perception are two critical tasks in visual semantic SLAM. However, many existing state-of-the-art methods address only one of these tasks in isolation. To address issues of functional limitations and insufficient information utilization in a single framework, we propose a unified visual semantic SLAM framework, SDS-SLAM, which tightly couples static and dynamic semantic information to handle the motion estimation of both the camera and observed objects in driving scenarios. A multi-task network for driving perception is employed to extract semantic information, including drivable areas, lanes, and vehicles. Based on various information obtained, we propose semantic local ground manifolds (SLGMs) to represent the geometric structure and semantic features, enabling the online generation of a lightweight semantic map. Subsequently, we integrate SLGM-based constraints such as lane alignment and planar motion to promote camera and object pose estimation. We evaluated our method on the public KITTI dataset and self-collected real-world data. The results demonstrate that our method effectively perceives both dynamic and static semantic elements in driving scenarios, achieving high accuracy in estimating the poses of the camera and objects.
Visual navigation requires the agent reasonably perceives the environment and effectively navigates to the given target. In this task, we present a Multimodal Adaptive Graph (MAG) for learning and grounding the visual clues based on the object relationships. MAG consists of key navigation elements: object relative position relationships, previous navigation actions, past training experience, and target objects. This enables the agent to accurately gather multimodal information and find the target faster. Technically, our framework performs continuous modeling of pre-trained vision-language grounding model to align the multimodal graph, text information with visual perception. For output, we introduce constraints on the graph's value estimation (GVE) functions to supervise the agent predict optimal actions, which can help it escape from deadlocks. With the MAG, the agent can effectively perceive the environment and get optimal actions. We train our framework with human demonstration and collision signals. Results demonstrate that our approach improves by 10.1% in SPL (Success weighted by Path Length) and 25.4% in success rate relative to the baseline method in the AI2THOR environment. Our code will be publicly released in the scientific community.
Pose estimation is a crucial problem in simultaneous localization and mapping (SLAM). However, developing a robust and consistent state estimator remains a significant challenge, as the traditional extended Kalman filter (EKF) struggles to handle the model nonlinearity, especially for inertial measurement unit (IMU) and light detection and ranging (LiDAR). To provide a consistent and efficient solution of pose estimation, we propose Eq-LIO, a robust state estimator for tightly coupled LIO systems based on an equivariant filter (EqF). Compared with the invariant Kalman filter based on the SE2 (3) group structure, the EqF uses the symmetry of the semi-direct product group to couple the system state including IMU bias, navigation state, and LiDAR extrinsic calibration state, thereby suppressing linearization error and improving the behavior of the estimator in the event of unexpected state changes. The proposed Eq-LIO owns natural consistency and higher robustness, which is theoretically proven with mathematical derivation and experimentally verified through a series of tests on both public and private datasets.
The standard extended Kalman filter (EKF) is widely used in integrated strapdown inertial navigation system and global navigation satellite system (SINS/GNSS). However, its state-space model, derived from traditional error definitions, leads to poor estimation consistency. Moreover, EKF fails to converge quickly and accurately in scenarios with large attitude misalignment. To address these issues, we derive the tightly coupled (TC) SINS/real-time kinematic (RTK) and the TC SINS/precise point positioning (PPP) with the tangent group-based equivariant global error. Furthermore, to mitigate the impact of large first-order approximation errors and the covariance distortion on filter performance, we proposed an iterated equivariant filter (I-EqF) that incorporates a covariance reset step. Field vehicle experiments with low-cost SINS/GNSS integration demonstrate that the proposed filter exhibits superior estimation accuracy and transient behavior, especially in the scenario involving large attitude misalignment.