Video stabilization remains a fundamental problem in computer vision, particularly pixel-level synthesis solutions for video stabilization, which synthesize full-frame outputs, add to the complexity of this task. These methods aim to enhance stability while synthesizing full-frame videos, but the inherent diversity in motion profiles and visual content present in each video sequence makes robust generalization with fixed parameters difficult. To address this, we present a novel method that improves pixel-level synthesis video stabilization methods by rapidly adapting models to each input video at test time. The proposed approach takes advantage of low-level visual cues available during inference to improve both the stability and visual quality of the output. Notably, the proposed rapid adaptation achieves significant performance gains even with a single adaptation pass. We further propose a jerk localization module and a targeted adaptation strategy, which focuses the adaptation on high-jerk segments for maximizing stability with fewer adaptation steps. The proposed methodology enables modern stabilizers to overcome the longstanding SOTA approaches while maintaining the full frame nature of the modern methods, and offers users with stability control mechanisms akin to classical approaches. Extensive experiments on diverse real-world datasets demonstrate the versatility of the proposed method. Our approach consistently improves the performance of various full-frame synthesis models in both qualitative and quantitative terms, including results on downstream applications.
Zero-shot Object Navigation demands the agent to explore unfamiliar environments and navigate to certain objects. In this paper, we propose a new framework that combines Explicit and Implicit Knowledge Navigation (EIK-Nav) to facilitate the overall zero-shot object navigation from both the exploration and detection process. The implicit knowledge is derived from the pre-trained vision-language model (VLM) that measures the semantic distance between the textual target category and RGB images, and thus provides a baseline value for determining the exploration potential of a specific area. The explicit knowledge is provided by the large language model (LLM), consisting of the co-occurrence relation among the target, the room, and other objects. EIK-Nav uses explicit knowledge as an extra constraint over the implicit value to decide where the most valuable area to explore is. In the detection process, the explicit co-occurrence relation can further offer additional evidence for determining the confidence of the detected target, which can elevate the detection process from being independent to considering the surrounding context. The experimental results on photo-realistic environments, including Matterport 3D (MP3D), Habitat-Matterport 3D (HM3D), and Gibson demonstrate that our proposed method achieves state-of-the-art results even compared with training-based methods.
Embodied Question Answering (EQA) requires agents to explore the environment actively and gather reasonable visual information to answer the questions raised by users. Despite the significant progress attained, previous works lack a comprehensive understanding of the observed environment. Particularly, in the question answering task, little attention is paid to the temporal context in multiple observations and the semantic understanding of the environment. Meanwhile, the navigator struggles when combining global spatial context and individual observations to comprehend the geometric environment. In this paper, we propose Context-aware Embodied Question Answering (CaEQA) to address the issues from two aspects. Firstly, to integrate temporal context into image embeddings obtained from the local observations at the navigation endpoint, we propose a local temporal context-aware visual question answering module (LTC-VQA). Specifically, we utilize the large language model in the question answering module to enhance the temporal and semantic understanding of complex unseen environments by introducing general knowledge obtained in the pretraining stage. Secondly, to model the complex interdependence within the overall navigation procedure, we propose a global-geometric-guided navigation module (G3-Nav), facilitating the interpretation of global historical information. In G3-Nav, we use a transformer that incorporates RGB images, depth maps and actions from all past positions to provide geometric hints from inter-image and intra-image perspectives. Experimental results on the Matterport3D and House3D datasets demonstrate that our method outperforms the SoTA methods in terms of navigation, question answering and the overall EQA system.
In embodied vision, Goal-Oriented Navigation (GON) requires robots to locate a specific goal within an unexplored environment. The primary challenge of GON arises from the need to construct a Bird's-Eye-View (BEV) map to understand the environment while simultaneously localizing an unobserved goal. Existing map-based methods typically employ self-centered semantic maps, often facing challenges such as reliance on complete maps or inconsistent semantic association. To this end, we propose Plug-and-Play Label Map Diffusion (PLMD), which defines a novel map completion diffusion model based on Denoising Diffusion Probabilistic Models (DDPM). PLMD generates obstacle and semantic labels for unobserved regions through a diffusion-based completion process, thereby enabling goal localization even in partially observed environments. Moreover, it mitigates inconsistent semantic association by leveraging structural consistency between known and unknown obstacle layouts and integrating obstacle priors into the semantic denoising process. By substituting predicted labels for unobserved regions, robots can accurately localize the specified objects. Extensive experiments demonstrate that PLMD \textbf{(I)} effectively expands the region of unknown maps, \textbf{(II)} integrates seamlessly into existing navigation strategies that rely on semantic maps, \textbf{(III)} achieves state-of-the-art performance on three GON tasks.
Absolute Pose Regression (APR) encompasses a spectrum of visual localization methods that directly regress the 6-DoF camera pose from input images. Previous APR methods typically rely on 3DGS rendered features that smooth out structural details, leading to ambiguous scene descriptors. To address this limitation, we introduce surface curvature as explicit multi-view epipolar geometric cues that capture stable, detailed variations across viewpoints and provide structurally reliable cues for pose estimation. Specifically, we adopt a 3D Gaussian Splatting (3DGS) representation equipped with surface curvature for the scene, and introduce a novel refinement framework termed CurvLoc. Within this framework, a Surface Curvature Extractor is designed to capture curvature information from rendered features along epipolar line directions. Additionally, we propose a Curvature-aware Sampling Strategy that prioritizes regions exhibiting the largest curvature, effectively exploiting multi-view information. This approach significantly enhances geometric detail awareness and delineates clear boundaries in complex regions, facilitating precise visual localization. Extensive experiments on indoor and outdoor visual localization benchmarks demonstrate that the proposed CurvLoc framework surpasses existing state-of-the-art methods in accuracy and robustness.
Real-time and precise traffic flow prediction serves as a vital enabler for modern intelligent transportation systems. In this domain, graph neural networks demonstrated remarkable proficiency in capturing the intricate spatiotemporal relationships embedded within traffic data. However, we identify a critical limitation in existing attention-based dynamic GNNs: their inadequate modeling of hub nodes, which maintain extensive connections and exhibit more intricate spatiotemporal patterns than ordinary nodes in transportation networks.To address this gap, we propose the Hub Node-enhanced Attention Fusion Dynamic Graph Convolutional Network (HN-AT-DGCN), a novel encoder-decoder framework comprising four essential components. The multi-scale feature fusion module first extracts traffic flow characteristics across varied temporal scales. The encoder then employs a dynamic graph convolutional gated recurrent unit to capture comprehensive spatiotemporal dependencies. A multi-head temporal attention mechanism further models long-range temporal patterns, while a 2D graph convolutional decoder ultimately generates future traffic flow predictions.Our key contributions are twofold. We introduce a hub node identification module that automatically detects critical nodes in the network. Our methodology incorporates a novel dynamic graph convolution scheme that facilitates discriminative feature learning across three semantic levels: node-wise properties, structural correlations, and hub-dominated propagations, comprehensively modeling multi-faceted relationships in traffic networks.The effectiveness of our method is evidenced by thorough evaluations across three authentic traffic datasets, where it attains leading performance and exceeds 15 baseline approaches.
Object detection in Unmanned Aerial Vehicle (UAV) imagery, a specialized subfield of visual perception, becomes substantially more challenging under nighttime conditions. In addition to inherent issues such as multi-scale targets and densely distributed small objects, low illumination reduces image visibility and affects detection accuracy. To address these challenges, we propose MFLDet, an end-to-end framework that integrates low-light image enhancement with object detection, specifically designed for nighttime UAV imagery. Unlike existing approaches, this work is the first to embed a low-light enhancement strategy that first restores brightness and then performs denoising within an end-to-end detection framework for nighttime UAV scenarios. To implement this strategy, the framework incorporates a Low-Light Enhancement Network (LNet) to generate informative enhancement features, which are fused with detection features at the feature level. Beyond the enhancement stage, a Local and Global Feature Enhancement (LGFE) module enhances attention to object regions while suppressing background interference. Using compact proxy representations, LGFE achieves a balance between computational efficiency and long-range dependency modeling. A Multi-Scale Feature Fusion Module (MSFM) further enriches semantic representations to improve small-object perception in complex nighttime scenes. Extensive experiments on curated nighttime subsets of the VisDrone and DroneVehicle datasets demonstrate that MFLDet outperforms state-of-the-art methods, achieving an mAP50 of 46.9
Object goal navigation, which involves autonomous navigation towards an object goal in an unseen environment, is a fundamental task for embodied AI agents. Recent works tackle this task mainly through solely reinforcement learning (RL) based methods or large language model (LLM) based methods. However, RL-based methods rely solely on historical trial-and-error experience, suffering from poor generalization to new environments. And LLM-based methods fall short in terms of reliable decision-making and cost-efficiency. To address these challenges, we propose RLLMNav, a method that integrates historical experience of RL with commonsense reasoning of LLM through a confidence-gated routing mechanism. Rather than replacing either paradigm, RLLMNav coordinates them so that each compensates for the other’s failure modes. Specifically, we utilize commonsense knowledge extracted from an LLM to suggest frontiers. We also use this form of knowledge in a dual-module strategy to select frontiers from suggestions as long-term goals to explore, where an LLM-based policy for accessing LLM commonsense knowledge improves generalization, and an RL-based policy for learning historical experience improves reliable decision-making and cost-efficiency. Extensive experiments on Gibson and Habitat-Matterport 3D (HM3D) demonstrate that RLLMNav achieves state-of-the-art results, validating the potential of combining RL and LLMs for ObjectNav. Code is available at https://github.com/ckx666/RLLMNav .
Vision-Language Models (VLMs) have demonstrated exceptional general reasoning capabilities. However, their performance in embodied navigation remains hindered by a scarcity of aligned open-world vision and robot control data. Despite simulators providing a cost-effective alternative for data collection, the inherent reliance on photorealistic simulations often limits the transferability of learned policies. To this end, we propose \textit{\textbf{S}andbox-\textbf{A}bstracted \textbf{G}rounded \textbf{E}xperience} (\textbf{\textit{SAGE}}), a framework that enables agents to learn within a physics-grounded semantic abstraction rather than a photorealistic simulation, mimicking the human capacity for mental simulation where plans are rehearsed in simplified physics abstractions before execution. \textit{SAGE} system operates via three synergistic phases: (1) \textit{Genesis}: constructing diverse, physics-constrained semantic environments to bootstrap experience; (2) \textit{Evolution}: distilling experiences through Reinforcement Learning (RL), utilizing a novel asymmetric adaptive clipping mechanism to stabilize updates; (3) \textit{Navigation}: bridging the abstract policy to real-world control. We demonstrate that \textit{SAGE} significantly improves navigation performance, achieving a 53.21\% LLM-Match Success Rate on A-EQA (+9.7\% over baseline) and generalizing widely to real-world deployments.
Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this challenge by employing flattened Gaussian primitives to better fit surface geometry, combined with depth regularization to alleviate geometric ambiguities under limited viewpoints. Nevertheless, the increased anisotropy inherent in flattened Gaussians exacerbates overfitting in sparse-view scenarios, hindering accurate surface fitting and degrading novel view synthesis performance. In this paper, we propose SparseSurf, a method that reconstructs more accurate and detailed surfaces while preserving high-quality novel view rendering. Our key insight is to introduce Stereo Geometry-Texture Alignment, which bridges rendering quality and geometry estimation, thereby jointly enhancing both surface reconstruction and view synthesis. In addition, we present a Pseudo-Feature Enhanced Geometry Consistency that enforces multi-view geometric consistency by incorporating both training and unseen views, effectively mitigating overfitting caused by sparse supervision. Extensive experiments on the DTU, BlendedMVS, and Mip-NeRF360 datasets demonstrate that our method achieves the state-of-the-art performance.
With advancements in robust stereo matching and optical flow estimation networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, their robustness can be seriously degraded when fine-tuning them in real-world scenarios. This paper investigates fine-tuning stereo matching and optical flow estimation networks without compromising their robustness to unseen domains. Specifically, we divide the pixels into consistent and inconsistent regions by comparing Ground Truth (GT) with Pseudo Label (PL) and demonstrate that the imbalance learning of consistent and inconsistent regions in GT causes robustness degradation. Based on our analysis, we propose the DKT framework, which utilizes PL to balance the learning of different regions in GT. The core idea is to utilize an exponential moving average (EMA) teacher to measure what the student network has learned and dynamically adjust the learning regions. We further propose the DKT++ framework, which improves target-domain performances and network robustness by applying slow-fast update teachers to generate more accurate PL, introducing the unlabeled data and synthetic data. We integrate our frameworks with state-of-the-art networks and evaluate their effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the robustness of stereo matching and optical flow networks during fine-tuning.
Recent advancements in Large Language Models (LLMs) offer powerful capabilities in common sense reasoning and planning, making them a promising tool for Object Goal Navigation (ObjectNav). However, existing LLM-based approaches face two significant challenges: the high computational cost of LLM inference, which limits real-time decision making, and a domain gap between the LLMs’ general-purpose knowledge and the specific demands of navigation scenarios. To overcome these challenges, we propose a Knowledge-Enhanced navigation framework with an Intuitive-Deliberate mechanism (KEID). KEID employs an Intuitive-Deliberate mechanism that mimics human cognition, using a lightweight intuition module to strategically invoke the LLM, which reduces computational overhead. Meanwhile, KEID enhances the LLM with two specialized knowledge bases: a Scene Description Tree that describes the complex spatial and semantic relationships of indoor environments within a hierarchical framework and a Navigation Example database for in-context learning adaptation. Evaluations on the HM3D dataset within the Habitat simulator validate our method’s efficacy, demonstrating that KEID achieves a 47.1% success rate and a competitive 18.8% success weighted by path length, significantly outperforming existing baselines. Our work not only improves navigation performance but also enhances decision-making efficiency, establishing an effective framework for developing practical, real-time LLM-based robotic agents.
Embodied Question Answering (EQA) is a task in artificial intelligence where an intelligent agent is required to answer questions about its environment. For example, to answer a question such as "Is the TV on or off?", the agent must navigate to the room with the TV and answer with either "On." or "Off." after recognizing the status. Unlike traditional question-answering systems that rely solely on text or static images, EQA involves agents that can move through a physical or simulated space, interact with the environment, and gather information to respond accurately. The agent must interpret both visual and linguistic inputs, navigate the environment, and complete tasks or locate objects based on the user’s questions. However, in the real world, the agent always faces unseen environments (i.e. different people’s houses), which makes the pre-trained model fail. Meanwhile, re-training in an unseen environment can cause high costs. Therefore, it is significant for the agent to learn continually by itself to cope with the challenges of unseen environments. In this work, we proposed a continual learning method based on generative adversarial imitation learning and self-supervision to support the agent when facing unseen environments. Besides, we designed a policy generator and policy quality discriminator to generate action policy sequences and evaluate the quality of the policy, respectively. Extensive experiments on the MP3D-EQA dataset demonstrate that our method reaches state-of-the-art performance.
In the scope of 3D human pose estimation, the task encompasses estimating the 3D positions of key skeletal points (i.e., wrists, elbows, and knees) from a 2D image or video sequence. This technology demonstrates widespread applicability across diverse domains, encompassing domains such as kinematic analysis, virtual reality, augmented reality, and medical imaging analysis. The common approach is divided into two stages: i) 2D Keypoint Detection: Detecting 2D keypoints from images. ii) 2D -to -3D Lifting: Converting 2D keypoints into 3D coordinates. Present research predominantly concentrates on Stage 2 and leverages sophisticated deep learning architectures, notably Transformers, yet two key challenges exist: i) suboptimal performance in predicting local details; ii) susceptibility to noise interference. In this work, we propose a time-principal component fusion model with limb segment property tracking is proposed to address the challenges above. By integrating time domain features with principal components extracted through the Karhunen-Loeve Transform (KLT), the model aims to address challenges related to feature extraction and noise reduction. Furthermore, to address the issue of suboptimal performance in predicting local details, we devise a property transformer to track the lengths of limb segments and predict the fixed property. Extensive experiments demonstrate that KLFormer showcases state-of-the-art performance on the standard benchmark dataset, Human3.6M.
Instance ImageGoal Navigation (IIN) entails an agent autonomously seeking out a specific object instance depicted by a goal image in an unknown environment. While Large Language Models (LLMs) have shown promise in navigation tasks similar to IIN, their application to IIN remains unexplored. Furthermore, existing LLM-based exploration faces challenges such as inaccurate reasoning due to informative environmental information available to the agent, especially in the early episode stages, and the inability of reinforcement learning(RL) exploration to fully leverage gathered information. Moreover, previous IIN methods did not productively verify potentially distant goal objects discovered during exploration. This work proposes MoPE–Mixture of Policy Experts for exploration and potential goal verification with multimodal information when exploring. Specifically, the hybrid exploration policy comprises an LLM and an RL-based Policy Network (RLPN) to generate an exploration goal to explore efficiently. Our MoPE model surpasses prior approaches on the HM3D datasets significantly.
Reinforcement learning(RL) has made significant advancements in autonomous driving(AD). However, the stochastic nature of dynamic traffic scenario and the diversity of road type make it challenging for autonomous vehicles to make safe and efficient decisions. To tackle these problems, this paper proposes a novel RL framework that incorporates the motion prediction model to enhance the agent’s decision-making capability. We first utilize Transformer to model driving scenarios and capture interaction-aware relationships between the ego vehicle and scenarios, then design a safety-constraint and integrate it into the Proximal Policy Optimization (PPO) algorithm so as to guarantee the safety and feasibility of the policy. To improve data efficiency and filter noisy samples, we construct a dual network to communicate and guide each other. Experimental results show that compared with popular RL algorithms, our method demonstrates superior performance in success rate, completion time, safety, and data efficiency.
Object Goal Navigation is a fundamental task for embodied agents to accomplish complex tasks in real-world environments. In this task, the agent is required to navigate to a specific target in unseen environments where prior maps are unavailable. Due to the target being invisible from the initial position, the agent needs to have the human-like ability to reason about the location of the target and conduct efficient exploration. To obtain this reasoning ability, reinforcement learning (RL) methods [ 1] adopted an end-to-end approach in the navigation process and modular-based methods [ 2] built an explicit map based on geometric and semantic observations.
Reward acts as a signal to guide the agent's learning process in Reinforcement Learning (RL), evaluating and assigning rewards to the agent's actions based on their alignment with goals. Designing reward is challenging in multi-agent environment such as StarCraft II benchmark since agents face credit allocation and role adaptation problems. Recent studies have successfully exploited the language understanding and reasoning capabilities of large language models (LLMs) to learn manipulation tasks. Impressed by the remarkable power of LLMs, this paper employs LLMs as role-specific reward designer for playing StarCraft II, making rewards more flexible and task-oriented. Firstly, we develop an interactive text and multi-agent RL environment to study real-time strategy generation in StarCraft II. Secondly, we use LLMs to interpret the game situation and understand agent roles from user instructions. Then, by assigning appropriate subtasks, LLMs quantify the completion of these subtasks to generate role-specific rewards. Further, credit assignment problem is addressed by introducing dynamic reward weights in value decomposition method. In StarCraft II maps, experiments show that role-aligned RL agents trained with our framework achieve superior policy performance, and win rate results demonstrates the effectiveness of our approach in decision-making for micromanagement and long-term planning.