Visual place recognition aims to identify previously visited locations from image representations that are robust to viewpoint and environmental changes. Deep convolutional neural network features perform well for this task, yet they are typically pretrained for classification or segmentation – objectives that do not account for the sequential nature of place recognition. In mobile robotics, exploiting image sequences can further improve performance. We propose softmax-based fine-tuning of a convolutional network extended with a recurrent model to enhance place recognition on image sequences. Experiments on two public datasets show that the proposed representation consistently outperforms competing approaches across all evaluated methods.
Robotic manipulation from instructional videos remains challenging because fine-grained hand actions are highly variable, context-dependent, and difficult to decompose into representations that are both perceptually robust and usable for planning. Subtle cues such as grasp strength, contact surface, and post-release object behavior are often temporally brief and sensitive to viewpoint and scene variability, making flat gesture recognition insufficient for learning-by-demonstration. We study hierarchical recognition of hand-object interactions and develop a gesture recognition system, called GCPR, that predicts a primary Gesture together with Power (grasp strength), Contact (pad/palm/side), and Release (stationary/passive/active). The system uses a time-aware transformer backbone with a configurable preprocessing stack comprising motion-aware frame selection, hand localization with Kalman stabilization, and a gated object-aware vector derived from YOLOv5 detections. Through controlled ablations of four preprocessing configurations, we isolate the effects of spatial alignment, temporal smoothing, and object context. Results show near-ceiling Gesture performance, consistent Contact gains from the full stack, Power favoring wider contextual views, and Release saturating when post-contact dynamics remain visible, while object gating provides selective improvements without brittleness. The system described in this paper builds on our previous exploratory work by developing a complex and mature pipeline, evaluated under stricter data and testing protocols, yielding an improved robustness and clearer attribution of module effects, and crucially adding an object awareness module based on YOLOv5.
The ability to selectively attend to relevant stimuli while filtering out distractions is essential for agents that process complex, high-dimensional sensory input. This paper introduces a model of covert and overt visual attention through the framework of active inference, utilizing dynamic optimization of sensory precisions to minimize free-energy. This work addresses the lack of active inference models that integrate visual attention with continuous sensory representations and deep generative models for robotics. Our proposed model determines visual sensory precisions based on both current environmental beliefs and sensory input, influencing attentional allocation in both covert and overt modalities. To test the effectiveness of the model, we analyze its behavior in the Posner cueing task and a simple target focus task using two-dimensional (2D) visual data. Reaction times are measured to investigate the interplay between exogenous and endogenous attention, as well as valid and invalid cueing. The results show that exogenous and valid cues generally lead to faster reaction times compared to endogenous and invalid cues. Finally, we show that reflexive saccades are faster than intentional ones, though less adaptable, and discuss the implications for robotic applications.
Bimanual manipulation enables complex tasks but introduces added complexity from the high number of degrees of freedom involved. When handling rigid objects, the relative transformation between the two end effectors must remain fixed throughout the motion, manifesting as a nonlinear equality constraint that confines the feasible configuration space to a measure-zero manifold and challenges conventional motion planners. We propose a fast bimanual motion planning pipeline that enforces this hard transformation constraint continuously along the entire path, using a leader-follower parameterization: the leader's configuration is treated as a free variable, while the follower's is determined via inverse kinematics to satisfy the constraint. We extensively evaluate the method in simulation across diverse environments, constraints and bimanual platforms, achieving 19.4x faster planning than prior work while guaranteeing continuous constraint satisfaction. Real-world experiments on a bimanual Kinova Gen3 setup, involving tray transport and elongated-object manipulation, validate direct transfer of planned trajectories to physical hardware.
Estimating an object’s distance from electromagnetic wave interactions is critical for robotic perception - especially in scenarios with occlusion or degraded visibility. Traditional radar systems rely on hand-crafted signal processing pipelines which limit adaptability to novel sensor configurations. In this work, we present a hybrid approach that combines lightweight signal processing with a data-driven inverse mapping from transmit-receive wave measurements to object distance. We first transform raw signals, FMCW radar intermediate-frequency signals and S-parameters, into the frequency and time domain, respectively, applying peak detection and Hamming windowing to extract propagation path features. These features are then used to train Gaussian Process Regression models that predict object coordinates with uncertainty estimates. We evaluate the method on both simulated S-parameters and measured radar data designed to emulate future flexible electromagnetic skins. The results demonstrate enhanced sensing modalities and inherently learned geometry of the sensor, highlighting the potential of combining minimal signal preprocessing with probabilistic learning for general-purpose robotic electromagnetic perception.
Learning-based monocular depth estimation leverages geometric priors present in the training data to enable metric depth perception from a single image, a traditionally ill-posed problem. However, these priors are often specific to a particular domain, leading to limited generalization performance on unseen data. Apart from the well studied environmental domain gap, monocular depth estimation is also sensitive to the domain gap induced by varying camera parameters, an aspect that is often overlooked in current state-of-the-art approaches. This issue is particularly evident in autonomous driving scenarios, where datasets are typically collected with a single vehicle-camera system, leading to a bias in the training data due to a fixed perspective geometry. In this paper, we challenge this trend and introduce GenDepth, a novel model capable of performing metric depth estimation for arbitrary vehicle-camera systems. To address the lack of data with sufficiently diverse camera parameters, we first create a bespoke synthetic dataset collected with different vehicle-camera systems. Then, we design GenDepth to simultaneously optimize two objectives: (i) equivariance of estimated depth to the camera parameter variations on synthetic data, (ii) transferring the learned equivariance to real-world environmental features using a single real-world dataset with a fixed vehicle-camera system. To achieve this, we propose a novel embedding of camera parameters as the ground plane depth and present a novel architecture that integrates these embeddings with adversarial domain alignment. We validate GenDepth on several autonomous driving datasets, demonstrating its state-of-the-art generalization capability.
Radar-Inertial Odometry (RIO) has emerged as a robust alternative to vision- and LiDAR-based odometry in challenging conditions such as low light, fog, featureless environments, or in adverse weather. However, many existing RIO approaches assume known radar-IMU extrinsic calibration or rely on sufficient motion excitation for online extrinsic estimation, while temporal misalignment between sensors is often neglected or treated independently. In this work, we present a RIO framework that performs joint online spatial and temporal calibration within a factor-graph optimization formulation, based on continuous-time modeling of inertial measurements using uniform cubic B-splines. The proposed continuous-time representation of acceleration and angular velocity accurately captures the asynchronous nature of radar-IMU measurements, enabling reliable convergence of both the temporal offset and extrinsic calibration parameters, without relying on scan matching, target tracking, or environment-specific assumptions.
The operational reliability of an autonomous robot depends crucially on extrinsic sensor calibration as a prerequisite for precise and accurate data fusion. Exploring the calibration of unscaled sensors (e.g., monocular cameras) and the effective utilization of uncertainties are difficult and often overlooked. The development of a solution for the simultaneous calibration of hand-eye sensors and scale estimation based on the Gauss-Helmert model aims to utilize the valuable information contained in the uncertainty of odometry. In this work, we propose a versatile and robust solution for batch calibration based on the analytical on-manifold approach for estimation. The versatility of our method is demonstrated by its ability to calibrate multiple unscaled and metric-scaled sensors while dealing with odometry failures and reinitializations. Importantly, all estimated parameters are provided with their corresponding uncertainties. The validation of our method and its comparison with five competing state-of-the-art calibration methods in both simulations and real-world experiments show its superior accuracy, with particularly promising results observed in high-noise scenarios.
In this paper we present a novel vehicle localization dataset captured with a pair of stereo rolling shutter cameras mounted on top a race car while driving at speeds of more than 200 km/h. The dataset contains six sequences recorded on two different racetrack coupled with accurate ground truth trajectories recorded with a real-time kinematics GNSS system. We make the dataset publicly available to the research community along with raw calibration sequences and our estimated calibration parameters. Besides the dataset, in the paper we also propose an extended version of our previously published stereo odometry, SOFT2, to the case of rolling shutter stereo cameras. The approach successfully models the rolling shutter effect and addresses the challenges both in the front- and back-end part of the odometry due to the line-by-line scanning nature of the sensor during image capture. Specifically, for the former the challenge of feature matching in a texture-poor environment of a racetrack, while for the later the challenge of not having a unique essential matrix for all feature correspondences as in the global shutter case. The proposed odometry is evaluated on the six sequences and achieves accuracy on par with the state-of-the-art methods of the KITTI benchmark.The dataset is available at https://unizgfer-lamor.github.io/rt-dataset.
Automated cargo transport is a frequent task in intralogistics and industrial scenarios, but is usually carried out by a single robot, thus limited to its payload capabilities. In this paper, we present a method for coordinated motion planning and control of two autonomous mobile robots for the joint transport of cargo that is too large for a single robot. We consider the two robots with a large cargo on top of them as an any-shape single kinematic chain with cargo. We plan an optimal collision free path for the entire differential drive kinematic chain that ensures the transport of the cargo to the desired target location. The implemented algorithm sends control commands to each robot separately, which achieves the coordinated execution of the planned path. Finally, the proposed method is validated in simulated and real-world scenarios.
Ego-motion estimation is an indispensable part of any autonomous system, especially in scenarios where wheel odometry or global pose measurement is unreliable or unavailable. In an environment where a global navigation satellite system is not available, conventional solutions for ego-motion estimation rely on the fusion of a LiDAR, a monocular camera and an inertial measurement unit (IMU), which is often plagued by drift. Therefore, complementary sensor solutions are being explored instead of relying on expensive and powerful IMUs. In this paper, we propose a method for estimating ego-motion, which we call MOVRO2, that utilizes the complementarity of radar and camera data. It is based on a loosely coupled monocular visual radar odometry approach within a factor graph optimization framework. The adoption of a loosely coupled approach is motivated by its scalability and the possibility to develop sensor models independently. To estimate the motion within the proposed framework, we fuse ego-velocity of the radar and scan-to-scan matches with the rotation obtained from consecutive camera frames and the unscaled velocity of the monocular odometry. We evaluate the performance of the proposed method on two open-source datasets and compare it to various mono-, dual- and three-sensor solutions, where our cost-effective method demonstrates performance comparable to state-of-the-art visual-inertial radar and LiDAR odometry solutions using high-performance 64-line LiDARs.
Advanced testing and validation in the automotive industry are predominantly conducted through real-world driving, while simulation is primarily used for scenarios that are difficult, dangerous, or impractical to reproduce on the road. The production of a well-designed simulation scenario includes numerous edge cases and relies on expert labor, but it can still miss real-world complexities. Given that, scenarios based on real-world driving data are more desirable, but are dominantly based on replaying the positions of all vehicles over time. For greater flexibility, real-world driving scenarios should be automatically converted into logic scenarios with adjustable segments such as lane keeping, lane changing, accelerating, etc. This paper introduces a new method for detecting and classifying lateral vehicle movements. The method is based on polynomial fitting and identifies the start and end points of vehicle lane changes with inflection points of a third-degree polynomial. For evaluation, two publicly available, manually labeled lane-change datasets are used, and a novel highway lane-change simulation dataset is introduced. The polynomial approach is compared against a bidirectional long short-term memory (Bi-LSTM) model introduced in this work, as well as three additional neural network architectures drawn from recent literature, followed by a discussion of the experimental validation for each.
Detecting decalibration is crucial to maintain perception accuracy and prevent performance degradation in long-term autonomous operations, where environmental changes, mechanical wear, and external forces can compromise sensor alignment. In this paper, we propose a motion-based framework that leverages uncertainty-aware optimization for accurate decalibration detection and correction. Key contributions include a batch calibration procedure based on the Gauss-Helmert model with variance component estimation and a probabilistic detection criterion using the Mahalanobis distance. Finally, we develop an online detection system employing a sliding-window approach for continuous monitoring, demonstrating particularly rapid response to rotational misalignment. Copyright (c) 2025 The Authors. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/)
Reinforcement learning has shown remarkable potential for autonomous skill acquisition. However, effective exploration-exploitation of possible actions and states remains a fundamental challenge, especially in soft robotic manipulation tasks where the continuous deformation of soft materials causes complex nonlinear dynamics and high-dimensional state spaces. A promising solution to this problem is intrinsic motivation, which allows learning agents to explore the environment more systematically by producing self-supervised reward signals. Existing intrinsic reward techniques, however, frequently approach the whole state space as a single entity, which may reduce their efficacy in robotic manipulation, where interactions can take place between the manipulated item and the robot. In this work, we introduce a new method of intrinsic reward decomposition which focuses exploration on task-relevant interactions. Our method implements a weighted combination of random network distillation rewards derived separately from robot observations and manipulation object states. Experimental results across various manipulation tasks in the soft robotic benchmark show that this attention-inspired decomposition enables more effective discovery of manipulation strategies and significantly enhances performance.
Accurate localization and scene reconstruction are essential for the autonomous navigation of mobile agents. Simultaneous Localization and Mapping (SLAM) algorithms address both challenges by formulating a unified optimization problem, offering an integrated solution to both objectives. Recent advances in learning-based scene understanding have significantly improved accuracy and robustness, particularly in adverse scenarios that are troublesome for traditional geometric methods. However, generating an accurate dense scene reconstruction remains an open challenge, largely due to the complexity of the optimization problem, making it unsuitable for real-time requirements on resource-constrained devices. Novel advances in 3D reconstruction such as implicit representations and Gaussian Splatting present an enticing formulation enabling offline reconstruction of large-scale scenes. While these approaches have been successfully adapted for online incremental reconstruction, particularly through Gaussian Splatting SLAM methods, they are hindered by significant computational complexity and convergence challenges due to the non-convex nature of photometric optimization. In this work we rethink this approach by combining the strengths of traditional feature-based methods with innovative reconstruction capability of Gaussian splatting. Specifically, we integrate feature-based pose estimation, relocalization and loop closure with 3D Gaussian-based scene reconstruction. This results in state-of-the-art tracking and mapping performance on the EuRoC and TUM datasets, while significantly reducing convergence iterations and improving real-time performance.
Generalizing metric monocular depth estimation presents a significant challenge due to its ill-posed nature, while the entanglement between camera parameters and depth amplifies issues further, hindering multi-dataset training and zero-shot accuracy. This challenge is particularly evident in autonomous vehicles and mobile robotics, where data is collected with fixed camera setups, limiting the geometric diversity. Yet, this context also presents an opportunity: the fixed relationship between the camera and the ground plane imposes additional perspective geometry constraints, enabling depth regression via vertical image positions of objects. However, this cue is highly susceptible to overfitting, thus we propose a novel canonical representation that maintains consistency across varied camera setups, effectively disentangling depth from specific parameters and enhancing generalization across datasets. We also propose a novel architecture that adaptively and probabilistically fuses depths estimated via object size and vertical image position cues. A comprehensive evaluation demonstrates the effectiveness of the proposed approach on five autonomous driving datasets, achieving accurate metric depth estimation for varying resolutions, aspect ratios and camera setups. Notably, we achieve comparable accuracy to existing zero-shot methods, despite training on a single dataset with a single-camera setup. Project website: https://unizgfer-lamor.github.io/gvdepth/
Accurate ego-motion estimation is a critical component of any autonomous system. Conventional ego-motion sensors, such as cameras and LiDARs, may be compromised in adverse environmental conditions, such as fog, heavy rain, or dust. Automotive radars, known for their robustness to such conditions, present themselves as complementary sensors or a promising alternative within the ego-motion estimation frameworks. In this paper we propose a novel Radar-Inertial Odometry (RIO) system that integrates an automotive radar and an inertial measurement unit. The key contribution is the integration of online temporal delay calibration within the factor graph optimization framework that compensates for potential time offsets between radar and IMU measurements. To validate the proposed approach we have conducted thorough experimental analysis on real-world radar and IMU data. The results show that, even without scan matching or target tracking, integration of online temporal calibration significantly reduces localization error compared to systems that disregard time synchronization, thus highlighting the important role of, often neglected, accurate temporal alignment in radar-based sensor fusion systems for autonomous navigation. Project website: https://rio-online-t.github.io/.
Addressing the challenges of autonomous racing in Formula Student competitions, this work proposes a novel LiDAR-based perception and trajectory optimization pipeline. Our key contributions include a robust cone detection algorithm for accurate track-edge identification from point clouds. These detections are integrated into a global map using an adapted FASTSLAM 2.0 framework, enhancing mapping accuracy in dynamic environments. Furthermore, we introduce a heuristic trajectory planner that minimizes path length and curvature trade-offs via cubic splines and a BFGS-based solver, and crucially, an offline regression model for real-time optimal weight estimation. The complete system is validated on real-world tracks with RTK-based ground truth, demonstrating superior mapping precision and efficient trajectory generation for high-performance autonomous navigation.