In stealth-constrained swarm robotics, visual communication provides a critical alternative to active radio transmissions, which might be jammed. This research investigates motion-based communication for non-active information exchange, utilizing modular, dynamically feasible planar trajectories as visual cues. On the receiver drone end, a pose estimator tracks the transmitting drone's pose, feeding it into our custom 3DTrajDecoder. The decoder is designed to classify and segment the spatiotemporal sequence while simultaneously regressing its size and normal vector. To robustly train the decoder on both communicative and non-communicative trajectories, we developed a configurable online procedural generation pipeline. We validate our system through real-world testing and simulation to define its operating domain, supported by an extensive ablation study detailing our architectural choices and system limitations.
This paper studies collaborative exploration of an initially uncharted environment, employing a tandem composed of an unmanned aerial vehicle and an unmanned ground vehicle. The proposed method harnesses the complementary capabilities of both platforms, which exhibit distinct heterogeneous characteristics, to enhance exploration performance. The UAV offers high maneuverability and a broad aerial vantage point but with restricted payload while the UGV, despite its kinematic limits, provides ground-level stability, extended sensing capabilities, and superior payload capacity. The approach takes advantage of their combined strengths to improve exploration and coverage with two distinct strategies based on Next Best View and frontier exploration. A frontier extraction and redefinition method is proposed in order to limit the number of candidate viewpoints to the most reachable ones, which also speeds up evaluation. Simulations were conducted to test and discuss the scenarios and to highlight the practical relevance of the proposed system.
Pose estimation plays a crucial role in robotics for prehension tasks or in augmented-reality application, yet its application on real-world far-range estimation has not been thoroughly studied. This study aims to evaluate pose estimators on a custom drone at distances from 0.5 m to 10 m, which is beyond the scope of existing datasets, that only contain objects close to less than 2 m. We created synthetic and real databases specific to our drone and compared various RGB pose estimators, evaluating their performance across different distances. PViT-6D, being one of the SoTA methods on the classic [0,2] m interval, also outperforms others estimators at greater distances, and proves robust with respect to detection inaccuracy. The results demonstrate the potential of PViT-6D to be used on a real time application embedded in the drone platform. This work aims to evaluate the potential of pose estimators for mutual perception and communication within a drone swarm.
The problem of safe navigation of a human-multi-robot system is addressed in this paper. More precisely, we propose a novel distributed algorithm to control a swarm of unmanned ground robots interacting with human operators in presence of obstacles. Contrary to many existing algorithms that consider formation control, the proposed approach results in non-rigid motion for the swarm, which more easily enables interactions with human operators and navigation in cluttered environments. Each vehicle calculates distributively and dynamically its own safety zone in which it generates a reference point to be tracked. The algorithm relies on purely geometric reasoning through the use of Voronoi partitioning and collision cones, which allows to naturally account for inter-robot, human-robot and robot-obstacle interactions. Different interaction modes have been defined from this common basis to address the following practical problems: autonomous waypoint navigation, velocity-guided motion, and follow a localized operator. The effectiveness of the algorithm is illustrated by outdoor and indoor field experiments.
This article introduces and evaluates two decentralized data sharing algorithms for multi-robot visualinertial simultaneous localization and mapping (VI- SLAM): Factor Sparsification for Visual-Inertial Packets (FS-VIP) and Min-K-Cover Selection for Visual-Inertial Packets (MKCS-VIP). Both methods make robots regularly build and exchange data packets which describe the successive portions of their map, but rely on distinct paradigms. While FS- VIP builds on consistent marginalization and sparsification techniques, MKCS-VIP selects raw visual and inertial information which can best help to perform a faithful and consistent re-estimation while reducing the communication cost. Performances in terms of accuracy and communication loads are evaluated on multi-robot scenarios built on both available (EUROC) and custom datasets (SOTTEVILLE).
Many applications of Visual SLAM, such as augmented reality, virtual reality, robotics or autonomous driving, require versatile, robust and precise solutions, most often with real-time capability. In this work, we describe OV(2)SLAM, a fully online algorithm, handling both monocular and stereo camera setups, various map scales and frame-rates ranging from a few Hertz up to several hundreds. It combines numerous recent contributions in visual localization within an efficient multi-threaded architecture. Extensive comparisons with competing algorithms shows the state-of-the-art accuracy and real-time performance of the resulting algorithm. For the benefit of the community, we release the source code: https://github.com/ov2slam/ov2slam.
This advance now makes possible the practical use of computer vision in civil drones or aircraft, replacing human pilots. The question that naturally arises is then to provide a way to certify those types of systems at a given level of safety. The aim of the article is, firstly, to understand the gap between today’s computer vision systems and the current certification standards; secondly, to identify the key activities that must be fulfilled in order to make computer-vision systems certifiable and, thirdly, to explore some recent works related to these key activities.
This paper presents a tele-operation system that enables a MAV to be controlled on virtual surfaces by an unskilled operator using high-level inputs. These virtual surfaces can be placed relatively to the infrastructure to be inspected, in order to ensure safety of the flight and repeatability of the acquisition conditions of the inspection data (e.g. at a constant distance from the infrastructure). The architecture, interface and embedded controller of the tele-operation system are described, and results from flight experiments in an industrial warehouse are provided for three typical inspection scenarios of infrastructures.
A new distributed algorithm is presented for waypoint navigation of a multi-robot system. The proposed two-level architecture (reference generator and local controller) exploits Voronoi partitioning and purely geometric considerations to distributively generate references for each robot in order to ensure collision avoidance and convergence of the fleet to the waypoint. Flexibility in the obtained formation pattern is made possible by the algorithm, by not pre-fixing as usually done its geometric form. In addition, the gain tuning is easy and the setting allows to naturally obtain certain formation patterns and adjust the rigidity of the fleet. Moreover the distributed nature of the algorithm also allows robustness to online modification of the number of vehicles (in the fleet or within range of communication), also addressing the 2-robot scenario. Field experiments on ground mobile robots are provided to illustrate the performance of the algorithm.
This paper introduces a new dataset dedicated to multi-robot stereo-visual and inertial Simultaneous Localization And Mapping (SLAM). This dataset consists in five indoor multi-robot scenarios acquired with ground and aerial robots in a former Air Museum at ONERA Meudon, France. Those scenarios were designed to exhibit some specific opportunities and challenges associated to collaborative SLAM. Each scenario includes synchronized sequences between multiple robots with stereo images and inertial measurements. They also exhibit explicit direct interactions between robots through the detection of mounted AprilTag markers [1]. Ground-truth trajectories for each robot were computed using Structure-from-Motion algorithms and constrained with the detection of fixed AprilTag markers placed as beacons on the experimental area. Those scenarios have been benchmarked on state-of-the-art monocular, stereo and visual-inertial SLAM algorithms to provide a baseline of the single-robot performances to be enhanced in collaborative frameworks.
This article introduces a decentralized multi-robot algorithm for Simultaneous Localization And Mapping (SLAM) inspired from previous work on collaborative mapping [1]. This method makes robots jointly build and exchange i) a collection of 3D dense locally consistent submaps, based on a Truncated Signed Distance Field (TSDF) representation of the environment, and ii) a pose-graph representation which encodes the relative pose constraints between the TSDF submaps and the trajectory keyframes, derived from the odometry, inter-robot observations and loop closures. Such loop closures are spotted by aligning and fusing the TSDF submaps. The performances of this method have been evaluated on multi-robot scenarios built from the EuRoC dataset [2].
Cet article presente et compare deux methodes decentrali-sees de partage de cartes et de trajectoires en SLAM visuel-inertiel multi-robot. La premiere methode repose sur le cal-cul de facteurs visuel-inertiels marginalises et sparsifies, associes a des informations visuelles locales, tandis que la seconde s'appuie sur l'echange de sous-cartes associees a des informations purement structurelles. Ces deux me-thodes resultent de la transposition de deux algorithmes developpes par Paull et al. [1] et Schuster et al. [2] res-pectivement pour du SLAM sous-marin acoustique et du SLAM terrestre stereo. Leurs performances en termes de precision et de quantite de donnees echangees sont eva-luees sur des scenarios multi-robot elabores sur le jeu de donnees EUROC [3].
Tractable algorithms used for 6DOF visual-inertial odometry have decades-long history of estimation consistency issues. Those arise in particular in two well-studied filters: namely the EKF-SLAM and MSCKF. Recently, strong theoretical works linked the error-state of these filters with their consistency properties; these results led to the synthesis of far more consistent filters. In previous works, we have shown that using similar filter for the fusion of magneto-inertial sensors with optical ones improved classical visual-inertial navigation systems. The consistency of those novel magneto-visual-inertial filters were, however, not addressed until now. This work does. We apply invariance theory findings to the specific case of magneto-inertial odometry and magneto-visual-inertial odometry for the synthesis of a filter with interesting consistency properties. We describe thoroughly such an invariant filter, implement it and conduct experiments on carefully captured data from real sensor. By comparing the results of non-invariant, observability-constrained and invariant versions of the filter, we find that the invariant version (i) shows an error estimate that is consistent with observability of the system, (ii) is applicable in case of unknown heading at initialization, (iii) improves long-term behavior of the filter and (iv) exhibits a lower normalized estimation error. We experiment on challenging scenarios for regular visual-inertial pedestrian navigation systems.
An assisted MAV tele-operation system is proposed, where user reference inputs in speed and yaw are fed to a control scheme for tracking a predefined trajectory, and filtered in case of collision risks. The on-board implementation uses embedded stereo-vision to obtain real-time localization and mapping in the context of navigation in a GPS-denied cluttered environment. A dedicated multi-platform HumanSystem Interface based on Rosbridge has been developed. Illustrative field results of typical inspection missions in a SNCF (French railway company) indoor facility are reported.
In this paper, the use of a Moving Horizon Estimator (MHE) is investigated to address a class of state estimation problems dealing with multi-rate sensor fusion in presence of time-delayed measurements. As it makes use of a batch of past measurement and state estimates, MHE is indeed a good candidate to deal with "missing" measurements. Nevertheless, since Moving Horizon Estimation relies on solving online an optimization problem to compute the state estimate, its computational load may be prohibitive for practical implementation to fast dynamical systems. Therefore this paper proposes a computationaly efficient implementation scheme for a variable structure linear MHE dealing with multi-rate time-delayed measurements, in the case where an analytical solution of the underlying optimization problem can be found. A simulation example is considered for performance comparison of the proposed MHE with respect to several state-of-the-art estimators, in terms of accuracy and computation time.
Visual-inertial Navigation Systems (VINS) are nowadays used for robotic or augmented reality applications. They aim to compute the motion of the robot or the pedestrian in an environment that is unknown and does not have specific localization infrastructure. Because of the low quality of inertial sensors that can be used reasonably for these two applications, state of the art VINS rely heavily on the visual information to correct at high frequency the drift of inertial sensors integration. These methods struggle when environment does not provide usable visual features, such than in low-light of texture-less areas. In the last few years, some work have been focused on using an array of magnetometers to exploit opportunistic stationary magnetic disturbances available indoor in order to deduce a velocity. This led to Magneto-inertial Dead-reckoning (MI-DR) systems that show interesting performance in their nominal conditions, even if they can be defeated when the local magnetic gradient is too low, for example outdoor. We propose in this work to fuse the information from a monocular camera with the MI-DR technique to increase the robustness of both traditional VINS and MI-DR itself. We use an inverse square root filter inspired by the MSCKF algorithm and describe its structure thoroughly in this paper. We show navigation results on a real dataset captured by a sensor fusing a commercial-grade camera with our custom MIMU (Magneto-inertial Measurment Unit) sensor. The fused estimate demonstrates higher robustness compared to pure VINS estimate, specially in areas where vision is non informative. These results could ultimately increase the working domain of mobile augmented reality systems.
This paper aims to leverage magnetic information from a Magneto-Inertial Measurement Unit — an IMU sensor augmented with an array of magnetometers, called MIMU hereafter — in a vision/inertial navigation system (VINS). This ego-motion estimation problem is formulated as an optimization over a sliding window fusing data from the MIMU with features tracked in a monocular camera image stream. The novelty of our approach lies in the formulation of preintegrated magnetic measurements that are computed from successive measurements of the local variations of the magnetic field, in the line of the preintegration of IMU data introduced in [1]. The resulting magnetic error terms participate to the minimized cost function along with the classical reprojection and IMU error terms. Our experiments show the benefits of this fusion. On the one hand, the added magnetic information from the MIMU allows to complement the VINS in cases where vision does not provide useful information during an extended period of time; on the other hand, vision does extend the operational domain of the navigation system compared to a pure MIMU solution, in particular for the outdoor portions of the trajectory.
This demonstrator paper describes a flight-tested, fully integrated perception-control loop for trajectory tracking with obstacle avoidance by micro-air vehicles (MAV) in indoor cluttered environments. For this purpose, a stereo-vision system is combined with an inertial measurement unit to estimate the vehicle localization and build a 3D model of the environment on-board. Emphasis is placed on a model predictive control (MPC) algorithm for safe guidance in unknown areas using the perception information. It combines an analytical linear quadratic solution for trajectory tracking and an efficient discretization strategy for collision avoidance. Experimental results in a flying arena and at an industrial site provide an overview of the demonstrator capabilities.
We propose a complete loop (detection, estimation, avoidance) for the safe navigation of an autonomous vehicle in presence of dynamical obstacles. For detecting moving objects from stereo images and estimating their positions, two algorithms are proposed. The first one is dense and has a high computational load but is designed to fully exploit GPU processing. The second one is lighter and can run on a standard embedded processor After a step of filtering, the estimated mobile objects are exploited in a model predictive control scheme for collision avoidance while tracking a reference trajectory. Experimental results with the complete loop are reported for a micro-air vehicle and a mobile robot in realistic situations, with everything computed on board.