Object tracking from Unmanned Aerial Vehicles (UAVs) is challenged by platform dynamics, camera motion, and limited onboard resources. Existing visual trackers either lack robustness in complex scenarios or are too computationally demanding for real-time embedded use. We propose an Modular Asynchronous Tracking Architecture (MATA) that combines a transformer-based tracker with an Extended Kalman Filter, integrating ego-motion compensation from sparse optical flow and an object trajectory model. We further introduce a hardware-independent, embedded oriented evaluation protocol and a new metric called Normalized time to Failure (NT2F) to quantify how long a tracker can sustain a tracking sequence without external help. Experiments on UAV benchmarks, including an augmented UAV123 dataset with synthetic occlusions, show consistent improvements in Success and NT2F metrics across multiple tracking processing frequency. A ROS 2 implementation on a Nvidia Jetson AGX Orin confirms that the evaluation protocol more closely matches real-time performance on embedded systems.
Visual odometry is the technique of determining a robot’s pose by analyzing images of its surroundings as it moves. Visual odometry can be categorized into monocular when using a single camera, or stereo when using two cameras or more. In this study, we investigate the use of light-field camera for visual odometry. Capitalizing on the distinctive capability of a light-field camera to record both the intensity and the direction of light, we propose an indirect visual odometry method able to estimate the scale of the translation similarly to stereo visual odometry, but using a single camera sensor. Our visual odometry framework combines light-field imaging with conventional odometry techniques to track the camera movements, using the depth insights provided by a light-field depth estimation approach. Additionally, this method differs from state-of-the-art methods by using a simplified calibration process and a new keypoints extraction method, which makes the use of the light-field cameras easier for robotics perception.
Visual object tracking using Unmanned Aerial Vehicles (UAVs) is a critical computer vision task for military applications, including reconnaissance, surveillance, or munition guidance. Visible and thermal (RGB-T) imaging modalities can be employed complementary by tracking algorithms to address the challenging conditions of such applications. Modern AI-based tracking models must be compatible with severe Size, Weight, and Power (SWaP) constraints for on-board processing. Existing RGB-T object tracking algorithms are typically evaluated on civilian datasets where results cannot be directly extrapolated to real-world military applications. This paper addresses these gaps by evaluating a selection of such algorithms from the state-of-the-art under conditions more representative of military environments and by proposing an efficient tracking algorithm compatible with real-time processing on embedded systems.
In this paper, a new deflectometry approach well suited for online inspection of specular reflective materials is proposed. Based on a simple hardware setup combining two linear light sources, a camera and a conveyor, our approach allows to detect, to localize, and to reconstruct in 3-D, surface aspect defects. It is easy to implement and particularly well suited for large objects in an industrial context (production line, for example). When the camera and light sources are fixed, the first step consists in calibrating the camera and estimating the light sources positions in the camera frame. Then, a full scanning of the object can be done, and the defects are detected in all images. In the last step, the defect is reconstructed in 3-D by slope integration. Thanks to a second light source added to the setup, the initial depth required by this kind of method is automatically estimated. Our approach has been tested in real experimental conditions with different plastic reflective parts and compared to the results provided by a metrology machine to validate the calibration and the reconstruction accuracy. The reconstructions of few millimeters defects on 4 different plastic parts show errors below 50 m. This accuracy meets the usual requirement in the industrial context as the 3D industrial metrology machine that we use to validate our method is more complex to handle with specular objects and has an accuracy of around 20 m.
In recent years, autonomous vehicles have become an axis of academic and industrial research. Localizing these vehicles without a GPS signal represents a challenge for researchers, because the other sensors are usually less accurate, less fast and require more computation. Among localization methods, dead reckoning ones do not need prior knowledge as they are easier to implement for real time purposes. However, their biggest flaw is the accumulation of errors over time. In this work, we present an onboard localization method dedicated to autonomous vehicles for short time navigation without GPS. We developed a method with a high rate inertial-visual data fusion module that allows locating the vehicle in real-time. This method has been validated offline and tested online in a path following control loop on an experimental vehicle.
Standard imaging techniques do not get as much information from a scene as light-field imaging. Light-field (LF) cameras can measure the light intensity reflected by an object and, most importantly, the direction of its light rays. This information can be used in different applications, such as depth estimation, in-plane focusing, creating full-focused images, etc. However, standard key-point detectors often employed in computer vision applications cannot be applied directly to plenoptic images due to the nature of raw LF images. This work presents an approach for key-point detection dedicated to plenoptic images. Our method allows using of conventional key-point detector methods. It forces the detection of this key-point in a set of micro-images of the raw LF image. Obtaining this important number of key-points is essential for applications that require finding additional correspondences in the raw space, such as disparity estimation, indirect visual odometry techniques, and others. The approach is set to the test by modifying the Harris key-point detector.
This paper investigates the impact of path tracking error definitions on vehicle lateral control behavior. Several existing error definitions are reviewed, highlighting the need for a assessment of their impacts. To this aim, new formulations of the vehicle lateral model in terms of tracking errors are proposed. These formulations consider the use of the center-of-gravity (COG) as well as the look-ahead distance (LAD) for lateral and/or orientation errors. Furthermore, four formulations are developed and a generic state space representation is derived. This generic model enables comparative analysis within an unified framework. Four robust state-feedback and feed-forward controllers are designed based on these formulations. Simulations are conducted to evaluate the controllers tracking performances. The results provide insights into the effects of different tracking error definitions on lateral control guidance.
The navigation and guidance of autonomous vehicles require precise and accurate localization. However, in some situations the use of dead reckoning localization is unavoidable, despite the fact that it introduces an accumulation of errors known as drift. This paper focuses on a comparison in simulation between three different localization drift models based on position measurement, including a proposed model. A guidance and control scheme is implemented, including the proposed drift model. Finally, results from closed loop simulation are exploited to highlight the discontinuity phenomenon caused by erroneous position measurements.
This work presents how deflectometry can be coupled with a light-field camera to better characterize and quantify the depth of anomalies on specular surfaces. In our previous work,1 we proposed a new scanning scheme for the detection and 3D reconstruction of defects on reflective objects. However, the quality of the reconstruction was strongly dependent on the object-camera distance which was required as an external input parameter. In this paper, we propose a new approach that integrates an estimation of this distance into our system by replacing the standard camera with a light-field camera.
Light-field and plenoptic cameras are widely available today. Compared with monocular cameras, these cameras capture not only the intensity but also the direction of the light rays. Due to this specificity, light-field cameras allow for image refocusing and depth estimation using a single image. However, most of the existing depth estimation methods using light-field cameras require a prior complex calibration phase and raw data preprocessing before the desired algorithm is applied. We propose a homography-based method with plenoptic camera parameters calibration and optimization, dedicated to our homography-based micro-images matching algorithm. The proposed method works on debayerred raw images with vignetting correction. The proposed approach directly links the disparity estimation in the 2D image plane to the depth estimation in the 3D object plane, allowing for direct extraction of the real depth without any intermediate virtual depth estimation phase. Also, calibration parameters used in the depth estimation algorithm are directly estimated, and hence no prior complex calibration is needed. Results are illustrated by performing depth estimation with a focused light-field camera over a large distance range up to 4 m. (C) 2021 Society of Photo-Optical Instrumentation Engineers (SPIE)
Light-Field (LF) cameras allow the extraction not only of the intensity of light but also of the direction of light rays in the scene, hence it records much more information of the scene than a conventional camera. In this paper, we present a novel method to detect key-points in raw LF images by applying key-points detectors on Pseudo-Focused images (PFIs). The main advantage of this method is that we don’t need to use complex key-points detectors dedicated to light-field images. We illustrate the method in two use cases: the extraction of corners in a checkerboard and the key-points matching in two view raw light-field images. These key-points can be used for different applications e.g. calibration, depth estimation or visual odometry. Our experiments showed that our method preserves the accuracy of detection by re-projecting the pixels in the original raw images.
In the 50s, biologists discovered that some electric fish is capable of discriminating the pose as well as the electric and geometric properties of surrounding objects by navigating and measuring the distortions of a self-generated electric field. In this article, we address the challenging issue of ellipsoidal objects pose and size estimation for underwater robots equipped with artificial electric sense. Unlike current methods, the approach can estimate both the position and size in parallel with a single straight trajectory. Neither multipolarization nor reactive self-alignment control are necessary to locate the object. The approach is a purely model-based heuristic that selects the best ellipsoid parameters among a set of potential candidates. It is based on a set of four electric measurements recorded at several positions along the robot trajectory along which the displacement is measured. The efficiency of the method is assessed over numerous experiments with different objects, several positions, and orientations, and two different kinds of water (fresh and salt water). Despite some model simplifications and experimental errors, location and size estimation errors are on average below $\text{1}\,$cm and $\text{15}\%$, respectively, while offering promising perspectives for real-time computation.
Autonomous vehicle navigation requires the desired trajectory and the current localization to be able to calculate the command that must be sent to the actuators. The localization of the vehicle (usually defined by a position vector and an orientation vector), can be provided by external systems. GPS localization is the most accurate solution but when it is no longer available or precise, an on-board localization estimation based on proprioceptive and exteroceptive sensors is needed. Visual odometry is a well-known approach to estimate the vehicle motion from a camera. Unfortunately, visual localization is subject to errors that increase over time (drift). In this paper, we provide a study of the impact of localization errors on the control of an autonomous vehicle. In order to validate visual odometry algorithms in simulation, a drift model is proposed. Real navigation experiments with errors on the localization are presented to characterize the drift model and the propagation of localization errors in the controller module and the associated command signal.
In the past few years, a new type of camera has been emerging on the market: a digital camera capable of capturing both the intensity of the light emanating from a scene and the direction of the light rays. This camera technology called a light-field camera uses an array of lenses placed in front of a single image sensor, or simply, an array of cameras attached together. An optical device is proposed: a four minilens ring that is inserted between the lens and the image sensor of a digital camera. This device prototype is able to convert a regular digital camera into a light-field camera as it makes it possible to record four subaperture images of the scene. It is a compact and cost-effective solution to perform both postcapture refocusing and depth estimation. The minilens ring makes also the plenoptic camera versatile; it is possible to adjust the parameters of the ring so as to reduce or increase the size of the projected image. Together with the proof of concept of this device, we propose a method to estimate the positions of each optical component depending on the observed scene (object size and distance) and the optics parameters. Real-world results are presented to validate our device prototype. (C) 2019 Society of Photo-Optical Instrumentation Engineers (SPIE)
In computer vision, the epipolar geometry embeds the geometrical relationship between two views of a scene. This geometry is degenerated for planar scenes as they do not provide enough constraints to estimate it without ambiguity. Nearly planar scenes can provide the necessary constraints to resolve the ambiguity. But classic estimators such as the 5-point or 8-point algorithm combined with a random sampling strategy are likely to fail in this case because a large part of the scene is planar and it requires lots of trials to get a nondegenerated sample. However, the planar part can be associated with a homographic model and several links exist between the epipolar geometry and homographies. The epipolar geometry can indeed be recovered from at least two homographies or one homgraphy and two noncoplanar points. The latter fits a wider variety of scenes, as it is unsure to be able to find a second homography in the noncoplanar points. This method is called plane-and-parallax. The equivalence between the parallax and the epipolar lines allows to recover the epipole as their common intersection and the epipolar geometry. Robust implementations of the method are rarely given, and we encounter several limitations in our implementation. Noisy image features and outliers make the lines not to be concurrent in a common point. Also off-plane features are unequally influenced by the noise level. We noticed that the bigger the parallax is, the lesser the noise influence is. We, therefore, propose a model for the parallax that takes into account the noise on the features location to cope with the previous limitations. We call our method the "parallax beam." The method is validated on the KITTI vision benchmark and on synthetic scenes with strong planar degeneracy. The results show that the parallax beam improves the estimation of the camera motion in the scene with planar degeneracy and remains usable when there is not any particular planar structure in the scene. (C) 2019 SPIE and IS&T
During the last two decades the number of visual odometry algorithms has grown rapidly. While it is straightforward to obtain a qualitative result, if the shape of the trajectory is in accordance with the movement of the camera, a quantitative evaluation is needed to evaluate the performances and to compare algorithms. In order to do so, one needs to establish a ground truth either for the overall trajectory or for each camera pose. To this end several datasets have been created. We propose a review of the datasets created over the last decade. We compare them in terms of acquisition settings, environment, type of motion and the ground truth they provide. The purpose is to allow researchers to rapidly identifies the datasets that best fit their work. While the datasets cover a variety of techniques to establish a ground truth, we provide also the reader with techniques to create one that were not present among the reviewed datasets.