This paper introduces a depth-independent Image-Based Visual Servoing (IBVS) framework, replacing the traditional inverse-Jacobian approach with a formulation based on homography-derived scaling factors. By exploiting the geometric relationships between corresponding image features, the method guides the control without requiring explicit depth measurements or 3D scene reconstruction. We provide a local stability analysis for the proposed Depth-Independent IBVS (DI-IBVS) and extend its applicability to non-planar scenes through a multi-homography approach. Experimental results on a real robot demonstrate that, when depth measurements are unavailable, DI-IBVS consistently outperforms classical IBVS and 2.5D methods, especially in non-planar environments where the latter are prone to depth estimation errors and homography decomposition failures.
This paper addresses vision based autonomous vine pruning with robots. The complex structure of vines makes visual servoing difficult due to challenges in feature extraction and 3D pose estimation. A novel approach interweaving visual servoing and vision based planning is proposed. An online planner features a Nonlinear Model Predictive Control to compute on-the-fly suitable 3D waypoints. These are navigated by a Position Based Visual Servoing integrating Iterative Closest Point based point-cloud alignment. Qualitative and quantitative evaluations are conducted on live experiments where a Franka Emika manipulator embedding a eye-in-hand stereo camera performs a sequence of pruning actions on seven increasingly complex vine stocks, in spite of unexpected motion of the vine and/or the robot's base. Outdoor experiments are also carried out in strong wind conditions.
Sound Source Localization algorithms estimate the Direction of Arrival of one or multiple (moving) sound sources in a 3-D space. Potential applications include environment acoustic mapping, spatial filtering of relevant sources out of acoustic clutter and noise, and headset-free human-robot speech interaction. Deep Learning algorithms have been widely used, with many solutions relying on Feed-forward Neural Networks, convolutional kernels, or attention mechanisms derived from the Transformer. In this paper, the hardware implementation of a Deep Learning model for Sound Source Localization is proposed on FPGA evaluation board (Digilent Zybo-7020). This enables Artificial Intelligence inference directly onto a FPGA with acceptable performances (70 % accuracy for an error < 10 degrees) through an energy-efficient full hardware implementation, without resorting to usual processing units (e.g., CPU, GPU, DSP, etc.) nor soft-core processor with hardwired accelerators/co-processors.
This paper addresses the estimation of trajectories of interacting vehicles at a microscopic scale, as a prerequisite to their prediction for risk assessment. A state-space solution is investigated, where both the Markov hidden state (continuous-valued, which captures the joint histories of vehicles) and the measurements (low-dimensional and noisy) admit a vehicle-wise structure. The vehicles' transition models are assumed independent of each other, time- and vehicle-invariant, and coequal to an “egocentric” prior dynamics pdf. To cope with the vehicles' interactions, this pdf is conditioned on the full state vector as the past time index, which imposes a centralized estimation/prediction of the fleet motion. The two fundamental pillars of the approach are developed: learning of a Gaussian mixture egocentric transition model by means of Deep Neural Networks; synthesis of a stochastic variational Bayes filtering algorithm which features a decentralized vehicle-wise structure but takes into account interactions. Tests on highway scenarios are presented.
In robotics, "person following" depicts the servoing of the relative situation of a robot w.r.t. a moving person. This property may be hard to achieve, especially when the estimation of the person ego-motion is weak (e.g., due to limited prior knowledge or computational resources). This paper introduces a nonholomic mobile robot controller, which ensures an intuitive and safe behavior through an insightful robot-centered problem statement. Under realistic bounded-error readings of hidden constant person velocities, ultimate boundedness of the state vector norm can be ensured in the neighborhood of its equilibrium.
Audio-motor binaural localization algorithms, which combine directional cues extracted from the sensed signals with the motion of the sensor, are known to overcome shortcomings such as font-back ambiguity and source range non-observability. They can be improved by closing the loop from their output to the control inputs of the sensor, i.e., the sensor motor commands. This paper presents an approach, coined “information-based feedback control”, which drives in real time a binaural head so as to gather information on the location of a static source. On the one hand, a “greedy” approach moves the head to its next best position. On the other hand, a multi-stepahead scheme determines its most effective path over a receding horizon of size N, by reasoning on average over yet uncollected audio data. Both methods internally entail the prediction of binaural cues, e.g., ITDs, ILDs or a combination of both. Some results can be given an elegant commonsense interpretation.
Robots are usually equipped with advanced capabilities in order to autonomously adapt to real and dynamic environments and to interact with humans. Robot Perception is being inspired by new embodied cognition approaches that redefine the notions of perception, cognition and action, basic processes of intelligent behaviour. Enactive approaches consider the perceptual act as a consequence in action of the structural coupling between the organism and its environment in seek of significance. There is a growing interest in the development of robot perception systems based on new architectures to materialize naturally these action-perception functions. In this direction, we propose an evolution of the EAR sensor [4], which fulfills the constraints of mobile robot audition, such as embeddability, synchronous multichannel acquisition and real-time execution. Using Systems-on-a-Programmable-Chip (SoPCs) methodology, this sensor incorporates a novel architecture that offers all the basic calculation blocks necessary to perform most binaural and array auditory functions, and allows to easily develop new functionalities and connections between motor and perceptual modules in order to implement enactive behaviour. Moreover, Microelectromechanical Systems (MEMS) microphones have been studied and implemented, enabling the acquisition of high-fidelity audio on inexpensive and portable devices. In this paper, performance results are presented for sound source detection and localization functions, and progress is shown towards the implementation of MEMS microphones in Human-Machine Interfaces and Robot Audition. Finally, evolutions towards an interdisciplinary design of enactive audition functions are discussed.
An Augmented Reality prototype is presented. Its hardware architecture is composed of a Head Mounted Display, a wide Field of View (FOV) stereo-vision passive system, a gaze tracker and a laptop. An associated software architecture is proposed to immerse the user in augmented environments where he/she can move freely. The system maps the unknown real-world (indoor or outdoor) environment and is localized into this map by means of binocular state-of-the-art Simultaneous Localization and Mapping techniques. It overcomes the FOV limitations of conventional augmented reality devices by using wide-angle cameras and associated algorithms. It also solves the parallax issue induced by the distinct locations of the two cameras and of the user's eyes by using Depth Image Based Rendering. An embedded gaze tracker, together with environment modeling techniques, enable gaze controlled interaction. A simple application is presented, in which a virtual object is inserted into the user's FOV and follows his/her gaze. While the targeted real time performance has not yet been achieved, the paper discusses ways to improve both frame rate and latency. Other future works are also overviewed.
A suboptimal algorithm to fixed-interval and fixed-lag smoothing for Markovian switching systems is proposed. It infers a Gaussian mixture approximation of the smoothing pdf by combining the statistics produced by an IMM filter into an original backward recursive process. The number of filters and smoothers is equal to the constant number of hypotheses in the posterior mixture. A comparison, conducted on simulated case studies, shows that the investigated method performs significantly better than equivalent algorithms.
In static scenarios, binaural sound localization is fundamentally limited by front-back ambiguity and distance non-observability. Over the past few years, “active” schemes have been shown to overcome these shortcomings, by combining spatial binaural cues with the motor commands of the sensor. In this context, given a Gaussian prior on the relative position to a source, this paper determines an admissible motion of a binaural head which leads, on average, to the one-step-ahead most informative audio-motor localization. To this aim, a constrained optimization problem is set up, which consists in maximizing the entropy of the next predicted measurement probability density function over a cylindric admissible set. The method is appraised through geometrical arguments, and validated in simulations and on real-life robotic experiments.
We present a database of binaural room impulse responses (BRIRs) measured in an apartment-like environment. The BRIRs were captured at four different sound source positions, each combined with four listener positions. A head and torso simulator (HATS) with varying head-orientation in the range of ±78 • with 2 • resolution was used. Additionally, BRIRs of 20 listener positions along a trajectory connecting two of the four positions were measured, each with a fixed head-orientation. The data is provided in the Spatially Oriented Format for Acoustics (SOFA) and it is freely available under the Creative Commons (CC-BY-4.0) license. It can be used to simulate complex acoustic scenes in order to study the process of auditory scene analysis for humans and machines.
Robots are usually equipped with advanced capabilities in order to autonomously adapt to real and dynamic environments and to interact with humans. Robot Perception is being inspired by new embodied cognition approaches that redefine the notions of perception, cognition and action, basic processes of intelligent behaviour. Enactive approaches consider the perceptual act as a consequence in action of the structural coupling between the organism and its environment in seek of significance. There is a growing interest in the development of robot perception systems based on new architectures to materialize naturally these action-perception functions. In this direction, we propose an evolution of the EAR sensor [4], which fulfills the constraints of mobile robot audition, such as embeddability, synchronous multichannel acquisition and real-time execution. Using Systems-on-a-Programmable-Chip (SoPCs) methodology, this sensor incorporates a novel architecture that offers all the basic calculation blocks necessary to perform most binaural and array auditory functions, and allows to easily develop new functionalities and connections between motor and perceptual modules in order to implement enactive behaviour. Moreover, Microelectromechanical Systems (MEMS) microphones have been studied and implemented, enabling the acquisition of high-fidelity audio on inexpensive and portable devices. In this paper, performance results are presented for sound source detection and localization functions, and progress is shown towards the implementation of MEMS microphones in Human-Machine Interfaces and Robot Audition. Finally, evolutions towards an interdisciplinary design of enactive audition functions are discussed.
En la actualidad existe un gran interes cientifico y tecnologico por el desarrollo de sistemas de percepcion robotica con arquitecturas que permitan materializar de manera mas “natural” -esto es, tal como un organismo se desempena en su propio medio- funciones percepcion de bajo nivel. Estas arquitecturas deben poseer una gran versatilidad y capacidad de computo que satisfaga los requerimientos de autonomia energetica y dimensiones para ser implementado en un sistema robotico movil. En particular, ciertos algoritmos de audicion robotica requieren la ejecucion secuencial de un gran numero de operaciones matriciales a variable compleja, sobre volumenes importantes de datos y en tiempo real. En este trabajo se presenta una metodologia de diseno para la implementacion de este tipo de funciones utilizando FPGA (del ingles Field Programmable Gate Array) a partir de la integracion de un procesador RISC (Reduced Instruction Set Computer) estandar y modulos de hardware especificamente disenados para la aceleracion de calculos trigonometricos y a variable compleja. De esta manera, se pueden implementar facilmente algoritmos mas complejos y se reduce el tiempo de desarrollo. En particular, se presentan resultados de mejora de la performance de la descomposicion propia generalizada de matrices hermitianas a traves de los metodos Jacobi y CORDIC (COordinate Rotation DIgital Computer).
Fundamental limitations of binaural localization, such as front-back ambiguity or distance non-observability, can be overcome by combining the sensed audio signals with the sensor motor commands into "active" schemes. Such strategies can rely on stochastic filtering. In this context, this paper addresses the determination of an admissible motion of a binaural head leading, on average, to the one-step-ahead most informative localization. To this aim, a constrained optimization problem is set up, which consists in maximizing the entropy of the next predicted measurement probability density function over a cylindric admissible set. The proposed optimum policy is validated on real-life robotic experiments.
This paper takes place within the field of sound source localization by combining the signals sensed by a binaural head with its motor commands. Such so-called "active" schemes are known to overcome limitations occurring in the static context, such as front-back ambiguities or distance non-observability. On the basis of a stochastic filter, which approximates the posterior probability density function of the sensor-to-source situation, a feedback controller of the sensor motion is proposed so as to reduce the associated uncertainty. An information-theoretic analysis of the effect of the sensor motion on the localization uncertainty is first conducted. Then, a gradient ascent scheme is used to drive the head towards the area of minimum uncertainty (maximum information). An evaluation on simulated scenarios, as well as on data coming from real experiments, is included.
This paper takes place within the field of binaural localization in robotics. The aim is to design “active” schemes, which combine the signals sensed by a binaural head with its motor commands so as to overcome limitations occurring in a static context: front-back confusion, non-observability of hidden variables, etc. A three-stage strategy is proposed, which entails: the short-term detection and localization of sources from the short-term analysis of the binaural stream; the assimilation of these data over time and the fusion with the motor commands of the binaural sensor; the improvement of this fusion through the feedback control of the binaural sensor. For each stage, the theoretical bases, some achievements and open problems are outlined.
Argos is a dedicated system for geo-localization and data collection of platform terminal transmitters (PTTs). The system exploits a constellation of polar-orbiting satellites recording the messages transmitted by the PTTs. The localization processing takes advantage of the Doppler effect on the carrier frequency of messages received by the satellites to estimate platform locations. It was recently demonstrated that the use of an Interacting Multiple Model (IMM) filter significantly increases the Argos location accuracy compared to the simple Least Square adjustment technique that had been used from the beginning of the Argos localization service in 1978. The accuracy gain is especially large in cases when the localization is performed from a small number of messages (n ≤ 3). The present paper shows how it is possible to further improve the Argos location accuracy if a processing delay is accepted. The improvement is obtained using a fixed-interval multiple-model smoothing technique.
This paper attempts to provide a state-of-the-art of sound source localization in robotics. Noticeably, this context raises original constraints—e.g. embeddability, real time, broadband environments, noise and reverberation—which are seldom simultaneously taken into account in acoustics or signal processing. A comprehensive review is proposed of recent robotics achievements, be they binaural or rooted in array processing techniques. The connections are highlighted with the underlying theory as well as with elements of physiology and neurology of human hearing.
This paper attempts to provide a state-of-the-art of sound source localization in robotics. Noticeably, this context raises original constraints—e.g. embeddability, real time, broadband environments, noise and reverberation—which are seldom simultaneously taken into account in acoustics or signal processing. A comprehensive review is proposed of recent robotics achievements, be they binaural or rooted in array processing techniques. The connections are highlighted with the underlying theory as well as with elements of physiology and neurology of human hearing.
An unknown intermittent sound source, at rest or in motion, is considered. We describe an original method to detect its activity and determine its azimuth and range from a moving binaural sensor subject to measurement noise. The proposed source localization scheme is said active, as it combines the sensed binaural signals with the sensor motion. It is composed of two stages: first, a pseudo-likelihood of the source azimuth is defined from the short-term time-frequency analysis of the binaural signals, and the source activity is detected by means of statistical identification; then, this information brought by the binaural signals is assimilated over time and fused with the sensor motor commands into a stochastic filtering strategy entailing a bank of noninteractive unscented Kalman filters. The method enjoys several important features. On the one hand, it is endowed with self-initialization and ensures the consistency of the covariances of the estimation errors. On the other hand, it can rigorously handle the effects induced by scatterers by explicitly exploiting the HRTFs to the microphones. A validation on simulated scenarios as well as on data coming from real experiments is included. This work has been partially supported by the EU FP7 FET-Open TWO!EARS project (2014-2016), whose goal is to develop a computational model of auditory perception and experience, in which binaural bottom up processing is interwoven with action and with top-down feedbacks originating from cortical levels. It takes place at the sensorimotor (reflex) level of this architecture. Current work concerns its extension to the multiple source case and to the definition of active motions which can improve perception.