Image-based localization in GNSS-denied environments is critical for UAV autonomy. Existing state-of-the-art approaches rely on matching UAV images to geo-referenced satellite images; however, they typically require large-scale, paired UAV-satellite datasets for training. Such data are costly to acquire and often unavailable, limiting their applicability. To address this challenge, we adopt a training paradigm that removes the need for UAV imagery during training by learning directly from satellite-view reference images. This is achieved through a dedicated augmentation strategy that simulates the visual domain shift between satellite and real-world UAV views. We introduce CAEVL, an efficient model designed to exploit this paradigm, and validate it on ViLD, a new and challenging dataset of real-world UAV images that we release to the community. Our method achieves competitive performance compared to approaches trained with paired data, demonstrating its effectiveness and strong generalization capabilities.
We propose a novel method for geolocalizing Unmanned Aerial Vehicles (UAVs) in environments lacking Global Navigation Satellite Systems (GNSS). Current state-of-the-art techniques employ an offline-trained encoder to generate a vector representation (embedding) of the UAV's current view, which is then compared with pre-computed embeddings of geo-referenced images to determine the UAV's position. Here, we demonstrate that the performance of these methods can be significantly enhanced by preprocessing the images to extract their edges, which exhibit robustness to seasonal and illumination variations. Furthermore, we establish that utilizing edges enhances resilience to orientation and altitude inaccuracies. Additionally, we introduce a confidence criterion for localization. Our findings are substantiated through synthetic experiments.
Neural implicit surfaces have become an important technique for multi-view 3D reconstruction but their accuracy remains limited. In this paper, we argue that this comes from the difficulty to learn and render high frequency textures with neural networks. We thus propose to add to the standard neural rendering optimization a direct photo-consistency term across the different views. Intuitively, we optimize the implicit geometry so that it warps views on each other in a consistent way. We demonstrate that two elements are key to the success of such an approach: (i) warping entire patches, using the predicted occupancy and normals of the 3D points along each ray, and measuring their similarity with a robust structural similarity (SSIM); (ii) handling visibility and occlusion in such a way that incorrect warps are not given too much importance while encouraging a reconstruction as complete as possible. We evaluate our approach, dubbed NeuralWarp, on the standard DTU and EPFL benchmarks and show it outperforms state of the art unsupervised implicit surfaces reconstructions by over 20% on both datasets. Our code is available at https://github.com/fdarmon/NeuralWarp
Deep multi-view stereo (MVS) methods have been developed and extensively compared on simple datasets, where they now outperform classical approaches.In this paper, we ask whether the conclusions reached in controlled scenarios are still valid when working with Internet photo collections.We propose a methodology for evaluation and explore the influence of three aspects of deep MVS methods: network architecture, training data, and supervision.We make several key observations, which we extensively validate quantitatively and qualitatively, both for depth prediction and complete 3D reconstructions.First, complex unsupervised approaches cannot train on data in the wild.Our new approach makes it possible with three key elements: upsampling the output, softmin based aggregation and a single reconstruction loss.Second, supervised deep depthmap-based MVS methods are state-of-the art for reconstruction of few internet images.Finally, our evaluation provides very different results than usual ones.This shows that evaluation in uncontrolled scenarios is important for new architectures.
Infrared systems are key to providing enhanced capability to military forces such as automatic control of threats and prevention from air, naval and ground attacks. Key requirements for such a system to produce operational benefits are real-time processing as well as high efficiency in terms of detection and false alarm rate. These are serious issues since the system must deal with a large number of objects and categories to be recognized (small vehicles, armored vehicles, planes, buildings, etc.). Statistical learning based algorithms are promising candidates to meet these requirements when using selected discriminant features and real-time implementation. This paper proposes a new decision architecture benefiting from recent advances in machine learning by using an effective method for level set estimation. While building decision function, the proposed approach performs variable selection based on a discriminative criterion. Moreover, the use of level set makes it possible to manage rejection of unknown or ambiguous objects thus preserving the false alarm rate. Experimental evidences reported on real world infrared images demonstrate the validity of our approach.
In this paper we present a novel approach for mimicking expre ssions in 3D from a monocular video sequence. To this end, first we construct a high res olution semantic mesh model through automatic global and local registration of a low resolution ra ge data. The MPEG-4 standard and radial basis functions are then considered to represent and an imate such a model using a predefined set of control points in a compact fashion. In order to recover th 2D positions of the 3D control points in the observed sequence, we use local cascade Adaboost-dri ven search constrained. The search space is reduced through the use of predictive expression mo deling. The optimal configuration of the Adaboost responses is determined using combinatorial l inear programming which enforces the anthropometric nature of the model defined from these points . 3D position is then deduced and the displacement can be reproduced on any version of the mode l, registered on another face. Our method doesn’t require dense stereo estimation and can then produce realistic animations, using any 3D model. Key-words: Face Reconstruction, Face Animation, Facial Feature Extra ction ∗ Ecole Centrale de Paris † Orange / France Telecom R&D Mime d’Expression : de la Séquence Monoculaire à l’Animation 3D Résumé :Nous présentons dans ce papier une nouvelle approche pour mi me les expressions en 3D à partir d’une séquence monoculaire. Pour cela, nous constru isons un modèle de visage sémantique de haute résolutions grâce à un recalage automatique global et local à partir de données basse résolution. Nous considérons le standard MPEG-4 et les fonctions à base r adiale pour représenter et animer le modèle en utilisant un ensemble de points de contrôle prédéfi nis. Pour retrouver la position 2D de ces points de contrôle 3D dans la séquence observée, nous uti lisons un l’aglorithme de classification Adaboost en cascade. L’espace de recherche est réduit grâce à un modèle de prédiction d’expression. La configuration optimale des réponses de Adaboost est déter minée en utilisant de la programmation linéaire combinatoire, qui contraint la nature anthropome trique du modèle. La position 3D est alors déduite et le déplacement peut être reproduit sur toute vers ion du modèle, recalé sur un autre visage. Notre méthode ne requière pas d’estimation stereo dense et p eut reproduire des animations réalistes, en utilisant n’importe quel modèle 3D. Mots-clés : Reconstruction de Visage, Animation de visage, Extraction de points d’intêret Expression Mimicking 3
This paper presents a new approach for automatic image color correction, based on statistical learning. The method both parameterizes color independently of illumination and corrects color for changes of illumination. The motivation for using a learning approach is to deal with changes of lighting typical of indoor environments such as home and office. The method is based on learning color invariants using a modified multi-layer perceptron (MLP). The MLP is odd-layered. The middle layer includes two neurons which estimate two color invariants and one input neuron which takes in the luminance desired in output of the MLP. The advantage of the modified MLP over a classical MLP is better performance and the estimation of invariants to illumination. The trained modified MLP can be applied using look-up tables (LUTs), yielding very fast processing. Results illustrate the approach.
Reproducing high quality facial expressions is an important challenge in human-computer interaction. Laser-scanners offer an expensive solution to such a problem with image based alternatives being a low-resolution alternative. In this paper, we propose a new method for stereo reconstruction from multiple video pairs that is capable of producing high resolution facial models. To this end, a combinatorial optimization approach is considered and is coupled in time to produce high resolution depth maps. Such optimization is addressed with the use of graph-cuts leading to precise reconstruction of facial expressions that can then be used for animation.
This paper presents a new statistical approach for learning automatic color image correction. The goal is to parameterize color independently of illumination and to correct color for changes of illumination. This is useful in many image processing applications, such as color image segmentation or background subtraction. The motivation for using a learning approach is to deal with changes of lighting typical of indoor environments such as home and office. The method is based on learning color invariants using a modified multi-layer perceptron (MLP). The MLP is odd-layered and the central bottleneck layer includes two neurons that estimates the color invariants and one input neuron proportional to the luminance desired in output of the MLP(luminance being strongly correlated with illumination). The advantage of the modified MLP over a classical MLP is better performance and the estimation of invariants to illumination. Results compare the approach with other color correction approaches from the literature.
Reproducing realistic facial expressions is an important challenge in human computer interaction. In this paper we propose a novel method of modeling and recovering the transitions between different expressions through the use of an autoregressive process. In order to account for computational complexity, we adopt a compact face representation inspired from MPEG-4 standards while in terms of expressions a well known Facial Action Unit System (FACS) comprising the six dominant ones is considered. Then, transitions between expressions are modeled through a time series according to a linear model. Explicit constraints driven from face anthropometry and points interaction are inherited in this model and minimize the risk of producing non-realistic configurations. Towards optimal animation performance, a particular hardware architecture is used to provide the 3D depth information of the corresponding facial elements during the learning stage and the Random Sampling Consensus algorithm for the robust estimation of the model parameters. Promising experimental results demonstrate the potential of such an approach.
In addition to being invariant to image rotation and translation, histograms have the advantage of being easy to compute. These advantages make histograms very popular in computer vision. However, without data quantization to reduce size, histograms are generally not suitable for realtime applications. Moreover, they are sensitive to quantization errors and lack any spatial information. This paper presents a way to keep the advantages of histograms avoiding their inherent drawbacks using local kernel histograms. This approach is tested for background subtraction using indoor and outdoor sequences.
In many applications, like surveillance, image sequences are of poor quality. Motion blur in particular introduces significant image degradation. An interesting challenge is to merge these many images into one high-quality, estimated still. We propose a method to achieve this. Firstly, an object of interest is tracked through the sequence using region based matching. Secondly, degradation of images is modelled in terms of pixel sampling, defocus blur and motion blur. Motion blur direction and magnitude are estimated from tracked displacements. Finally, a high-resolution deblurred image is reconstructed. The approach is illustrated with video sequences of moving people and blurred script.
Although large displays could allow several users to work together and to move freely in a room, their associated interfaces are limited to contact devices that must generally be shared. This paper describes a novel interface called SHIVA (Several-Humans Interface with Vision and Audio) allowing several users to interact remotely with a very large display using both speech and gesture. The head and both hands of two users are tracked in real time by a stereo vision based system. From the body parts position, the direction pointed by each user is computed and selection gestures done with the second hand are recognized. Pointing gesture is fused with n-best results from speech recognition taking into account the application context. The system is tested on a chess game with two users playing on a very large display.
This article presents two new approaches, one parametric and one non-parametric, to the linear grouping of image features. They are based on the Bayesian Hough Transform, which takes into account feature uncertainty. Our main contribution are two new ways to detect the most significant modes of the Hough Transform. Traditionally, this is done by non-maximum suppression. However, in truth, Hough bins measure the likelihoods not of single lines but of collection of lines. Therefore finding lines by non-maxima suppression is not appropriate. This article presents two alternatives. The first method uses bin integration, automatic pruning and fusion to perform mode detection. The second approach detects dominant modes using variable bandwidth mean shift. The advantages of these algorithms are that: (1) the uncertainties associated with feature measurements are taken into account during voting and mode estimation (2) dominant modes are detected in ways that are more correct and less sensitive to errors and biases than non-maxima suppression. The methods can be used with any feature type and any associated feature detection algorithm provided that it outputs a feature position, orientation and covariance matrices. Results illustrate the approaches.
PURPOSETo discuss the technical details of a head mounted display with an augmented reality (AR) system and to describe a first pre-clinical evaluation in interventional MRI.METHODThe AR system consists of a video-see-through head mounted display (HMD), mounted with a mini video camera for tracking and a stereo pair of mini cameras that capture live images of the scene. The live video view of the phantom/patient is augmented with graphical representations of anatomical structures from MRI image data and is displayed on the HMD. The application of the AR system with interventional MRI was tested using a MRI data set of the head and a head phantom.RESULTSThe HMD enables the user to move around and observe the scene dynamically from various viewpoints. Within a short time the natural hand-eye coordination can easily be adapted to the slightly different view. The 3D perception is based on stereo and kinetic depth cues. A circular target with a diameter of 0.5 square centimeter was hit in 19 of 20 attempts. In a first evaluation the MRI image data augmented reality scene of a head phantom allowed good planning and precise simulation of a puncture.CONCLUSIONThe HMD in combination with AR provides a direct, intuitive guidance for interventional MR procedures.
Olivier Bernier合作论文数France Telecom R&D, Lannion Cedex, France4