In this paper, we address common error sources for 3D Gaussian Splatting (3DGS) including blur, imperfect camera poses, and color inconsistencies, with the goal of improving its robustness for practical applications like reconstructions from handheld phone captures. Our main contribution involves modeling motion blur as a Gaussian distribution over camera poses, allowing us to address both camera pose refinement and motion blur correction in a unified way. Additionally, we propose mechanisms for defocus blur compensation and for addressing color in-consistencies caused by ambient light, shadows, or due to camera-related factors like varying white balancing settings. Our proposed solutions integrate in a seamless way with the 3DGS formulation while maintaining its benefits in terms of training efficiency and rendering speed. We experimentally validate our contributions on relevant benchmark datasets including Scannet++ and Deblur-NeRF, obtaining state-of-the-art results and thus consistent improvements over relevant baselines.
Reconstruction 3D par Deep Learning : supervision et représentation La reconstruction 3D est un problème classique en vision par ordinateur. Pourtant, les meilleures méthodes ne fonctionnent toujours pas parfaitement lorsque les images utilisées présentent de grands changements d'illumination et de nombreuses occlusions. L'apprentissage profond (Deep Learning) promet d'améliorer la reconstruction 3D dans de telles configurations, mais les méthodes classiques produisent encore les meilleurs résultats aujourd'hui. Dans cette thèse, nous analysons la spécificité de l'apprentissage profond appliqué à la reconstruction 3D multi-vues et nous introduisons de nouvelles méthodes basées sur l'apprentissage profond.La première contribution de cette thèse est une analyse des différentes supervisions possibles pour l’entraînement de modèles d'apprentissage profond pour l’appariement d'images. Nous introduisons un algorithme en deux étapes qui calcule d'abord des correspondances à basse résolution en utilisant l'apprentissage profond, puis des correspondances de points d'intérêt classiques à l'intérieur des régions appariées. Nous analysons plusieurs niveaux de supervision et montrons que notre nouvelle supervision épipolaire donne les meilleurs résultats.La deuxième contribution est également une étude de la supervision pour l'apprentissage profond mais appliquée à un autre scénario : la reconstruction 3D calibrée à partir d’image non contraintes. Nous montrons que les méthodes non supervisées existantes ne fonctionnent pas sur de telles données et nous introduisons une nouvelle technique d’apprentissage qui résout ce problème. Nous comparons ensuite de manière exhaustive l'approche non supervisée et l'approche supervisée avec différentes architectures de réseau et différentes données d'entraînement.Enfin, notre troisième contribution concerne la représentation des données. Les représentations implicites ont été récemment utilisées pour le rendu d'images. Nous adaptons cette représentation au problème de la reconstruction multi-vues et nous introduisons une nouvelle méthode qui, comme les techniques classiques de reconstruction 3D, optimise la photo-consistance entre les projections de plusieurs images. Notre approche améliore largement les performances de l'état de l'art.
Neural implicit surfaces have become an important technique for multi-view 3D reconstruction but their accuracy remains limited. In this paper, we argue that this comes from the difficulty to learn and render high frequency textures with neural networks. We thus propose to add to the standard neural rendering optimization a direct photo-consistency term across the different views. Intuitively, we optimize the implicit geometry so that it warps views on each other in a consistent way. We demonstrate that two elements are key to the success of such an approach: (i) warping entire patches, using the predicted occupancy and normals of the 3D points along each ray, and measuring their similarity with a robust structural similarity (SSIM); (ii) handling visibility and occlusion in such a way that incorrect warps are not given too much importance while encouraging a reconstruction as complete as possible. We evaluate our approach, dubbed NeuralWarp, on the standard DTU and EPFL benchmarks and show it outperforms state of the art unsupervised implicit surfaces reconstructions by over 20% on both datasets. Our code is available at https://github.com/fdarmon/NeuralWarp
Epipolar rectification of a stereo pair is the process of resampling a pair of stereo images so that the apparent motion of corresponding points is horizontal. This is an important preliminary step in depth estimation, substituting depth by disparity estimation. Most methods rely on a perspective transform of both images, which has the advantage to simulate a different attitude of the pinhole cameras. A limitation is that when an epipole is inside the image domain, it has to be sent to infinity by the perspective transform, producing a strong distortion. On the contrary, relying on a polar transform centered at the epipole provides a method applicable universally to a pair of pinhole camera views. We present in detail the algorithm, filling in the information important for its implementation and missing in published articles.
Deep multi-view stereo (MVS) methods have been developed and extensively compared on simple datasets, where they now outperform classical approaches.In this paper, we ask whether the conclusions reached in controlled scenarios are still valid when working with Internet photo collections.We propose a methodology for evaluation and explore the influence of three aspects of deep MVS methods: network architecture, training data, and supervision.We make several key observations, which we extensively validate quantitatively and qualitatively, both for depth prediction and complete 3D reconstructions.First, complex unsupervised approaches cannot train on data in the wild.Our new approach makes it possible with three key elements: upsampling the output, softmin based aggregation and a single reconstruction loss.Second, supervised deep depthmap-based MVS methods are state-of-the art for reconstruction of few internet images.Finally, our evaluation provides very different results than usual ones.This shows that evaluation in uncontrolled scenarios is important for new architectures.
We tackle the problem of finding accurate and robust keypoint correspondences between images. We propose a learning-based approach to guide local feature matches via a learned approximate image matching. Our approach can boost the results of SIFT to a level similar to state-of-the-art deep descriptors, such as Superpoint, ContextDesc, or D2-Net and can improve performance for these descriptors. We introduce and study different levels of supervision to learn coarse correspondences. In particular, we show that weak supervision from epipolar geometry leads to performances higher than the stronger but more biased point level supervision and is a clear improvement over weak image level supervision. We demonstrate the benefits of our approach in a variety of conditions by evaluating our guided keypoint correspondences for localization of internet images on the YFCC100M dataset and indoor images on theSUN3D dataset, for robust localization on the Aachen day-night benchmark and for 3D reconstruction in challenging conditions using the LTLL historical image data.
This paper considers the generic problem of dense alignment between two images, whether they be two frames of a video, two widely different views of a scene, two paintings depicting similar content, etc. Whereas each such task is typically addressed with a domain-specific solution, we show that a simple unsupervised approach performs surprisingly well across a range of tasks. Our main insight is that parametric and non-parametric alignment methods have complementary strengths. We propose a two-stage process: first, a feature-based parametric coarse alignment using one or more homographies, followed by non-parametric fine pixel-wise alignment. Coarse alignment is performed using RANSAC on off-the-shelf deep features. Fine alignment is learned in an unsupervised way by a deep network which optimizes a standard structural similarity metric (SSIM) between the two images, plus cycle-consistency. Despite its simplicity, our method shows competitive results on a range of tasks and datasets, including unsupervised optical flow on KITTI, dense correspondences on Hpatches, two-view geometry estimation on YFCC100M, localization on Aachen Day-Night, and, for the first time, fine alignment of artworks on the Brughel dataset. Our code and data are available at http://imagine.enpc.fr/~shenx/RANSAC-Flow/
BACKGROUND:We describe the epidemiological, clinical, and prognostic aspects of 177 tularemia cases diagnosed at the National Reference Center for rickettsioses, coxiellosis, and bartonelloses between 2008 and 2017. METHODS:All patients with a microbiological diagnosis of tularemia made in the laboratory were included. Clinical and epidemiological data were collected retrospectively from clinicians in charge of patients using a standardized questionnaire. Diagnostic methods used were indirect immunofluorescence serology, real-time polymerase chain reaction (PCR), and universal PCR targeting the 16S ribosomal ribonucleic acid gene. RESULTS:The series included 54 females and 123 males (sex ratio, 2.28; mean age, 47.38 years). Eighty-nine (50.2%) were confirmed as having tularemia on the basis of a positive Francisella tularensis PCR or seroconversion, and 88 (49.8%) were considered as probable due to a single positive serum. The regions of France that were most affected included Pays de la Loire (22% of cases), Nouvelle Aquitaine (18.6% of cases), and Grand Est (12.4% of cases). Patients became infected mainly through contact with rodents or game (38 cases, 21.4%), through tick-bites (23 cases, 12.9%), or during outdoor leisure activities (37 cases, 20.9%). Glandular and ulceroglandular forms were the most frequent (109 cases, 61.5%). Two aortitis, an infectious endocarditis, a myocarditis, an osteoarticular infection, and a splenic hematoma were also diagnosed. Tularemia was discovered incidentally in 54.8% of cases. Seventy-eight patients were hospitalized, and no deaths were reported. CONCLUSIONS:Our data suggest that in an endemic area and/or in certain epidemiological contexts, tularemia should be sought to allow an optimized antibiotic therapy and a faster recovery.