Image-based 3D reconstruction is a powerful method for accurately reconstructing an object’s geometry and texture from images. A crucial factor for the accuracy and completeness of the resulting reconstructed model is the choice of poses for capturing images, which is called view planning. One possible view planning strategy uses an iterative feedback loop that switches between planning poses and an incremental reconstruction to autonomously digitize an object without prior knowledge. However, this approach requires identifying which parts of an object are “poorly reconstructed” and thus would benefit from being part of additional images. This work explores the use of point cloud quality metrics to provide this feedback by comprehensively comparing a set of existing and newly introduced metrics in terms of their time-dependent behavior, similarity, and their applicability to view planning. Among the newly proposed metrics this work introduces the Reconstruction Quality Feedback (RQF), which shows a significantly improved performance in simulations when being used for view planning. The effectiveness of RQF is also demonstrated for real objects on an autonomous robotic 3D digitization system.
Estimating the 6D pose of objects is a critical challenge for robotics and augmented reality applications. The problem is aggravated by the fact that critical attributes, such as an object's texture and material, as well as the specific lighting conditions under which it must be identified, are often unknown. Neural Radiance Fields (NeRFs) and 3D Gaussian splatting (3DGS) are techniques that enable high-quality reconstruction of real-world scenes. By revising the scene fitting function, these representations can facilitate the estimation of an object's pose within a given environment. However, a major complication is that the unique textures, materials, and lighting conditions are fixed within the scene, which can impair the accuracy of pose estimation. To address this, we adopt two alterations to the standard NeRF framework that enhance its ability to handle greatly varied object appearances such as material and texture. Our modified approaches are evaluated on the prevalent YCB-V object dataset, demonstrating their effectiveness. Our two proposed algorithms achieve mesh-free 6D Object Pose Estimation for objects with previously unseen appearances, requiring only a collection of input images to train the NeRF model.
Semantic Image Segmentation facilitates a multitude of real-world applications ranging from autonomous driving over industrial process supervision to vision aids for human beings. These models are usually trained in a supervised fashion using example inputs. Distribution Shifts between these examples and the inputs in operation may cause erroneous segmentations. The robustness of semantic segmentation models against distribution shifts caused by differing camera or lighting setups, lens distortions, adversarial inputs and image corruptions has been topic of recent research. However, robustness against spatially varying radial distortion effects that can be caused by uneven glass structures (e.g. windows) or the chaotic refraction in heated air has not been addressed by the research community yet. We propose a method to synthetically augment existing datasets with spatially varying distortions. Our experiments show, that these distortion effects degrade the performance of state-of-the-art segmentation models. Pretraining and enlarged model capacities proof to be suitable strategies for mitigating performance degradation to some degree, while fine-tuning on distorted images only leads to marginal performance improvements.
Computer vision techniques are on the rise for industrial applications, like process supervision and autonomous agents, e.g., in the healthcare domain and dangerous environments. While the general usability of these techniques is high, there are still challenging real-world use-cases. Especially transparent structures, which can appear in the form of glass doors, protective casings or everyday objects like glasses, pose a challenge for computer vision methods. This paper evaluates the combination of transparent objects in conjunction with (naturally occurring) contamination through environmental effects like hazing. We introduce a novel publicly available dataset containing 489 images incorporating three grades of water droplet contamination on transparent structures and examine the resulting influence on transparency handling. Our findings show, that contaminated transparent objects are easier to segment and that we are able to distinguish between different severity levels of contamination with a current state-of-the art machine-learning model. This in turn opens up the possibility to enhance computer vision systems regarding resilience against, e.g., datashifts through contaminated protection casings or implement an automated cleaning alert.
Interactive deformation via control handles is essential in computer graphics for the modeling of 3D geometry. Deformation control structures include lattices for free-form deformation and skeletons for character articulation, but this report focuses on cage-based deformation. Cages for deformation control are coarse polygonal meshes that encase the to-be-deformed geometry, enabling high-resolution deformation. Cage-based deformation enables users to quickly manipulate 3D geometry by deforming the cage. Due to their utility, cage-based deformation techniques increasingly appear in many geometry modeling applications. For this reason, the computer graphics community has invested a great deal of effort in the past decade and beyond into improving automatic cage generation and cage-based deformation. Recent advances have significantly extended the practical capabilities of cage-based deformation methods. As a result, there is a large body of research on cage-based deformation. In this report, we provide a comprehensive overview of the current state of the art in cage-based deformation of 3D geometry. We discuss current methods in terms of deformation quality, practicality, and precomputation demands. In addition, we highlight potential future research directions that overcome current issues and extend the set of practical applications. In conjunction with this survey, we publish an application to unify the most relevant deformation methods. Our report is intended for computer graphics researchers, developers of interactive geometry modeling applications, and 3D modeling and character animation artists.
Image-based 3D reconstruction is a commonly used technique for measuring the geometry and color of objects or scenes based on images. While the geometry reconstruction of state-of-the-art approaches is mostly robust against varying lighting conditions and outliers, these pose a significant challenge for calculating an accurate texture map. This work proposes a deep-learning based texturing approach called "DeepTex" that uses a custom learned blending method on top of a traditional mosaic-based texturing approach. The model was trained using a custom synthetic data generation workflow and showed a significantly increased accuracy when generating textures in the presence of outliers and non-uniform lighting.
Real-time delivery of captured optical surface materials has several uses, ranging from commercial showcasing of products, or product configurators, on the web, to dissemination of and research on cultural heritage by providing easy access to collections. In this work we describe two contributions: First, the development of a real-time 3D viewer for captured optical surface materials mapped to 3D geometry on the web, which is demonstrated using measured Approximate Bidirectional Texturing Function (ABTF) surface materials as an example. Second, we solve the general problem of the initial wait time for the user when viewing 3D renderings coupled with large measured material data by progressively enhancing the material details presented to the user from the start, following one of several methods proposed to define the order in which the individual surface fragments are sent. This reordering prevents the retransmission of fragments and continuously enhances the perceived increase in quality over time, making each step noticeably better in terms of perceived quality without producing visible pop-in effects or other artifacts. The benefit is the ability to visualize captured optical surface materials in their full detail without need for lossy compression, while retaining interactive loading times and thus improving user experience.
Neural radiance fields (NeRFs) have revolutionized novel view synthesis, leading to an unprecedented level of realism in rendered images. However, the reconstruction quality of NeRFs suffers significantly from out-of-focus regions in the input images. We propose NeRF-FF, a plug-in strategy that estimates image masks based on Focus Frustums (FFs), i.e., the visible volume in the scene space that is in-focus. NeRF-FF enables a subsequently trained NeRF model to omit out-of-focus image regions during the training process. Existing methods to mitigate the effects of defocus blurred input images often leverage dynamic ray generation. This makes them incompatible with the static ray assumptions employed by runtime-performance-optimized NeRF variants, such as Instant-NGP, leading to high training times. Our experiments show that NeRF-FF outperforms state-of-the-art approaches regarding training time by two orders of magnitude—reducing it to under 1 min on end-consumer hardware—while maintaining comparable visual quality.
Deep Neural Networks (DNNs) require large amounts of annotated training data for good performance. Often this data is generated using manual labeling (error-prone and time-consuming) or rendering (requiring geometry and material information). Both approaches make it difficult or uneconomic to apply them to many small-scale applications. A fast and straightforward approach of acquiring the necessary training data would allow the adoption of deep learning to even the smallest of applications. Chroma keying is the process of replacing a color (usually blue or green) with another background. Instead of chroma keying, we propose luminance keying for fast and straightforward training image acquisition. We deploy a black screen with high light absorption (99.99\%) to record roughly 1-minute long videos of our target objects, circumventing typical problems of chroma keying, such as color bleeding or color overlap between background color and object color. Next we automatically mask our objects using simple brightness thresholding, saving us from manual annotation. Finally, we automatically place the objects on random backgrounds and train a 2D object detector. We do extensive evaluation of the performance on the widely-used YCB-V object set and compare favourably to other conventional techniques such as rendering without needing 3D meshes, materials or any other information of our target objects and in a fraction of the time needed for other approaches. Our work demonstrates highly accurate training data acquisition allowing to start training state-of-the-art networks within minutes.
Many tasks in computer graphics and engineering involve unstructured tetrahedral meshes. Numerical methods such as the finite element method (FEM) oftentimes use tetrahedral meshes to compute a solution for complex problems such as physically-based simulation or shape deformation. As each tetrahedron costs computationally, coarsening tetrahedral meshes typically reduces the overhead of numerical methods, which is attractive for interactive applications. In order to enable reduction of the tetrahedron count, we present a quick adaptive coarsening method for unstructured tetrahedral meshes. Our method collapses edges using the massively parallel processing power of
A cross polarization could be indispensable in certain applications when scanning and digitizing highly reflective materials or when certain applications couldn’t afford following the recommended imaging geometry 00/450 | 45o/0o for some technical reasons. However, that puts very much color fidelity in question, to which extent a cross polarization may impact the source illuminant in the first place that is consequently impacting the color appearance during the imaging and the color correction procedures. In this research we show how certain cross polarization setups are adding a chroma tint to the light source, D50 in this study, causing by that undesirable color shift of the color of the light source. Consequently, a shift in its color correlated temperature moving, in worst case scenario, from ~5000K to ~4500K and resulting in an increased DE00 as a result of the added chroma when compared against a standard D50; nearly doubled in best case scenario and nearly tripled in worst case scenario.
Image segmentation (or masking) finds a very useful use case within 3D reconstruction of cultural heritage objects. The 3D re-constructions can be accelerated, reconstructing the object without any background noise. Conventional segmentation methods can calculate erroneous masks for certain objects and environments, which can lead to errors within the reconstruction: Parts of the 3D reconstruction may be missing or are incorrectly reconstructed, which contradicts adequate archiving. The automated iterative Multi-View Stereo (MVS) scanning process makes it necessary to obtain masks that reconstruct the object in the best possible way, regardless of the environment, the stabilizing mount, the color of the background and the object. In addition, it should not be necessary to tweak the best possible parameters for conventional masking procedures and to create masks manually. State-of-the-art artificial intelligence (AI) segmentation networks will be trained and applied to the MVS scans to verify the behavior of the associated 3D reconstructions and the automated iterative scanning process. In addition, a comparison between different AI segmentation networks and a comparison between conventional masking methods and AI segmentation networks is performed.
Neural Radiance Fields have revolutionized Novel View Synthesis by providing impressive levels of realism. However, in most in-the-wild scenes they suffer from floater artifacts that occur due to sparse input images or strong view-dependent effects. We propose an approach that uses neighborhood based clustering and a consistency metric on NeRF models trained on different scene scales to identify regions that contain floater artifacts based on Instant-NGPs multiscale occupancy grids. These occupancy grids contain the position of relevant optical densities in the scene. By pruning the regions that we identified as containing floater artifacts, they are omitted during the rendering process, leading to higher quality resulting images. Our approach has no negative runtime implications for the rendering process and does not require retraining of the underlying Multi Layer Perceptron. We show on a qualitative base, that our approach is suited to remove floater artifacts while preserving most of the scenes relevant geometry. Furthermore, we conduct a comparison to state-of-the-art techniques on the Nerfbusters dataset, that was created with measuring the implications of floater artifacts in mind. This comparison shows, that our method outperforms currently available techniques. Our approach does not require additional user input, but can be be used in an interactive manner. In general, the presented approach is applicable to every architecture that uses an explicit representation of a scene's occupancy distribution to accelerate the rendering process.
Color targets come in different designs, sizes and surface finishes. A high quality color target such as the Next Generation Target (NGT)1, designed for the Library of Congress, has a glossy finish that makes it sensitive to the light-setup geometry. When the NGT color target is to be captured orthogonally, i.e. both the camera and the light share the same plane and lie on the normal of the target’s surface, even with cross-polarization in place it is not possible to completely eliminate the high reflections caused by the camera/light geometry – unlike for less glossy color targets such as the X-Rite SG CC- not even if the camera/light setup were to be tilted at different angles. We are demonstrating in this paper that it is possible, however, to deploy a mosaic approach to capture the NGT color target at a tilted angle, masking out the reflections, and composing a rectified mosaic image out of only the clear parts of the target. The resultant ICC color correction profile for the mosaic image is proved to be viable to put in use and it satisfies all the necessary metrics for ISO level ”A” when it comes to color calibration and color accuracy.
Transparency detection is a hard problem, as suggested by animals and humans flying or running into glass. However, humans seem to be able to learn and improve on the task with experience, begging the question, whether computers are able to do so too. Making a computer learn and understand transparency would be beneficial for moving agents, such as robots or autonomous vehicles. Our contributions are threefold: First, we conducted a perception study to obtain insights about human transparency detection methods, when borders of transparent objects are not visible. Second, based on our study insights we created a novel synthetic dataset called DISTOPIA , which focuses on the warping properties of transparent objects, placed in a variety of natural scenes and contains over 140 000 high resolution images. Third, we modified and trained a deep neural network classification model with an attention module to detect transparency through warping. Our results show that a neural network trained on synthetic data depicting only distortion effects can solve the transparency detection problem and surpasses human performance.
The use of computer-aided methods for the design of parts that must meet functional or stability requirements typically consists of an iterative cycle of design, physical simulation and testing or analysis, followed by redesign, etc. Each step is often performed with a domain-specific tool, e.g., a specific CAD modeling suite. This results in the need to convert the model representation between steps, such as meshing for finite element simulation for example. In recent work, a distributed application framework has been proposed that allows for the interactive modification and simulation of tetrahedral meshes derived from existing CAD models, e.g., to create customized versions of parts that were designed for mass production. This shortens the design cycle by eliminating the need for conversion and switching between tools. In this paper, we present a more detailed description and improvements to this architecture by using GPU parallelization not only for simulation but also for mesh editing, which leads to even shorter iteration cycles.
Pedro Santos合作论文数Instituto Superior Tecnico;Departamento de Matematica13
Dirk Burkhardt合作论文数Darmstadt University of Applied Sciences, Research Group on Human-Computer Interaction and Visual Analytics11