
This paper describes an innovative extension of morphological operators to the analysis and processing of three dimensional images represented by voxels. In particular, the morphological skeleton well approximates the Medial Axis of a 3D object rendered using a net of Kinect devices: it has proven to be a really effective tool to obtain a compact representation of the object surface that is accurate and very inexpensive in terms of computational time. The reconstructed surface is representable either as a union of balls or as an iso-surface of a 3D function defined as a linear combination of elementary functions with radial support. Furthermore, the representation is easily processable and hierarchically interpretable. The morphological operators neither require any prior information on the reconstructed volume, nor impose any resolution requirement on the object under analysis. The only input for morphological operators is a volume of voxels. Information on the volume nature or resolution requirements are related only to the particular application to which the morphological operators are applied, but do not limit the applicability of any of the proposed morphological operators.
We present an algorithm to detect and track color calibration charts in images robustly. A model of the calibration chart describes the geometry of the chart regions by polygons along with the related reference colors. A homography maps the model into the image and an optimizer adapts the mapping parameters to align the model regions to the chart in the image. We propose a cost function that evaluates the quality of model alignment based on color statistics and an efficient method to extract color statistics in polygonal image regions using the integral image. The algorithm measures the colors of the chart in the image and determines correction parameters. Without loss of generality, we describe the method using the X-Rite ColorChecker Classic as a typical chart. Experiments show the robustness to noise and blur and the real-time capability of the system.
This paper describes the development of a new texture based segmentation algorithm which uses a set of features extracted from Grey-Level Co-occurrence Matrices. The proposed method segments different textures based on noise reduced features which are effective texture descriptor. Each of the features is processed including normalisation and noise removal. Principal Component Analysis is used to reduce the dimensionality of the resulting feature space. Gaussian Mixture Modelling is used for the subsequent segmentation and false positive regions are removed using morphology. The evaluation includes a wide range of textures (more than 80 Brodatz textures) and in comparison (both qualitative and quantitative) with state of the art techniques very good segmentation results have been obtained.
In a stereo camera set-up for 3D cinema, much care is taken to ensure that the cameras are both the same model and use identical parameter values. But an imperfect adjustment, the use of beam splitters, as well as unavoidable differences in the characteristics of the cameras and lens systems, often cause visible differences in color among the two views. Many of these differences appear to be local and therefore cannot be fully removed by common color matching approaches, which are typically global methods. We propose a simple yet effective method for local color matching, that operates in three steps. First, one view is chosen as the target and morphed so that it becomes locally registered with the source view, the one whose colors will be modified. Next, color matching is performed on the source view by aligning each local histogram with the corresponding local histogram in the warped target view. Finally, Poisson editing [9] is applied on the source regions which have no correspondence in the target. For several professional, high quality examples of 3D sequences, we show that our algorithm outperforms four different conventional color matching approaches.
Human motion estimation is a topic receiving high attention during the last decades. There is a vast range of applications that employ human motion tracking, while the industry is continuously offering novel motion tracking systems, which are opening new paths compared to traditionally used passive cameras. Motion tracking algorithms, in their general form, estimate the skeletal structure of the human body and consider it as a set of joints and limbs. However, human motion tracking systems usually work on a single sensor basis, hypothesizing on occluded parts. We hereby present a methodology for fusing information from multiple sensors (Microsoft's Kinect sensors were utilized in this work) based on a series of factors that can alleviate from the problem of occlusion or noisy estimates of 3D joints' positions.
The success of matching algorithms relies on the definition of features which are both invariant against the geometric distortions to be considered, and distinctive enough to avoid ambiguities. This paper addresses the problem of color feature points matching under photometric and geometric changes. Considering the popular SURF descriptor, it analyzes its state-of-the-art color versions, and proposes a new extension by using local histogram equalization (LHE). While most existing descriptors stem from color conversions and apply to standard lighting variations acquired by the same device, the proposed feature is device-independent and could fit to very generic changes. The experimental results show that the proposed color descriptors outperform the existing ones under some types of distortions, and are more precise and invariant to different color variations. The paper considers Projector-based Augmented Reality (PAR) as an application field, where one of the evaluation criteria is homography accuracy between real and estimated distorted images. The results show that the proposed method gives the most stable results over all the other techniques and therefore they justify its use for robust color feature matching and its application to geometric correction.
This paper presents two learning based algorithms that are designed for the purpose of extracting and processing suitable information in images for the visually impaired. Both algorithms are developed to be used within a specific modular sonification system. This system is designed to allow visually impaired people to explore images, actively on a touch screen, and to receive an auditory response about the image content at any current finger position. The first algorithm presented in this paper therefore addresses the problem of labeling regions within images, incorporating spatial dependencies. The second algorithm strives to alleviate the rejection of false object detections before sonification. This is crucial to avoid confusion on the side of the blind user, who can not check for a correct image labeling or object detection visually. Due to the modular design principle of the modular sonification system, both algorithms can be incorporated easily and efficiently.
We present a novel 3D-image-based client-server streaming approach, as an alternative to current video-based streaming for engineering applications on mobile devices. A render server generates multi-layered panoramic 3D impostors on-the-fly. The impostors are compressed and then transmitted to the client. Since impostors can be reused for several frames of a camera path, the necessary bandwidth can be reduced to a level much below video compression. Furthermore, this approach completely eliminates latencies, which are a major problem of 2D-image-based streaming methods. Also, client rendering speed is decoupled from scene size and complexity and server rendering load can be reduced many times compared to continuous 2D streaming. We demonstrate navigating through a CAD model of a Boeing 777 passenger aircraft with 350 million triangles, on a smartphone, remotely connected via an EDGE network, with zero latency, at 60 frames per second. This makes the approach suitable also for Augmented Reality applications.
We propose a method for creating flexible-viewpoint facial video from monocular input employing highly detailed static 3D reconstructions of an actor's head. We use the term flexible-viewpoint to indicate that the viewpoint can be arbitrarily chosen (as in free-viewpoint video), but from a restricted set of viewing directions. Our method enables dynamic changes of the viewpoint without requiring estimation of the head geometry form the video sequences which is hard with the methods typically used for creating free-viewpoint video. Alongside video capture of a certain actor, we record static high resolution stereo images of the actor's head and face. From these images, we create a detailed 3D model of the head by image-based reconstruction methods. We propose two methods to register this 3D model to the a starting frame of the video stream following the actors head and facial action. Furthermore, we show how model-based tracking over the whole video sequence provides precise head pose estimates for each video frame. Once the registration is complete, the 3D model serves as geometric proxy for image-based rendering techniques in order to create novel viewpoints using the video stream as texture.
We present a solution to the people tracking problem using a monocular vision approach from a bird's eye view and Sequential Monte-Carlo Filtering. Each tracked human is represented by an individual Particle Filter using spheroids as a three-dimensional approximation to the shape of the upstanding human body. We use the bearings-only model as the state update function for the particles. Our measurement likelihood function to estimate the probability of each particle is imitating the image formation process. This involves also partial occlusion by dynamic movements from other humans within neighbored areas. Due to algorithmic optimization the system is real-time capable and therefore not only limited to surveillance or human motion analysis. It could rather be used for Human-Computer-Interaction (HCI) and indoor location. To demonstrate this capabilities we evaluated the accuracy of the system and show the robustness in different levels of difficulty.
This work presents an evaluation of different detectors used for the Edge Error measure, which is one of the most used features in Image Quality Assessment (IQA). Among detectors, the proposed Wavelet--based Edge Detector (WEE) is also evaluated. the test is carried out for the quality assessment of images which are distorted by either JPEG, JPEG2000, Fast Fading, Gaussian Blur or White Noise process. The measures with high correlation values are considered as more accurate for human--based evaluations, namely perceptual . As a result, we found the following: firstly, the measure can be considered as perceptual for images distorted by JPEG and White Noise processes using any of the considered detectors, specially for the measure using wavelet--based edge detector (0:94 Spearman, 0:81 Kendall and 0:93 Pearson); secondly, Edge Error can be considered perceptual by using a Canny detector to evaluate Fast--Fading, Gaussian Blur, JPEG and White Noise distortion types. Finally, WEE achieve high correlation values for all considered distortion types, over performing the other considered edge error measures. As a conclusion, we select the WEE and Canny--based measures as perceptual, confirming that an edge enhancement method can improve the measurement accuracy.
Meeting, socializing and conversing online with a group of people using teleconferencing systems is still quite different from the experience of meeting face to face. We are abruptly aware that we are online and that the people we are engaging with are not in close proximity. Analogous to how talking on the telephone does not replicate the experience of talking in person. Several causes for these differences have been identified and we propose inspiring and innovative solutions to these hurdles in attempt to provide a more realistic, believable and engaging online conversational experience. We present the distributed and scalable framework REVERIE that provides a balanced mix of these solutions. Applications build on top of the REVERIE framework will be able to provide interactive, immersive, photo-realistic experiences to a multitude of users that for them will feel much more similar to having face to face meetings than the experience offered by conventional teleconferencing systems.
This paper introduces a novel framework for full body human motion reconstruction from 2D video data using a motion capture database as knowledge base containing information on how people move. By extracting suitable two-dimensional features from both, the input video sequence and the motion capture database, we are able to employ an efficient retrieval technique to run a data-driven optimization. Only little preprocessing is needed by our method, the reconstruction process runs close to real time. We evaluate the proposed techniques on synthetic two-dimensional input data obtained from motion capture data and on real video data.
In this paper, we present an innovative analytic algorithm for tomographic reconstruction from fewer numbers of projections. Back-projection has been customized to make it work even when the projections are not uniformly distributed, and (or) are missing at certain orientation(s). Contour information of the object has been used efficiently to ignore all points/pixels that lie outside the objects boundary. Aiming successful reconstruction with minimum number of projections an innovative interpolation methodology has been proposed to figure out all the missing projections. Based on the experiments on simulated and real medical images it has been shown that the proposed modality is capable of producing better reconstruction than the state-of-the-art methods with comparatively less number of projections.
This paper deals with the problem of computing a semantic segmentation of an image via label transfer from an already labeled image set. In particular it proposes a method that takes advantage of sparse 3D structure to infer the category of superpixel in the novel image. The label assignment is computed by a Markov random field that has the superpixels of the image as nodes. The data term combines labeling proposals from the appearance of the superpixel and from the 3D structure, while the pairwise term incorporates spatial context, both in the image and in 3D space. Exploratory results indicate that 3D structure, albeit sparse, improves the process of label transfer.
This paper presents a complete system for computer graphics compression and transmission showing MPEG-4 graphics compression standard tool performance to encode 3D graphics objects. We employed a web based receiver to decode and render 3D objects. We provide a compression evaluation of user generated scanned 3D body models, which are huge in size and presented in XML format. A web based daemon is developed that can receive, decode and render such encoded 3D graphics objects.
This paper presents a novel way to embed local binary texture information in the form of local binary patterns (LBP) into the covariance descriptor. Contrary to previous publications, our method is not based on the LBP decimal values where arithmetic operations have no texture meaning. Our method uses the angles described by the uniform LBP patterns and includes them into the set of features used to build the covariance descriptor. Our representation is not only more compact but more robust because it is less affected by noise and small neighborhood rotations. Experimental evaluations corroborate the performance of our descriptor for texture analysis and tracking applications. Our descriptor rivals with state-of-the-art methods and beats other covariance-based descriptors.
Free-viewpoint video renderers (FVVR) allow a user to view captured video footage from any position and direction. Despite the obvious appeal of such systems, they have yet to make a major impact on digital entertainment. Current FVVR implementations have been on desktop computers. Media consumption is increasingly through mobile devices, such as smart phones and tablets; adapting FVVR to mobile platforms will open this new form of media up to a wider audience. An efficient, high-quality FVVR, which runs in real time with user interaction on a mobile device, is presented. Performance is comparable to recent desktop implementations. The FVVR supports relighting and integration of relightable free-viewpoint video (FVV) content into computer-generated scenes. A novel approach to relighting FVVR content is presented which does not require prior knowledge of the scene illumination or accurate surface geometry. Surface appearance is separated into a detail component, and a set of materials with properties determining surface colour and specular behaviour. This allows plausible relighting of the dynamic FVV for rendering on mobile devices.
Snap Composition broadens the applicability of interactive image composition. Current tools, like Adobe's Photomerge Group Shot, do an excellent job when the background can be aligned and objects have limited motion. Snap Composition works well even when the input images include different objects and the backgrounds cannot be aligned. The power of Snap Composition comes from the ability to assign for every output pixel a source pixel in any input image, and from any location in that image. An energy value is computed for each such assignment, representing both the user constraints and the quality of composition. Minimization of this energy gives the desired composition.Composition is performed once a user marks objects in the different images, and optionally drags them into a new location in the target canvas. The background around the dragged objects, as well as the final locations of the objects themselves, will be automatically computed for seamless composition. If the user does not drag the selected objects to a desired place, they will automatically snap into a suitable location. A video describing the results can be seen in www.vision.huji.ac.il/shiftmap/SnapVideo.mp4.
Active contours or snakes are widely used for segmentation and tracking. These techniques require the minimization of an energy function, which is generally a linear combination of a data fit term and a regularization term. This energy function can be adjusted to exploit the intrinsic object and image features. This can be done by changing the weighting parameters of the data fit and regularization term. There is, however, no rule to set these parameters optimally for a given application. This results in trial and error parameter estimation. In this paper, we propose a new active contour framework defined using probability theory. With this new technique there is no need for ad hoc parameter setting, since it uses probability distributions, which can be learned from a given training dataset.