A huge fraction of cameras used nowadays is based on CMOS sensors with a rolling shutter that exposes the image line by line. For dynamic scenes/cameras this introduces undesired effects like stretch, shear and wobble. It has been shown earlier that rotational shake induced rolling shutter effects in hand-held cell phone capture can be compensated based on an estimate of the camera rotation. In contrast, we analyse the case of significant camera motion, e.g.\ where a bypassing street level capture vehicle uses a rolling shutter camera in a 3D reconstruction framework. The introduced error is depth dependent and cannot be compensated based on camera motion/rotation alone, invalidating also rectification for stereo camera systems. On top, significant lens distortion as often present in wide angle cameras intertwines with rolling shutter effects as it changes the time at which a certain 3D point is seen. We show that naive 3D reconstructions (assuming global shutter) will deliver biased geometry already for very mild assumptions on vehicle speed and resolution. We then develop rolling shutter dense multiview stereo algorithms that solve for time of exposure and depth at the same time, even in the presence of lens distortion and perform an evaluation on ground truth laser scan models as well as on real street-level data.
Google Street View has provided millions of users with the ability to visually locate businesses around the world using 360° panoramic imagery. Due to the bulky, custom hardware required to precisely geo-locate the imagery, such as laser scanners and high precision GPS devices, Street View experiences have been limited to large areas that can provide cost-effective collections. This has prevented users from discovering places such as the interiors of small businesses.
We present a novel algorithm that takes as input an uncalibrated unordered set of spherical panoramic images and outputs their relative pose up to a global scale. The panoramas contain both indoor and outdoor shots and each set was taken in a particular indoor location e.g. a bakery or a restaurant. The estimated pose is used to build a map of the location, and allow easy visual navigation and exploration in the spirit of Google's Street View. We also present a dataset of 9 sets of panoramas, together with an annotation tool and ground truth point correspondences. The manual annotations were used to obtain ground truth relative pose, and to quantitatively evaluate the different parameters of our algorithm, and can be used to benchmark different approaches. We show excellent results on the dataset and point out future work.
Camera calibration is essential to many computer vision applications. Most existing calibration algorithms in the literature consist of estimating both intrinsic and extrinsic camera parameters through non-linear iterative minimizations using only point features. We show that using lines and planes in a consistant framework allows to rst entirely decouple intrinsic and extrinsic camera parameters, and then extract closed form solutions for them making optimal use of the geometrical constraints existing in the scene. We introduce a new formalism, which we calìdual space', that allows one to represent cleanly and intuitively many useful geometrical relashionships between planes, lines and points in space, and lines and points on the image plane. Calibration results are presented on real images retrieved from a standard image database.
With the advent and proliferation of low cost and high performance digital video recorder devices, an increasing number of personal home video clips are recorded and stored by the consumers. Compared to image data, video data is lager in size and richer in multimedia content. Efficient access to video content is expected to be more challenging than image mining. Previously, we have developed a content-based image retrieval system and the benchmarking framework for personal images. In this paper, we extend our personal image retrieval system to include personal home video clips. A possible initial solution to video mining is to represent video clips by a set of key frames extracted from them thus converting the problem into an image search one. Here we report that a careful selection of key frames may improve the retrieval accuracy. However, because video also has temporal dimension, its key frame representation is inherently limited. The use of temporal information can give us better representation for video content at semantic object and concept levels than image-only based representation. In this paper we propose a bottom-up framework to combine interest point tracking, image segmentation and motion-shape factorization to decompose the video into spatiotemporal regions. We show an example application of activity concept detection using the trajectories extracted from the spatio-temporal regions. The proposed approach shows good potential for concise representation and indexing of objects and their motion in real-life consumer video.
While image clustering has many important applications ranging from personal to web image management, its use is often limited by the difficulty of extracting reliable semantics from low level image features. The image clusters can be improved by using features extracted from image regions rather than the whole image. Region segmentation can be improved in turn, by considering all images within the same cluster rather than segmenting each image independently. This observation leads to the unified Bayesian framework for image clustering and segmentation presented in this paper. The experimental results, reported using several types of visual feature extractors on a database of web documents containing over 6000 images, illustrates a significant improvement over existing techniques.
With the proliferation of digital cameras, the size of personal media data such as digital photos, videos, etc. has grown extremely large. The personal nature of the data has heightened the demands for a media management system on personal desktops. Existing solutions for media management target mostly server-based Web databases and rely on extensive metadata (i.e., labels) generation to aid retrieval. Personal media databases, on the other hand, have very limited labels generated by the end users themselves. This paper introduces a method for learning concept templates from web images to query personal image databases. The proposed method has the advantage of leveraging Web resources to ease personal photo retrieval in order to avoid costly annotation of personal image databases.
With the advent and proliferation of digital cameras and computers, the number of digital photos created and stored by consumers has grown extremely large. This created increasing demand for image retrieval systems to ease interaction between consumers and personal media content. Active learning is a widely used user interaction model for retrieval systems, which learns the query concept by asking users to label a number of images at each iteration. In this paper, we study sampling strategies for active learning in personal photo retrieval. In order to reduce human annotation efforts in a content-based image retrieval setting, we propose using multiple sampling criteria for active learning: informativeness, diversity and representativeness. Our experimental results show that by combining multiple sampling criteria in active learning, the performance of personal photo retrieval system can be significantly improved
Most programs are repetitive, where similar behavior can be seen at different execution times. Algorithms have been proposed that automatically group similar portions of a program's execution into phases, where samples of execution in the same phase have homogeneous behavior and similar resource requirements. In this paper, we examine applying these phase analysis algorithms and how to adapt them to parallel applications running on shared memory processors. Our approach relies on a separate representation of each thread's activity. We first focus on showing its ability to identify similar intervals of execution across threads for a single run. We then show that it is effective at identifying similar behavior of a program when the number of threads is varied between runs. This can be used by developers to examine how different phases scale across different number of threads. Finally, we examine using the phase analysis to pick simulation points to guide multithreaded simulation.
Cast indexing is a very important application for content-based video browsing and retrieval, since the characters in feature-length films and TV series are always the major focus of interest to the audience. By cast indexing, we can discover the main cast list from long videos and further retrieve the characters of interest and their relevant shots for efficient browsing. This paper proposes a novel cast indexing approach based on hierarchical clustering, semi-supervised learning and linear discriminant analysis of the facial images appearing in the video sequence. The method first extracts local SIFT features from detected frontal faces of each shot, and then utilizes hierarchical clustering and Relevant Component Analysis (RCA) to discover main cast. Furthermore, according to the user's feedback, we project all the face images to a set of the most discriminant axes learned by Linear Discriminant Analysis (LDA) to facilitate the retrieval of relevant shots of specified person. Extensive experimental results on movie and TV series demonstrate that the proposed approach can efficiently discover the main characters in such videos and retrieve their associated shots.
It is now common to have accumulated tens of thousands of personal pictures. Efficient access to that many pictures can only be done with a robust image retrieval system. This application is of high interest to Intel processor architects. It is highly compute intensive. and could motivate end users to upgrade their personal computers to the next generations of processors. A key question is how to assess the robustness of it personal image retrieval system. Personal image databases are very different front digital libraries that have been used by many Content Based Image Retrieval Systems.(1) For example a personal image database has a lot of pictures of people, but a small set of different people typically family, relatives, and friends. Pictures are taken ill it limited set of places like home, work, school, and vacation destination. The most frequent queries are. searched for people. and for places. These attributes, and many others affect how a personal image retrieval system should be benchmarked, and benchmarks need to be different from existing ones based on art images, or medical images for examples. The attributes of the data set do not change the list of components needed for the benchmarking of such systems as specified in(2):center dot data sets center dot query tasks center dot ground truth center dot evaluation measures center dot benchmarking events.This paper proposed it way to build these components to be representative of personal image databases, and of the corresponding usage models.
Simulating chip-multiprocessor systems (CMP) can take a long time. For single-threaded workloads, earlier work has shown the utility of phase analysis , that is identification of repetitive program behaviors, in reducing overall simulation time while maintaining an acceptable loss in accuracy. To cope with multithreaded workloads, a combination of phases from all executing threads must be taken into consideration since inter-thread interference may distort the homogeneity of each phases' true performance. Unfortunately, phase analysis does not work for multithreaded (MT) workloads because the possible phase combinations in an inherently nondeterministic execution model grows exponentially with the number of threads. To this end, we propose a new technique to reduce the number of simulation samples by synthesizing samples from similar phase combinations. We present a simple cost function for measuring the similarity between phase combinations and by using the individual thread samples from the similar phase combinations, a new sample can be constructed. This cost function provides a convenient control knob for exploiting tradeoffs between simulation speed and accuracy. Our experimental results show that in most cases, properly setting the cost function's threshold can yield a reduction in sampling by 90%, while maintaining error to less than 5%.
Igor Kozintsev合作论文数Intel Microprocessor Research Lab7
Ehud Rivlin合作论文数 Technion-Israel Institute of Technology;Computer Science Department 1
Salih Burak Göktürk合作论文数Ojos, Inc.1