We introduce 3D Moments, a new computational photography effect. As input we take a pair of near-duplicate photos, i.e., photos of moving subjects from similar viewpoints, common in people's photo collections. As output, we produce a video that smoothly interpolates the scene motion from the first photo to the second, while also producing camera motion with parallax that gives a heightened sense of 3D. To achieve this effect, we represent the scene as a pair of feature-based layered depth images augmented with scene flow. This representation enables motion interpolation along with independent control of the camera viewpoint. Our system produces photorealistic space-time videos with motion parallax and scene dynamics, while plausibly recovering regions occluded in the original views. We conduct extensive experiments demonstrating superior performance over baselines on public datasets and in-the-wild photos. Project page: https://3d-moments.github.io/.
Single image 3D photography enables viewers to view a still image from novel viewpoints. Recent approaches combine monocular depth networks with inpainting networks to achieve compelling results. A drawback of these techniques is the use of hard depth layering, making them unable to model intricate appearance details such as thin hair-like structures. We present SLIDE, a modular and unified system for single image 3D photography that uses a simple yet effective soft layering strategy to better preserve appearance details in novel views. In addition, we propose a novel depth-aware training strategy for our inpainting module, better suited for the 3D photography task. The resulting SLIDE approach is modular, enabling the use of other components such as segmentation and matting for improved layering. At the same time, SLIDE uses an efficient layered depth formulation that only requires a single forward pass through the component networks to produce high quality 3D photos. Extensive experimental analysis on three view-synthesis datasets, in combination with user studies on in-the-wild image collections, demonstrate superior performance of our technique in comparison to existing strong baselines while being conceptually much simpler. Project page: https://varunjampani.github.io/slide
We present a multi-modal approach for automatically generating hierarchical tutorials from instructional makeup videos. Our approach is inspired by prior research in cognitive psychology, which suggests that people mentally segment procedural tasks into event hierarchies, where coarse-grained events focus on objects while fine-grained events focus on actions. In the instructional makeup domain, we find that objects correspond to facial parts while fine-grained steps correspond to actions on those facial parts. Given an input instructional makeup video, we apply a set of heuristics that combine computer vision techniques with transcript text analysis to automatically identify the fine-level action steps and group these steps by facial part to form the coarse-level events. We provide a voice-enabled, mixed-media UI to visualize the resulting hierarchy and allow users to efficiently navigate the tutorial (e.g., skip ahead, return to previous steps) at their own pace. Users can navigate the hierarchy at both the facial-part and action-step levels using click-based interactions and voice commands. We demonstrate the effectiveness of segmentation algorithms and the resulting mixed-media UI on a variety of input makeup videos. A user study shows that users prefer following instructional makeup videos in our mixed-media format to the standard video UI and that they find our format much easier to navigate.
We present a new framework for sketch-based modeling and animation of 3D organic shapes that can work entirely in an intuitive 2D domain, enabling a playful, casual experience. Unlike previous sketch-based tools, our approach does not require a tedious part-based multi-view workflow with the explicit specification of an animation rig. Instead, we combine 3D inflation with a novel rigidity-preserving, layered deformation model, ARAP-L, to produce a smooth 3D mesh that is immediately ready for animation. Moreover, the resulting model can be animated from a single viewpoint - and without the need to handle unwanted inter-penetrations, as required by previous approaches. We demonstrate the benefit of our approach on a variety of examples produced by inexperienced users as well as professional animators. For less experienced users, our single-view approach offers a simpler modeling and animating experience than working in a 3D environment, while for professionals, it offers a quick and casual workspace for ideation.
Video summaries are a popular way to share important events, but creating good summaries is hard. It requires expertise in both capturing and editing footage. While hiring a professional videographer is possible, this is too costly for most casual events. An alternative is to place 360 video cameras around an event space to capture footage passively and then extract regular field-of-view (RFOV) shots for the summary. This paper focuses on the problem of extracting such RFOV shots. Since we cannot actively control the cameras or the scene, it is hard to create "ideal" shots that adhere strictly to traditional cinematography rules. To better understand the tradeoffs, we study human preferences for static and moving camera RFOV shots generated from 360 footage. From the findings, we derive design guidelines. As a secondary contribution, we use these guidelines to develop automatic algorithms that we demonstrate in a prototype user interface for extracting RFOV shots from 360 videos.
Evaluation methodologies provide a better understanding of the relationship between a technique and the image attributes. Metrics are used to evaluate the similarities between images. They may use different approaches depending on what needs to be achieved. If only objective values need to be compared statistics-based metrics are suitable. A number of image comparison metrics have been proposed in the literature that is based mostly on images statistics. This chapter shows the MATLAB code for the root mean square error (RMSE) calculation. It presents the MATLAB code for the computation of mean square error (MSE) between two images. Peak signal to noise ratio (PSNR) is another widely used metric, which takes into account the maximum value of the signal, and can be defined based on MSE. Metrics enable automation and the use of metrics can also be applied directly to methods, for example, to …
course Share on How to write a SIGGRAPH paper: a guide to choosing a good research topic, doing the research, and writing it up Author: David Salesin View Profile Authors Info & Claims SA '16: SIGGRAPH ASIA 2016 CoursesNovember 2016 Article No.: 3Pages 1–103https://doi.org/10.1145/2988458.2988471Published:28 November 2016Publication History 0citation623DownloadsMetricsTotal Citations0Total Downloads623Last 12 Months28Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
This paper introduces and presents a solution to the “Escherization” problem: given a closed figure in the plane, find a new closed figure that is similar to the original and tiles the plane.... Introduction ... Problem (“ESCHERIZATION”): Given a closed plane figure S (the “goal shape”), find a new closed figure T such that: 1. T is as close as possible to S; and 2. copies of T fit together to form a tiling of the plane. Problem (“IMAGE ANALOGIES”): Given a pair of images A and A’ (the unfiltered and filtered source images, respectively), along with some additional unfiltered target image B, synthesize a new filtered target image B’ such that
We describe a new vector-based primitive for creating smooth-shaded images, called the diffusion curve . A diffusion curve partitions the space through which it is drawn, defining different colors on either side. These colors may vary smoothly along the curve. In addition, the sharpness of the color transition from one side of the curve to the other can be controlled. Given a set of diffusion curves, the final image is constructed by solving a Poisson equation whose constraints are specified by the set of gradients across all diffusion curves. Like all vector-based primitives, diffusion curves conveniently support a variety of operations, including geometry-based editing, keyframe animation, and ready stylization. Moreover, their representation is compact and inherently resolution-independent. We describe a GPU-based implementation for rendering images defined by a set of diffusion curves in realtime. We then demonstrate an interactive drawing system for allowing artists to create artworks using diffusion curves, either by drawing the curves in a freehand style, or by tracing existing imagery. The system is simple and intuitive: we show results created by artists after just a few minutes of instruction. Furthermore, we describe a completely automatic conversion process for taking an image and turning it into a set of diffusion curves that closely approximate the original image content.
A simple recognition system is proposed which clusters the gray scale images using K-means algorithm based on wavelet features. The method is based on information extracted from the images known as features extraction. The features are extracted by using the following process: the image is decomposed by applying 2Ddiscrete wavelet transform (DWT) for one, two, three and four levels. The dimensionality of the image data is reduced up to desired level by the application of wavelets. The decomposed coefficients of an image are considered as the feature sets. The four methods of reducing dimensions are applied on a specific set of images to obtain four different data sets which serve as input to the k-means algorithm for clustering. The number of clusters is fixed prior to the experiments. The relative performances of k-means based on four different data sets are evaluated in terms of clustering accuracy and CPU time consumed.
Incorporating the individual and collective problem solving skills of non-experts into the scientific discovery process could potentially accelerate the advancement of science. This paper discusses the design process used for Foldit, a multiplayer online biochemistry game that presents players with computationally difficult protein folding problems in the form of puzzles, allowing ordinary players to gain expertise and help solve these problems. The principle challenge of designing such scientific discovery games is harnessing the enormous collective problem-solving potential of the game playing population, who have not been previously introduced to the specific problem, or, often, the entire scientific discipline. To address this challenge, we took an iterative approach to designing the game, incorporating feedback from players and biochemical experts alike. Feedback was gathered both before and after releasing the game, to create the rules, interactions, and visualizations in Foldit that maximize contributions from game players. We present several examples of how this approach guided the game's design, and allowed us to improve both the quality of the gameplay and the application of player problem-solving.
This chapter describes the Web Summaries system, which is designed to aid people in accomplishing exploratory Web research. Web Summaries enables users to produce automation artifacts, such as extraction patterns, relations, and personalized task-specific search templates, in the context of existing tasks. By leveraging the growing amount of structured Web pages and pervasive search capabilities, Web Summaries provides a set of semiautomatic interaction techniques for collecting and organizing personal Web content. First, it takes advantage of the growth in template material on the Web and designs semiautomatic interactive extraction of content from similar Web pages using structural and content extraction patterns. Second, it employs layout templates and user labeling to create rich displays of heterogeneous Web content collections. Third, it uses search technology for proactive retrieval of content from different, related Web sites through user-defined relations. Fourth, it let users define their own personalized and aesthetic views of heterogeneous content from any number of Web sites through cards. And finally, it introduces a new template-based search paradigm that combines the user-defined relations and cards into a search template. Search templates present a goal-driven search mechanism that creates visual personalized summaries of the content users need to accomplish their task.
This chapter describes the Web Summaries system, which is designed to aid people in accomplishing exploratory Web research. Web Summaries enables users to produce automation artifacts, such as extraction patterns, relations, and personalized task-specific search templates, in the context of existing tasks. By leveraging the growing amount of structured Web pages and pervasive search capabilities Web Summaries provides a set of semiautomatic interaction techniques for collecting and organizing personal Web content.
Destination maps are navigational aids designed to show anyone within a region how to reach a location (the destination). Hand-designed destination maps include only the most important roads in the region and are non-uniformly scaled to ensure that all of the important roads from the highways to the residential streets are visible. We present the first automated system for creating such destination maps based on the design principles used by mapmakers. Our system includes novel algorithms for selecting the important roads based on mental representations of road networks, and for laying out the roads based on a non-linear optimization procedure. The final layouts are labeled and rendered in a variety of styles ranging from informal to more formal map styles. The system has been used to generate over 57,000 destination maps by thousands of users. We report feedback from both a formal and informal user study, as well as provide quantitative measures of success.
Taking into account human visual characteristics, when isolated distortion pixels disperse sufficiently in an image, these pixels have minimal effect on the visual images. This study presents a color and brightness correction method based on lαβ color space transformation and excluding isolated pixels with brightness distortion. The experimental results show that the method can obtain better brightness and color uniformity in the vision, as well as make the histograms of corrected images in three channels get closer to normal distribution. Furthermore, the method suggested in this study can better reflect the transformation of surface features in the images.
This paper presents an approach to render novel views from input photographs, a task which is commonly referred to as image based rendering. We first compute dense view dependent depthmaps using consistent segmentation. This method jointly computes multiview stereo and segments input photographs while accounting for mixed pixels (matting). We take the images with depth as our input and then propose two rendering algorithms to render novel views using the segmentation, both realtime and off-line. We demonstrate the results of our approach on a wide variety of scenes.
We present the results of a ten-week field study that explored the use of automatic Web tools for collecting and organizing Web content in the context of users' personal tasks. Our find- ings show that people welcome automatic gathering of struc- tured information, such as job or rental listings, and are eager to use rich visualizations and displays of content they find on the Web. We also found that users collect a variety of Web content including a large amount of unstructured informa- tion and are interested in using automation not just for long- term content intensive tasks but also for short-lived transient tasks. Finally, we present a first exploration of an online col- laborative repository of user-defined semantic content. Our study participants used this repository and modified the col- laborative content to accomplish tasks.