
We present a sketch-based image retrieval system, designed to answer arbitrary queries that may go beyond searching for predefined object or scene categories. While sketching is fast and intuitive to formulate visual queries, pure sketch-based image retrieval often returns many outliers because it lacks a semantic understanding of the query. Our key idea is to combine sketch-based queries with inter-active, semantic re-ranking of query results. We leverage progress in deep learning and use a feature representation learned for image classification for re-ranking. This allows us to cluster semantically similar images, re-rank based on the clusters, and present more meaningful query results to the user. We report on two large-scale benchmarks and demonstrate that our re-ranking approach leads to significant improvements over the state of the art. Finally, a user study designed to evaluate a practical use case confirms the benefits of our approach.
From a user interaction perspective, speech and sketching make a good couple for describing motion. Speech allows easy specification of content, events and relationships, while sketching brings in spatial expressiveness. Yet, we have insufficient knowledge of how sketching and speech can be used for motion-based video retrieval, because there are no existing retrieval systems that support such interaction. In this paper, we describe a Wizard-of-Oz protocol and a set of tools that we have developed to engage users in a sketch- and speech-based video retrieval task. We report how the tools and the protocol fit together using "retrieval of soccer videos" as a use case scenario. Our software is highly customizable, and our protocol is easy to follow. We believe that together they will serve as a convenient and powerful duo for studying a wide range of multi-modal use cases.
We present a drawing assistant for sketching and for assisting users in shading a hand drawn sketch. The augmented reality based system uses a sketch made by a professional and uses it to help inexperienced users to do sketching and shading. The input image is converted to a set of points based on simple heuristics for providing a "connect the dots" interface for a user to aid sketching. With the help of a 2.5D mesh generated by our algorithm, the system assists the user by providing information about the colors that can be given in different parts of the sketch. The system was tested with users of different age groups and skill levels, indicating its usefulness.
In this manuscript, we describe a process that can be used to create still and/or animated portrait paintings to be shown in Expressive Art Exhibit. Our process consists of two stages: (1) Creation of control textures for a Barycentric shader by using color information gathered from photographs to provide realistic looking skin rendering; (2) Filtering and compositing the layers of images that are obtained by control textures, which correspond to effects such as diffuse, specular and ambient. To demonstrate proof-of-concept, we have created a few rigid body animations of painterly portraits under different lighting conditions.
Flowcharts play an important role when learning to program by conveying algorithms graphically and making them easy to read and understand. Computer-based flowchart design requires the user to learn the software first, which often results in a steep learning curve. Paper-drawn flowcharts don't provide feedback. We propose a system that allows users to draw their flowcharts directly on paper combined with a mobile phone app that takes a photo of the flowchart, interprets it, and generates and executes the resulting code. Flow2Code uses off-line sketch recognition and computer vision algorithms to recognize flowcharts drawn on paper. To gain practice and feedback with flowcharts, the user needs only a pencil, white paper, and a mobile device. The paper describes a tested system and algorithmic model for recognizing and interpreting offline flowcharts as well as a novel geometric feature, Axis Aligned Score (AAS), that enables fast accurate recognition of various quadrilaterals.
In this work we describe an interactive technique which enables gestural curve and surface design in an immersive virtual environment. We use a pair of motion tracked controllers to allow the user to intuitively control a Hermite spline curve which can be swept through space to create surfaces, or to define the location and orientation of cloned meshes. A head mounted display and tracked controllers replace the traditional keyboard and mouse for view selection and object interaction. Natural and expressive body motions allow the user to specify the parameters which define the curves and surfaces. Results are demonstrated using the HTC Vive VR system.
We propose a new taxonomy that explains the roles of motion in data visualization, focusing especially on their communicative aspects. Our taxonomy clarifies the main axis in how visualization designers can employ motion in data portrayal.
This article presents an easy to use mobile application which allows users to create 3D digital copies of their interested objects anywhere and anytime. An advanced 3-sweep modeling technique is developed to construct 3D primitives not only from generalized cylinder and cuboid, but also objects with symmetrical or non-uniformly scaled profiles. In addition, our system supports the texture and structure refinement which combine results created from multiple source images. The constructed 3D model will be the combination of our 3D primitives. The combined result can preserve more features which may not be seen from a single photo.
Design sketching is a powerful tool for expressing ideas from pen and paper effectively and becoming a more well-rounded communicator. Sketching instructors conventionally employ pen and paper in their classrooms to convey these fundamentals to students. However this traditional approach limits the bandwidth and capability of instructors to give timely and individualized feedback. An intelligent tutoring system can leverage the knowledge of domain expert design sketching instructors so that students can practice and receive real-time feedback outside of classroom hours. Our system leverages consulted instructor insights and observed pedagogical practices of an active university design sketching curriculum, and applies them in a mastery-based progression of exercises that utilize sketch recognition to give real-time feedback. An evaluation of our system's usability in a class of engineering students studying design sketching showed that it performed very well, was seen by the students as a motivating and intuitive practice tool, and allowed the students to improve the accuracy and speed of their sketches.
Sketching is one of the simplest ways to visualize ideas. Its key advantage is its easy availability and accessibility, as it require the user to have neither deep knowledge of a particular drawing program nor any advanced drawing skills. In practice, however, all these skills become necessary to improve the visual fidelity of the resulting drawing. In this paper, we present ShipShape—a general beautification assistant that allows users to maintain the simplicity and speed of freehand sketching while still taking into account implicit geometric relations to automatically rectify the output image. In contrast to previous approaches ShipShape works with general Bézier curves, enables undo/redo operations, is scale independent, and is fully integrated into Adobe Illustrator. We show various results to demonstrate the capabilities of the proposed method.
We introduce an interactive modeling tool for designing a smooth 3D normal field from the isophotes of a discretely shaded 2D image. Block or cartoon shading is a visual style in which artists depict a smoothly shaded 3D object using a small number of discrete brightness values, manifested as regions or bands of constant color. In our approach, artists trace isophotes, or curves of constant brightness, along the boundaries between constant color bands. Our algorithm first estimates light directions and computes 3D normals along the object silhouette and at intersections between isophotes from different light sources. We then propagate these 3D normals smoothly along isophotes, and subsequently throughout the interior of the shape. We describe our user interface for editing isophotes and correcting unintended normals produced by our algorithm. We validate our approach with a perceptual experiment and comparisons to ground truth data. Finally, we present a set of 3D renderings created using our interface.
Recently there has been a growing interest in sketch recognition technologies for facilitating human-computer interaction. Existing sketch recognition studies mainly focus on recognizing pre-defined symbols and gestures. However, just as there is a need for systems that can automatically recognize symbols and gestures, there is also a pressing need for systems that can automatically recognize pen-based manipulation activities (e.g. dragging, maximizing, minimizing, scrolling). There are two main challenges in classifying manipulation activities. First is the inherent lack of characteristic visual appearances of pen inputs that correspond to manipulation activities. Second is the necessity of real-time classification based upon the principle that users must receive immediate and appropriate visual feedback about the effects of their actions. In this paper (1) an existing activity prediction system for pen-based devices is modified for real-time activity prediction and (2) an alternative time-based activity prediction system is introduced. Both systems use eye gaze movements that naturally accompany pen-based user interaction for activity classification. The results of our comprehensive experiments demonstrate that the newly developed alternative system is a more successful candidate (in terms of prediction accuracy and early prediction speed) than the existing system for real-time activity prediction. More specifically, midway through an activity, the alternative system reaches 66% of its maximum accuracy value (i.e. 66% of 70.34%) whereas the existing system reaches only 36% of its maximum accuracy value (i.e. 36% of 55.69%).
Cartoonists and animators often use lines of action to emphasize dynamics in character poses. In this paper, we propose a physically-based model to simulate the line of action's motion, leading to rich motion from simple drawings. Our proposed method is decomposed into three steps. Based on user-provided strokes, we forward simulate 2D elastic motion. To ensure continuity across keyframes, we re-target the forward simulations to the drawn strokes. Finally, we synthesize a 3D character motion matching the dynamic line. The fact that the line can move freely like an elastic band raises new questions about its relationship to the body over time. The line may move faster and leave body parts behind, or the line may slide slowly towards other body parts for support. We conjecture that the artist seeks to maximize the filling of the line (with the character's body)---while respecting basic realism constraints such as balance. Based on these insights, we provide a method that synthesizes 3D character motion, given discontinuously constrained body parts that are specified by the user at key moments.
The interpretation of user sketches generates research interest in the product design community since the computer interpretation of sketches may reduce the design-to-market time while giving the designer greater flexibility and control of the design process. This paper describes how cues, namely shadows and table lines used to express structural form in the drawing, may be used in a line-labelling algorithm to obtain a drawing interpretation that matches some design intent. To this extent, this paper describes canonical forms of the cues from which a combined junction and cue dictionary is created and used within a genetic algorithm framework to label the drawing. This paper also describes how such cues may be identified from the sketch.
Hyperparameters are among the most crucial factors that affect the performance of machine learning algorithms. In general, there is no direct method for determining a set of satisfactory parameters, so hyperparameter search needs to be conducted each time a model is to be trained. In this work, we analyze how similar hyperparameters perform across various datasets from the sketch recognition domain. Results show that hyperparameter search space can be reduced to a subspace despite differences in characteristics of datasets.
Mosaic is a sketch-based system that simplifies and automates the creation of digital decorative mosaics from scratch. The creation of each tile piece of unique shape, color and orientation, in a complex mosaic is a tedious process. Our core contribution is two-fold: first, we present a new tile growing algorithm, that balances the shape and placement of tiles with need for uniform grout; second, we develop a suite of sketch-based tools on top of this algorithm to create and clone tiles and tile patterns along sketched paths, and color them efficiently. A user evaluation shows that our system makes the creation of mosaics fast and accessible to a broad audience.
Although many online shops allow users to search for clothes by categories or keywords, it is usually impossible to specify the details of the design. This paper presents a new technology for retrieving skirt images based on sketches. We first conducted a user study to investigate the typical features illustrated in a sketch. Then algorithms have been developed for automatically extracting those features from both the skirt images and the sketches. A prototype system has been implemented to retrieve and present the best matched skirts in real time when a user interactively sketches her imagined skirt on the canvas.
Sketching is a natural way to input chemical structures that can be used to query information from a large chemical structure database. Based on a user's incomplete sketch of a chemical structure, sketch prediction becomes a challenging problem not only due to arbitrary drawings orders among users but also similarities among chemical structure layouts. In this paper, we present a graph-based approach to handle the sketch prediction problem. We use multisets as the data representation of hand-drawn chemical structures and create an undirected graph to handle data in all multisets. This approach transforms the sketch prediction problem into a search problem to find a hamiltonian path in the corresponding sub-graph with polynomial time complexity. We introduce mixed heuristics to guide the search procedure. Through an initial experiment on a hand-drawn chemical structure dataset, we demonstrate that in comparison with a baseline method, the proposed approach improves the prediction accuracy and efficiently predicts chemical structures from only partially sketched drawings.
Constructing 3D geological models is a fundamental task in oil/gas exploration and production. A critical stage in the existing 3D geological modeling workflow is moving from a geological interpretation (usually 2D) to a 3D geological model. The construction of 3D geological models can be a cumbersome task mainly because of the models' complexity, and inconsistencies between the interpretation and modeling tasks. To narrow the gap between interpretation and modeling tasks, we propose a sketched based approach. Our main goal is to mimic how domain experts interpret geological structures and allow the creation of models directly from the interpretation task, therefore avoiding the drawbacks of a separate modeling stage. Our sketch-based modeler is based on standard annotations of 2D geological maps and on geologists' interpretation sketches. Specific geological rules and constraints are applied and evaluated during the sketch-based modeling process to guarantee the construction of a valid 3D geologic model.
In this paper we present a new method for automatically constructing 3D meshes from a single input image. With the increasing content demands of modern digital entertainment and the expectation of involvement from users, automatic artist-free systems are an important step in allowing user generated content and rapid game prototyping. Our system proposes a novel heuristic for the creation of a 3D mesh from a single piece of non-occluding 2D concept art. By extracting a skeleton structure, approximating the 3D orientation and analysing line curvature properties, appropriate centrepoints can be found around which to create the cross-sectional slices used to build a final triangle mesh. Our results show that a single 2D input image can be used to generate a rigged 3D low-polygon model suitable for use in realtime applications.