While touchscreen input is the de facto standard for information access on modern smartwatches, issues like accuracy and occlusion have encouraged researchers to explore various off-screen forms of input, including on-body, in-air, and hand-pose interactions. However, past research has typically explored these alternative input spaces in isolation one from the other. Additionally, it has received no attention to design the space leveraging multi-finger interaction considering accidental triggering of unwanted interaction. In this paper, we introduce Watch+AD, an explicit form of smartwatch interaction that combines touchscreen input with the around-device inputs on a smartwatch for disambiguing interactions. We present the parameters which can be used to create Watch+AD gestures, then explore the comfort and reachability of AD space when a finger is encumbered by the touchscreen, and present an early probe of command mappings to drive a set of applications that leverage results from ergonomics to guide design of Watch+AD applications. Together, our results support combined on-screen and arounddevice smartwatch input as a way to expand smartwatches' gestural input vocabulary.
Given the computing power of mobile devices, porting feature-rich applications to these devices is increasingly feasible. However, feature-rich applications include large command sets, and providing access to these commands through screen-based widgets results in issues of occlusion and layering. To address this issue, we introduce Ether-Mark, a hierarchical, gesture-based, marking menu inspired, around-device menu for mobile devices enabling both on- and near-device interaction. We investigate the design of such menus and their learnability through three experiments. We first design and contrast three variants of Ether-Mark, yielding a zigzag menu design. We then refine input accuracy via a deformation model of the menu. And, we evaluate the learnability of the menus and the accuracy of the deformation model, revealing an accuracy rate up to 98.28%. We finally, compare in-air Ether-Mark with marking menus.Our results argue for Ether-Mark as a promising effective mechanism to leverage proximal around-device space.
We compare the performance and level of distraction of expressive directional gesture input in the context of in-vehicle system commands. Center console touchscreen swipes and midair swipe-like movements are tested in 8-directions, with 8-button touchscreen tapping as a baseline. Participants use these input methods for intermittent target selections while performing the Lane Change Task in a virtual driving simulator. Input performance is measured with time and accuracy, cognitive load with deviation of lane position and speed, and distraction from frequency of off-screen glances. Results show midair gestures were less distracting and faster, but with lower accuracy. Touchscreen swipes and touchscreen tapping are comparable across measures. Our work provides empirical evidence for vehicle interface designers and manufacturers considering midair or touch directional gestures for centre console input.
Extracting underlying data from rasterized charts is tedious and inaccurate; values might be partially occluded or hard to distinguish, and the quality of the image limits the precision of the data being recovered. To address these issues, we introduce a semi-automatic system leveraging vector charts to extract the underlying data easily and accurately. The system is designed to make the most of vector information by relying on a drag-and-drop interface combined with selection, filtering, and previsualization features. A user study showed that participants spent less than 4 minutes to accurately recover data from charts published at CHI with diverse styles, thousands of data points, a combination of different encodings, and elements partially or completely occluded. Compared to other approaches relying on raster images, our tool successfully recovered all data, even when hidden, with a 78% lower relative error.
Recent computer vision advances can enable television control using midair barehand input, for which selecting targets using hand positions is a primary task. Yet, the size and position of a preferred 2D interaction region has not been specifically investigated for this setting. We report on a field study with people in front of their own television in their own home. Controlled variations of target position stimuli were presented while a camera records the natural hand position used by the participant in response. Based on hand and face landmarks, density plots define a preferred input region location and size using a human-scaled unit of face-widths. Distribution of hand positions relative to target stimuli reveals consistency and precision, suggesting an ability for users to map their input space to display space. We elicit an ideal input region from participants and provide a range of dimensions for designers implementing barehand input techniques.
Emerging input techniques that rely on sensing and recognition can misinterpret a user's intention, resulting in errors and, potentially, a negative user experience. To enhance the development of such input techniques, it is valuable to understand implications of these errors, but they can very costly to simulate. Through two controlled experiments, this work explores various low-cost methods for evaluating error acceptability of freehand mid-air gestural input in virtual reality. Using a gesture-driven game and a drawing application, the first experiment elicited error characteristics through text descriptions, video demonstrations, and a touchscreen-based interactive simulation. The results revealed that video effectively conveyed the dynamics of errors, whereas the interactive modalities effectively reproduced the user experience of effort and frustration. The second experiment contrasts the interactive touchscreen simulation with the target modality - a full VR simulation - and highlights the relative costs and benefits for assessment in an alternative, but still interactive, modality. These findings introduce a spectrum of low-cost methods for evaluating recognition-based errors in VR and a series of characteristics that can be understood in each.
Pointing is an elementary interaction in virtual and augmented reality environments, and, to effectively support selection, techniques must deal with the challenges of occlusion and depth specification. Most of the previous techniques require two explicit steps to handle occlusion. In this paper, we propose Conductor, an intuitive, plane-ray, intersection-based, 3D pointing technique where users leverage bimanual input to control a ray and intersecting plane. Conductor allows users to use the non-dominant hand to adjust the cursor distance on the ray while pointing with the dominant hand. We evaluate Conductor against Raycursor, a state-of-the-art VR pointing technique, and show that Conductor outperforms Raycursor for selection tasks. Given our results, we argue that bimanual selection techniques merit additional exploration to support object selection and placement within virtual environments.
While many smartphones now include a waterfall (or curved edge) display, this edge is rarely exploited for input. In this paper, we introduce EdgeMark, a thumb-operated, one-handed, gestural menu to provide rapid and subtle access to commands. In a first study, we assess the movement range of the user during uni-manual input to define a spatial range for our EdgeMark menu. In a second study, we compare three potential EdgeMark designs to a marking menu variant to assess accuracy and performance. Our findings indicate that EdgeMark menus can serve as a reliable, effective mechanism for subtle, one-hand command invocation in modern, waterfall-edged smartphones.
One common task when controlling smart displays is the manipulation of video timelines. Given current examples of smart displays that support distant bare hand control, in this paper we explore CD Gain functions to support both seeking and scrubbing tasks. Through a series of experiments, we demonstrate that a linear CD Gain function provides performance advantages when compared to either a constant function or generalised logistic function (GLF). In particular, linear gain is faster than a GLF and has lower error rate than Constant gain. Furthermore, Linear and GLF gains’ average temporal error when targeting a one second interval on a two hour timeline (±5 s) is less than one third the error of a Constant gain.
Alongside vision and sound, hardware systems can be readily designed to support various forms of tactile feedback; however, while a significant body of work has explored enriching visual and auditory communication with interactive systems, tactile information has not received the same level of attention. In this work, we explore increasing the expressivity of tactile feedback by allowing the user to dynamically select between several channels of tactile feedback using variations in finger speed. In a controlled experiment, we show that a user can learn the dynamics of eyes-free tactile channel selection among different channels, and can reliable discriminate between different tactile patterns during multi-channel selection with an accuracy up to 90% when using two finger speed levels. We discuss the implications of this work for richer, more interactive tactile interfaces.
Virtual hand or pointer metaphors are among the key approaches for target selection in immersive environments. However, targeting moving objects is complicated by factors including target speed, direction, and depth, such that a basic implementation of these techniques might fail to optimize user performance. We present results of two empirical studies comparing characteristics of virtual hand and pointer metaphors for moving target acquisition. Through a first study, we examine the impact of depth on users' performance when targets move beyond and within arms' reach. We find that movement in depth has a great impact on both metaphors. In a follow-up study, we design a reach-bounded Go-Go (rbGo-Go) technique to address challenges of virtual hand and compare it to Ray-Casting. We find that target width and speed are significant determinants of user performance and we highlight the pros and cons for each of the techniques in the given context. Our results inform the UI design for immersive selection of moving targets.
One common task when controlling smart displays is the manipulation of menu items. Given current examples of smart displays that support distant bare hand control, in this paper we explore menu item selection tasks with three different mappings of barehand movement to target selection. Through a series of experiments, we demonstrate that Positional mapping is faster than other mappings when the target is visible but requires many clutches in large targeting spaces. Rate-based mapping is, in contrast, preferred by participants due to its perceived lower effort, despite being slightly harder to learn initially. Tradeoffs in the design of target selection in smart tv displays are discussed.
Target acquisition in an occluded environment is challenging given the omni-directional and first-person view in virtual reality (VR). We propose Solar-Casting, a global scene filtering technique to manage occlusion in VR. To improve target search, users control a reference sphere centered at their head through varied occlusion management modes: Hide, SemiT (Semi-Transparent), Rotate. In a preliminary study, we find SemiT to be better suited for understanding the context without sacrificing performance by applying semi-transparency to targets within the controlled sphere. We then compare Solar-Casting to highly efficient selection techniques to acquire targets in a dense and occluded VR environment. We find that Solar-Casting performs competitively to other techniques in known environments, where the target location information is revealed. However, in unknown environments, requiring target search, Solar-Casting outperforms existing approaches. We conclude with scenarios demonstrating how Solar-Casting can be applied to crowded and occluded environments in VR applications.
This paper presents a qualitative multi-phase study seeking to identify patterns in users' anthropomorphized perceptions of conversational agents. Through a comparative analysis of behavioral perceptions and visual conceptions of three agents - Alexa, Google Assistant, and Siri - we first show that the perceptions of an agent's character are structured according to five categories: approachability, sentiment toward a user, professionalism, intelligence, and individuality. We then explore visualizations of the agents' appearance and discuss the specifics assigned to each agent. Finally, we analyze associative explanations for these perceptions. We demonstrate that the anthropomorphized behavioral and visual perceptions of agents yield structural consistency and discuss how these perceptions are linked with each other and system features.
In exploratory search tasks, alongside information retrieval, information representation is an important factor in sensemaking. In this paper, we explore a multi-layer extension to knowledge graphs, hierarchical knowledge graphs (HKGs), that combines hierarchical and network visualizations into a unified data representation asa tool to support exploratory search. We describe our algorithm to construct these visualizations, analyze interaction logs to quantitatively demonstrate performance parity with networks and performance advantages over hierarchies, and synthesize data from interaction logs, interviews, and thinkalouds on a testbed data set to demonstrate the utility of the unified hierarchy+network structure in our HKGs. Alongside the above study, we perform an additional mixed methods analysis of the effect of precision and recall on the performance of hierarchical knowledge graphs for two different exploratory search tasks. While the quantitative data shows a limited effect of precision and recall on user performance and user effort, qualitative data combined with post-hoc statistical analysis provides evidence that the type of exploratory search task (e.g., learning versus investigating) can be impacted by precision and recall. Furthermore, our qualitative analyses find that users are unable to perceive differences in the quality of extracted information. We discuss the implications of our results and analyze other factors that more significantly impact exploratory search performance in our experimental tasks.
Documents such as presentations, instruction manuals, and research papers are disseminated using various file formats, many of which barely support the incorporation of interactive content. To address this lack of interactivity, we present Chameleon, a system-wide tool that combines computer vision algorithms used for image identification with an open database format to allow for the layering of dynamic content. Using Chameleon, static documents can be easily upgraded by layering user-generated interactive content on top of static images, all while preserving the original static document format and without modifying existing applications. We describe the development of Chameleon, including the design and evaluation of vision-based image replacement algorithms, the new document-creation pipeline as well as a user study evaluating Chameleon.
Writing is a complex non-linear process that begins with a mental model of intent, and progresses through an outline of ideas, to words on paper (and their subsequent refinement). Despite past research in understanding writing, Web-scale consumer and enterprise collaborative digital writing environments are yet to greatly benefit from intelligent systems that understand the stages of document evolution, providing opportune assistance based on authors' situated actions and context. In this paper, we present three studies that explore temporal stages of document authoring. We first survey information workers at a large technology company about their writing habits and preferences, concluding that writers do in fact conceptually progress through several distinct phases while authoring documents. We also explore, qualitatively, how writing stages are linked to document lifespan. We supplement these qualitative findings with an analysis of the longitudinal user interaction logs of a popular digital writing platform over several million documents. Finally, as a first step towards facilitating an intelligent digital writing assistant, we conduct a preliminary investigation into the utility of user interaction log data for predicting the temporal stage of a document. Our results support the benefit of tools tailored to writing stages, identify primary tasks associated with these stages, and show that it is possible to predict stages from anonymous interaction logs. Together, these results argue for the benefit and feasibility of more tailored digital writing assistance.
Delayed display of menu items is a core design component of marking menus, arguably to prevent visual distraction and foster the use of mark mode. We investigate these assumptions, by contrasting the original marking menu design with immediately-displayed marking menus. In three controlled experiments, we fail to reveal obvious and systematic performance or usability advantages to using delay and mark mode. Only in very constrained settings – after significant training and only two items to learn – did traditional marking menus show a time improvement of about 260~ms. Otherwise, we found an overall decrease in performance with delay, whether participants exhibited practiced or unpracticed behaviour. Our final study failed to demonstrate that an immediately-displayed menu interface is more visually disrupting than a delayed menu. These findings inform the costs and benefits of incorporating delay in marking menus, and motivate guidelines for situations in which its use is desirable.
Eric Saund合作论文数Palo Alto Research Center4