We developed the Motion-Simulation Platform, a platform running within a game engine that is able to extract both RGB imagery and the corresponding intrinsic motion data (i.e., motion field). This is useful for motion-related computer vision tasks where large amounts of intrinsic motion data are required to train a model. We describe the implementation and design details of the Motion-Simulation Platform. The platform is extendable, such that any scene developed within the game engine is able to take advantage of the motion data extraction tools. We also provide both user and AI-bot controlled navigation, enabling user-driven input and mass automation of motion data collection.
This paper presents a novel solution for estimating simulator sickness in HMDs using machine learning and 3D motion data, informed by user-labeled simulator sickness data and user analysis. We conducted a novel VR user study, which decomposed motion data and used an instant dial-based sickness scoring mechanism. We were able to emulate typical VR usage and collect user simulator sickness scores. Our user analysis shows that translation and rotation differently impact user simulator sickness in HMDs. In addition, users' demographic information and self-assessed simulator sickness susceptibility data are collected and show some indication of potential simulator sickness. Guided by the findings from the user study, we developed a novel deep learning-based solution to better estimate simulator sickness with decomposed 3D motion features and user profile information. The model was trained and tested using the 3D motion dataset with user-labeled simulator sickness and profiles collected from the user study. The results show higher estimation accuracy when using the 3D motion data compared with methods based on optical flow extracted from the recorded video, as well as improved accuracy when decomposing the motion data and incorporating user profile information.
Despite the increasing popularity of VR games, one factor hindering the industry's rapid growth is motion sickness experienced by the users. Symptoms such as fatigue and nausea severely hamper the user experience. Machine Learning methods could be used to automatically detect motion sickness in VR experiences, but generating the extensive labeled dataset needed is a challenging task. It needs either very time consuming manual labeling by human experts or modification of proprietary VR application source codes for label capturing. To overcome these challenges, we developed a novel data collection tool, VRhook, which can collect data from any VR game without needing access to its source code. This is achieved by dynamic hooking, where we can inject custom code into a game's run-time memory to record each video frame and its associated transformation matrices. Using this, we can automatically extract various useful labels such as rotation, speed, and acceleration. In addition, VRhook can blend a customized screen overlay on top of game contents to collect self-reported comfort scores. In this paper, we describe the technical development of VRhook, demonstrate its utility with an example, and describe directions for future research.
Eye typing, which utilizes eye gaze input to interact with computers, provides an indispensable means for people with severe disabilities to write, to talk, and to communicate. Despite more than two decades' of research into eye typing (which, for the most part, focused on the hardware/technical aspects associated with implementing a system), there lacks a well designed solution that has incorporated the key research findings and integrated them into a unified system. Hence, we designed and developed a novel user interface AVIN (Assisted Visual Interactive Notepad) for eye writing that expedites the writing workflow to enhance the overall user experience. Our preliminary user testing results showed that a novice user can achieve 7 wpm with an hour's typing practice, and a more experienced user can achieve 15 wpm with ten hours' practice; whereas an expert user can only reach 6-8 wpm using the standard QWERTY design.
In the present paper, we report two experiments out of series of studies designed to examine various aspects of visual search for icons of differing spatial frequencies. Specifically, the present experiments explore whether there exists a search asymmetry between high and low spatial frequency icons (A amongst B > B amongst A), and whether observers can limit their search to the relevant set of items in a display containing both types of icons. Our results show that a classic search asymmetry does not exist for spatial frequency; that, rather, both types of targets ‘pop out’; that search for a high spatial frequency target amongst high spatial frequency distractors is less efficient than search for a low spatial frequency target amongst low spatial frequency distractors; and that observers are partially able to limit their search to the relevant subset in mixed displays. Implications for the design of touch screen user interfaces are discussed.
In this paper, we report a study that examines the relationship between image-based computational analyses of web pages and users' aesthetic judgments about the same image material. Web pages were iteratively decomposed into quadrants of minimum entropy (quadtree decomposition) based on low-level image statistics, to permit a characterization of these pages in terms of their respective organizational symmetry, balance and equilibrium. These attributes were then evaluated for their correlation with human participants' subjective ratings of the same web pages on four aesthetic and affective dimensions. Several of these correlations were quite large and revealed interesting patterns in the relationship between low-level (i.e., pixel-level) image statistics and design-relevant dimensions.
Users' perceptions of an interface play an important role in influencing their acceptance of a product. Currently, however, designers lack an objective method to predict users' perceptions and to guide decisions during the design process. We propose a statistical image analysis method that takes the user interface as an input and computes image features that can be used to inform design. Based on these metrics, we then implemented a supervised learning algorithm to predict users' perceptions. We tested this method with a case study involving car infotainment systems. The results showed that the algorithm can capture users' perceptual ratings quite well, in particular for the perceptions of clutter vs. organized, and aggressive vs. calm. We discuss the implications of this computational method for UI design.
Users' perceptions of the appearance and the usability of an interactive system are two integral parts that contribute to the users' experience of the system. "Actual usability" represents a system value that is revealed either during usability testing and related methods by experts or during use by the target users. Perceived usability is an assumption about a systems' usability that has been made prior to, or independent of, its use. The appearance of a product can inadvertently affect its perceived usability; however, their relationship has not been systematically explored. We describe an approach that uses "perceptual maps" to visualize the relationship between perceived usability and subjective appearance. A group of professional designers rated representative car infotainment systems for their subjective appearance; a group of usability experts rated the same models for their perceived usability. We applied multidimensional scaling (MDS) to project the ratings into the same Euclidean space. The results show certain overlap between the perceptions of product appearance and usability. The implications of this approach for designing interactive systems are discussed.
User experience in virtual environments including presence, enjoyment, and Simulator Sickness (SS) was modeled based on the effects of field-of-view (FOV), stereopsis, visual motion frequency, interactivity, and predictability of motion orientation. We developed an instrument to assess the user experience using multivariate statistics and Item Response Theory. Results indicated that (1) presence was increased with a large FOV, stereo display, visual motion in low frequency ranges (.03 Hz), and high levels of interactivity; (2) more SS was reported with increasing FOV, stereo display, .05-.08 Hz visual motion frequency, lack of interactivity and predictability to visual motion; (3) enjoyment was increased with visual motion in low frequency ranges (.03 Hz) and high levels of interactivity. The resulting response surface model visualizes the complex relationships between presence, enjoyment, and SS. Overall, increasing interactivity was found to be the most profound way to enhance user experience in virtual environments.
A sedentary lifestyle is a contributing factor to chronic diseases, and it is often correlated with obesity. To promote an increase in physical activity, we created a social computer game, Fish'n'Steps, which links a player's daily foot step count to the growth and activity of an animated virtual character, a fish in a fish tank. As further encouragement, some of the players' fish tanks included other players' fish, thereby creating an environment of both cooperation and competition. In a fourteen-week study with nineteen participants, the game served as a catalyst for promoting exercise and for improving game players' attitudes towards physical activity. Furthermore, although most player's enthusiasm in the game decreased after the game's first two weeks, analyzing the results using Prochaska's Transtheoretical Model of Behavioral Change suggests that individuals had, by that time, established new routines that led to healthier patterns of physical activity in their daily lives. Lessons learned from this study underscore the value of such games to encourage rather than provide negative reinforcement, especially when individuals are not meeting their own expectations, to foster long-term behavioral change.
Two studies used virtual environments (VEs) to examine the role of simulated motion frequency in visually-induced motion and simulator sickness (SS). Predictions were derived from the “crossover” hypothesis, which suggests that restricting simulated motion frequency in the region where conflicting motion cues to the visual and vestibular self-motion systems are readily detected (around 0.07 Hz) should reduce SS. Results from both studies were consistent with predictions.
This study developed a new procedure, a Virtual Guiding Avatar (VGA), which combined self-motion prediction cues and an independent visual background (IVB) to alleviate simulator sickness (SS). The VGA, which was embodied as an abstract airplane, was designed to lead the participant along a horizontal motion trajectory through a virtual environment. Both motion prediction cues and IVBs, which provide an earth-fixed reference frame, reduced SS in separate previous studies. Participants were exposed to complex visual motion through a cartoon-like simulated environment in a very wide field of view driving simulator. Participants' responses to avatars with varying motion properties - fixed, rotation only or rotation plus translation - were assessed using a within-subjects experimental design. Results indicated that SS was reduced by a VGA that presented rotational cues alone or rotation plus translation. The VGA also increased participants' sense of presence and enjoyment relative to conditions lacking a VGA. The VGA procedure can be used to enhance user experiences in immersive virtual environments as well as to improve motion simulator design.
To address the problem of information overload in today's world, we have developed START, a natural language question answering system that provides users with high-precision information access through the use of natural language annotations. To address the difficulty of accessing large amounts of heterogeneous structured and semistructured data, we have developed Omnibase, which assists START by integrating Web databases into a single, uniformly structured "virtual database." To address the sheer amount of unstructured information available electronically, we have developed techniques for distilling large amounts of free text into relations that capture the salient aspects of the text. The combination of natural language annotation technology, object-property-value data model, and relation extraction technology allows us to rapidly develop and deploy smart applications for knowledge intensive domains. Our ultimate goal is to develop a computer system that acts like a "smart reference librarian," providing users with "just the right information" in response to questions posed in natural language.
Traditional information retrieval systems based on the "bag-of-words" paradigm cannot completely capture the semantic content of documents. Yet it is impossible with current technology to build a practical information access system that fully analyzes and understands unrestricted natural language. However, if we avoid the most complex and processing-intensive natural language understanding techniques, we can construct a large-scale information access system which is capable of processing...