Correctly estimating the surrounding illumination is essential for creating visually coherent Mixed Reality (MR) experiences. The most accurate results can be achieved by utilizing a light probe, a dedicated object with known reflectance parameters that is placed into the scene. However, the need for a dedicated object placed in the area where the illumination is estimated presents a severe limitation. Building on the increasing popularity of gestural interaction in MR, we present HandLight, an approach to estimating the illumination from the user's hands during interaction. Contrary to static light probes, HandLight does not require preparation of the environment and generates an atlas of light probes while the user moves in the world, thus reflecting variable illumination. Our system utilizes a neural network that learns the environment lighting from images of the hand. We train the network on a dataset depicting three common gestures (pinch, fist, bloom) under varying light conditions. We show that our approach can provide believable illumination estimations for a variety of illuminations on a dataset of real hand images.
Extended Reality (XR) is increasingly used as a productivity tool and recent commercial XR devices have even been specifically designed as productivity tools, or, at least, are heavily advertised for such purposes, such as the Apple Vision Pro (AVP), which has now been available for more than one year. In spite of what marketing suggests, research still lacks an understanding of the long-term usage of such devices in ecologically valid everyday settings, as most studies are conducted in very controlled environments. Therefore, we conducted interviews with ten AVP users to better understand how experienced users engage with the device, and which limitations persist. Our participants report that XR can increase productivity and that they got used to the device after some time. Yet, a range of limitations persist that might hinder the widespread use of XR as a productivity tool, such as a lack of native applications, difficulties when integrating XR into current workflows, and limited possibilities to adapt and customize the XR experience.
Autonomous robot operations in the industry are becoming increasingly complex. It is therefore a significant challenge to comprehend the fundamental processes and to gain an understanding of the status of these systems. The RoboCup Logistics League (RCLL) represents a small smart factory environment with several workstations and operating robots. Despite its small scale the processes that occur within the league are very complex. Even with live commentary, observers have difficulties to follow the processes and game progress. This results in low interest in the RCLL and only few of visitors at competitions. To address this, we want to present RCLL-AR, an augmented reality (AR) solution visualizing highly relevant information of the RoboCup Logistics League. By using RCLL-AR, spectators of the game can see the current progress of the game, receive additional information about different workstations and understand future robot movements. To gain insights into the benefits of RCLL-AR for different stakeholders, we conducted expert interviews, a novice user study and an HMD study. Our findings showcase challenges AR faces in complex autonomous systems but also indicate benefits for novices and experts.
Placing a transparent liquid crystal display (LCD) into the light path is a simple approach to create occlusion-capable optical seethrough head-mounted displays (OST-HMDs) that suffers from defocused (soft-edge) occlusion where the mask leakage partially occludes surrounding content as well. Creating a focused (hard-edge) occlusion that does not suffer from mask leakage requires complicated, bulky optical setups. We present X-Mask, a pinhole-arraybased OST-HMD that creates a sharp occlusion mask without the need for a bulky setup requiring only two transparent LCD layers. By rendering a pinhole array on the layer closer to the user's eye, our system functions as a programmable aperture layer that extends the effective depth of field and improves the sharpness of the occlusion mask rendered on the second LCD layer. Utilizing a conventional circular pinhole would result in non-uniform brightness and contrast. By changing the pinhole shape to a cross enables nearoptimal retinal tiling with reduced overlaps and gaps. To accommodate pupil size variation, focus distance, and gaze direction, our system design allows for gaze-contingent adjustment of both LCD layers. We validate X-Mask in simulations and a physical prototype showing improved occlusion sharpness and visual uniformity.
Systems with occlusion capabilities, such as those used in vision augmentation, image processing, and optical see-through head-mounted display (OST-HMD), have gained popularity. Achieving precise (hard-edge) occlusion in these systems is challenging, often requiring complex optical designs and bulky volumes. On the other hand, utilizing a single transparent liquid crystal display (LCD) is a simple approach to create occlusion masks. However, the generated mask will appear defocused (soft-edge) resulting in insufficient blocking or occlusion leakage. In our work, we delve into the perception of soft-edge occlusion by the human visual system and present a preference-based optimal expansion method that minimizes perceived occlusion leakage. In a user study involving 20 participants, we made a noteworthy observation that the human eye perceives a sharper edge blur of the occlusion mask when individuals see through it and gaze at a far distance, in contrast to the camera system's observation. Moreover, our study revealed significant individual differences in the perception of soft-edge masks in human vision when focusing. These differences may lead to varying degrees of demand for mask size among individuals. Our evaluation demonstrates that our method successfully accounts for individual differences and achieves optimal masking effects at arbitrary distances and pupil sizes.
Automatic Augmented Reality (AR)-based work support systems seamlessly display the next step instruction when the current task step completion is detected, all without requiring explicit user interaction. However, whenever the progress estimation fails it can lead to confusion and loss of trust in the system. To support the user’s understanding of the reliability of the displayed instructions, we display the confidence of the current step. To explore the effects the presented confidence has on user interaction with the system, we conducted a Wizard-of-Oz style user study with 12 participants. Participants had to assemble a LEGO model with and without confidence visualization, where the confidence for each step was predefined. Our results showed that while the visualization did not increase the user’s performance, participants preferred having the confidence displayed to them, especially when it was low. Our findings provide insights into the design of enhancing the user’s perception of uncertainty in current step detection in automated instruction systems.
In an optical see-through augmented reality system, virtual and real information can be viewed at different distances. This requires users to frequently change their eye focus from one distance to another, and only one piece of information is in sharp focus while the other is out of focus. Previous studies have found that due to out-of-focus virtual information, when integrating information between the distances, users suffer fatigue and miss important information. Therefore, this paper introduces a novel font, termed a SharpView font, which looks sharper and more legible than standard fonts when seen out of focus. Our method models out-of-focus blur with Zernike polynomials and coefficients, develops a focus correction algorithm based on constrained total variation optimization, and proposes a novel gradient-based algorithm to quantify the sharpness of textual information. We have evaluated the SharpView font through simulation and optically viewed camera-based measurement. When seen out of focus, our proposed font are significantly sharper than standard fonts, as assessed both visually and quantitatively through simulation (40%-44%), as well as the optics of an augmented reality display (24%-32%).
Providing attention guidance, such as assisting in search tasks, is a prominent use for Augmented Reality. Typically, this is achieved by graphically overlaying geometrical shapes such as arrows. However, providing visual guidance can cause side effects such as attention tunnelling or scene occlusions, and introduce additional visual clutter. Alternatively, visual guidance can adjust saliency but this comes with different challenges such as hardware requirements and environment dependent parameters. In this work we advocate for using flicker as an alternative for real-world guidance using Augmented Reality. We provide evidence for the effectiveness of flicker from two user studies. The first compared flicker against alternative approaches in a highly controlled setting, demonstrating efficacy (N = 28). The second investigated flicker in a practical task, demonstrating feasibility with higher ecological validity (N = 20). Finally, our discussion highlights the opportunities and challenges when using flicker to provide real-world visual guidance using Augmented Reality.
The vergence-accommodation conflict (VAC) presents a major perceptual challenge for head-mounted displays with a fixed image plane. Varifocal and layered display designs can mitigate the VAC. However, the image quality of varifocal displays is affected by imprecise eye tracking, whereas layered displays suffer from reduced image contrast as the distance between layers increases. Combined designs support a larger workspace and tolerate some eye-tracking error. However, any layered design with a fixed layer spacing restricts the amount of error compensation and limits the in-focus contrast. We extend previous hybrid designs by introducing confidence-driven volume control, which adjusts the size of the view volume at runtime. We use the eye tracker's confidence to control the spacing of display layers and optimize the trade-off between the display's view volume and the amount of eye tracking error the display can compensate. In the case of high-quality focus point estimation, our approach provides high in-focus contrast, whereas low-quality eye tracking increases the view volume to tolerate the error. We describe our design, present its implementation as an optical-see head-mounted display using a multiplicative layer combination, and present an evaluation comparing our design with previous approaches.
With innovations in the field of gaze and eye tracking, a new concentration of research in the area of gaze-tracked systems and user interfaces has formed in the field of Extended Reality (XR). Eye trackers are being used to explore novel forms of spatial human–computer interaction, to understand human attention and behavior, and to test expectations and human responses. In this article, we review gaze interaction and eye tracking research related to XR that has been published since 1985, which includes a total of 215 publications. We outline efforts to apply eye gaze for direct interaction with virtual content and design of attentive interfaces that adapt the presented content based on eye gaze behavior and discuss how eye gaze has been utilized to improve collaboration in XR. We outline trends and novel directions and discuss representative high-impact papers in detail.
This paper presents guitARhero, an Augmented Reality application for interactively teaching guitar playing to beginners through responsive visualizations overlaid on the guitar neck. We support two types of visual guidance, a highlighting of the frets that need to be pressed and a 3D hand overlay, as well as two display scenarios, one using a desktop magic mirror and one using a video see-through head-mounted display. We conducted a user study with 20 participants to evaluate how well users could follow instructions presented with different guidance and display combinations and compare these to a baseline where users had to follow video instructions. Our study highlights the trade-off between the provided information and visual clarity affecting the user's ability to interpret and follow instructions for fine-grained tasks. We show that the perceived usefulness of instruction integration into an HMD view highly depends on the hardware capabilities and instruction details.
Augmented Reality has traditionally been used to display digital overlays in real environments. Many AR applications such as remote collaboration, picking tasks, or navigation require highlighting physical objects for selection or guidance. These highlights use graphical cues such as outlines and arrows. Whilst effective, they greatly contribute to visual clutter, possibly occlude scene elements, and can be problematic for long-term use. Substituting those overlays, we explore saliency modulation to accentuate objects in the real environment to guide the user’s gaze. Instead of manipulating video streams, like done in perception and cognition research, we investigate saliency modulation of the real world using optical-see-through head-mounted displays. This is a new challenge, since we do not have full control over the view of the real environment. In this work we provide our specific solution to this challenge, including built prototypes and their evaluation.
Accurate camera localization is an essential part of tracking systems. However, localization results are greatly affected by illumination. Including data collected under various lighting conditions can improve the robustness of the localization algorithm to lighting variation but it is time consuming. Synthetic images are easy to accumulate and we can control varying illumination. However, synthetic images do not perfectly match real images of the same scene, i.e., there exists a gap between real and synthetic images that also affects the accuracy of camera localization. To reduce the impact of this gap, we introduce ''real-to-synthetic feature transform (REST)." REST is a fully connected neural network that converts real features to their synthetic counterpart. The converted features can then be matched against the accumulated database for robust camera localization. Our experimental results show that REST improves matching accuracy by approximately 28% compared to a naïve method.
Triangle meshes are used in many important shape-related applications including geometric modeling, animation production, system simulation, and visualization. However, these meshes are typically generated in raw form with several defects and poor-quality elements, obstructing them from practical application. Over the past decades, different surface remeshing techniques have been presented to improve these poor-quality meshes prior to the downstream utilization. A typical surface remeshing algorithm converts an input mesh into a higher quality mesh with consideration of given quality requirements as well as an acceptable approximation to the input mesh. In recent years, surface remeshing has gained significant attention from researchers and engineers, and several remeshing algorithms have been proposed. However, there has been no survey article on remeshing methods in general with a defined search strategy and article selection mechanism covering the recent approaches in surface remeshing domain with a good connection to classical approaches. In this article, we present a survey on surface remeshing techniques, classifying all collected articles in different categories and analyzing specific methods with their advantages, disadvantages, and possible future improvements. Following the systematic literature review methodology, we define step-by-step guidelines throughout the review process, including search strategy, literature inclusion/exclusion criteria, article quality assessment, and data extraction. With the aim of literature collection and classification based on data extraction, we summarized collected articles, considering the key remeshing objectives, the way the mesh quality is defined and improved, and the way their techniques are compared with other previous methods. Remeshing objectives are described by angle range control, feature preservation, error control, valence optimization, and remeshing compatibility. The metrics used in the literature for the evaluation of surface remeshing algorithms are discussed. Meshing techniques are compared with other related methods via a comprehensive table with indices of the method name, the remeshing challenge met and solved, the category the method belongs to, and the year of publication. We expect this survey to be a practical reference for surface remeshing in terms of literature classification, method analysis, and future prospects.
In optical see-through augmented reality (AR), information is often distributed between real and virtual contexts, and often appears at different distances from the user. To integrate information, users must repeatedly switch context and change focal distance. If the user's task is conducted under time pressure, they may attempt to integrate information while their eye is still changing focal distance, a phenomenon we term transient focal blur. Previously, Gabbard, Mehra, and Swan (2018) examined these issues, using a text-based visual search task on a one-eye optical see-through AR display. This paper reports an experiment that partially replicates and extends this task on a custom-built AR Haploscope. The experiment examined the effects of context switching, focal switching distance, binocular and monocular viewing, and transient focal blur on task performance and eye fatigue. Context switching increased eye fatigue but did not decrease performance. Increasing focal switching distance increased eye fatigue and decreased performance. Monocular viewing also increased eye fatigue and decreased performance. The transient focal blur effect resulted in additional performance decrements, and is an addition to knowledge about AR user interface design issues.
Colour vision deficiency is a common visual impairment that cannot be compensated for using optical lenses in traditional glasses, and currently remains untreatable. In our work, we report on research on Computational Glasses for compensating colour vision deficiency. While existing research only showed corrected images within the periphery or as an indirect aid, Computational Glasses build on modified standard optical see-through head-mounted displays and directly modulate the user’s vision, consequently adapting their perception of colours. In this work, we present an exhaustive literature review of colour vision deficiency compensation and subsequent findings; several prototypes with varying advantages—from well-controlled bench prototypes to less controlled but higher application portable prototypes; and a series of studies evaluating our approach starting with proving its efficacy, comparing to the state-of-the-art, and extending beyond static lab prototypes looking at real world applicability. Finally, we evaluated directions for future compensation methods for computational glasses.
Physical objects are usually not designed with interaction capabilities to control digital content. Nevertheless, they provide an untapped source for interactions since every object could be used to control our digital lives. We call this the missing interface problem: Instead of embedding computational capacity into objects, we can simply detect users' gestures on them. However, gesture detection on such unmodified objects has to date been limited in the spatial resolution and detection fidelity. To address this gap, we conducted research on micro-gesture detection on physical objects based on Google Soli's radar sensor. We introduced two novel deep learning architectures to process range Doppler images, namely a three-dimensional convolutional neural network (Conv3D) and a spectrogram-based ConvNet. The results show that our architectures enable robust on-object gesture detection, achieving an accuracy of approximately 94% for a five-gesture set, surpassing previous state-of-the-art performance results by up to 39%. We also showed that the decibel (dB) Doppler range setting has a significant effect on system performance, as accuracy can vary up to 20% across the dB range. As a result, we provide guidelines on how to best calibrate the radar sensor.
Adding virtual information that is indistinguishable from reality has been a long-awaited goal in Augmented Reality (AR). While already demonstrated in the 1960s, only recently have Optical See-Through Head-Mounted Displays (OST-HMDs) seen a reemergence, partially thanks to large investments from industry, and are now considered to be the ultimate hardware for augmenting our visual perception. In this article, we provide a thorough review of state-of-the-art OST-HMD-related techniques that are relevant to realize the aim of an AR interface almost indistinguishable from reality. In this work, we have an initial look at human perception to define requirements and goals for implementing such an interface. We follow up by identifying three key challenges for building an OST-HMD-based AR interface that is indistinguishable from reality: spatial realism, temporal realism, and visual realism. We discuss existing works that aim to overcome these challenges while also reflecting against the goal set by human perception. Finally, we give an outlook into promising research directions and expectations for the years to come.
Haruo Takemura合作论文数Osaka University;Infomedia Education Division;Cybermedia Center14