
In this paper, we propose a vision-based haptic sensor (VHS) capable of acquiring force at the required speed for softness discrimination during palpation, along with a force measurement algorithm. Palpation requires the simultaneous acquisition of surface imagery from the affected area and haptic information, such as softness and surface texture. Additionally, the sensor must exhibit softness comparable to human skin to avoid causing discomfort to the patient. By designing a force sensor that tracks markers embedded in transparent gel using a camera, we enable the concurrent capture of visual and haptic data. An algorithm is also presented for calculating normal forces based on the extension of the markers' image plane. Accurate force modeling was achieved by training a normal force estimation model using an asymmetric stiffness coefficient matrix, which effectively mitigates cross-talk effects. Furthermore, the process was optimized by employing sparse search techniques with narrow marker search ranges between frames during high-speed imaging, enabling rapid detection of circular force markers and achieving force acquisition at 601.25 Hz. Compared to previous methods, the proposed approach offers higher measurement accuracy and speed within the force range required for palpation. It can measure at 500 Hz or higher, which is crucial for discriminating the five levels of softness important in dermatological palpation. Therefore, the proposed haptic sensor shows promise for use in robotic palpation.
Users can interact in virtual reality (VR) spaces through avatars that differ markedly from their real-world looks. These avatars can be customized to any appearance and size, whether they are based on real entities or are entirely fictitious. These avatars include non-humanoid avatars as well. Some non-humanoid avatars do not have hands, in which case the problem arises that they cannot reference using gesture. In this case, the interlocutor must determine the object from the direction of the referent's gaze and the context. Given the impact of avatar characteristics on the visual communication process of joint attention among users, it is essential to elucidate the connection between avatar traits and the range of reference to facilitate smooth interaction. In this study, the influence of avatar looks on the referential range of demonstrative indicators was elucidated. Experiments were conducted in a VR spaces using avatars of different appearances and sizes, with the aim of understanding how these differences impact the ability to refer to objects using both distal and proximal indicators. Specifically, the study aimed to identify the transition point from the proximal to the distal referential field for each type of avatar. This research seeks to deepen the understanding of how avatars, as proxies for humans in VR spaces, influence communication dynamics. Looking forward, it is anticipated that the findings will enhance the VR experience by improving referential communication among avatars of diverse appearances and sizes. This enhancement is expected to foster richer user interactions, thereby contributing to the future growth of the VR market.
This paper proposes an eye-tracking system using a CNN-LSTM network that utilizes only event data. This method holds potential for future applications in a wide range of fields, including AR/VR headsets, healthcare, and sports. Compared to traditional frame-based camera methods, our proposed approach achieves high FPS and low power consumption by utilizing event cameras. To improve the estimation accuracy, our gaze estimation system incorporates a blink detection, which was absent in existing systems. Our results shows that our method achieves better performance compared to existing studies.
This paper introduces our investigation into driving behaviors during emergency operations, such as entering an intersection on a red traffic light, to compare and analyze behavior differences based on the driver's emergency driving experience and skills. As a preliminary step, we developed a VR-based simulator that replicates emergency driving scenarios, in consultation with firefighters. Drivers with varied levels of experience and skill in emergency driving executed emergency driving maneuvers using this VR simulator, during which the system recorded their behaviors, including eye movement, head orientation, throttle and brake positions, and steering angle. Our analysis revealed distinct behavioral differences based on experience and skill; for example, professional firefighters repeatedly decelerated, briefly stopped, and checked both left and right sides before entering the intersection.
Discrepancies between an avatar's movements in virtual space and participants' movements in the real world can degrade the quality of the virtual reality (VR) experience. One prominent form of such a discrepancy is delay. Many previous studies have investigated the acceptable delay between head-tracking and landscape rendering, or the delay of the seen user's hand movements. However, the minimum detectable delay during full-body movements, particularly those involving significant changes in viewpoint, has not yet been fully investigated. In this study, the detection threshold for delays between participants' real-world movements, where the head and viewpoint positions move substantially, and corresponding avatar movements in virtual space was investigated. In the experiment, participants wearing VR goggles performed stand-up motions. Corresponding stand-up motion of the avatar in the VR space involved the delay of up to 300 ms. Participants looked at the avatar's movements through the mirror placed in front of him/her in the VR space. The detection thresholds of five individuals were investigated using psychophysical method of constant stimuli. In the experiment, the participants answered whether the avatar's movements delayed or did not delay comparing with their own movements. The mean detection threshold, at which the participant reports the presence of delay for 50% of all the time, was found to be 129.70 ms, with a 95% confidence interval of 31.59 ms. These findings provide some insights for designers of VR applications.
Transition methods that seamlessly connect real environments (REs) and virtual environments (VEs) using head-mounted displays are known to enhance user experiences, particularly the sense of presence. However, transitions relying solely on visual cues often fall short in making the VE feel convincingly real. To address this limitation, we developed a multi-modal transition method that integrates a physical door, combining tactile (e.g., turning a doorknob), and auditory (e.g., hearing a squeaky sound) stimuli with video see-through augmented reality. This approach seamlessly bridges an RE and a VE, offering a richer, more immersive experience. To validate the effectiveness of our method, we constructed a VE allowing users to move between a real office environment and a forest VE. We hypothesized that our multi-modal transition would lead to a greater sense of self-experience, presence, relaxation, and a higher physical movement level than traditional transition methods like portal and fade methods. Our results demonstrated that the total IPQ (Igroup Presence Questionnaire) scores for the proposed method and the portal were significantly higher than those for the fade method. Moreover, users exhibited significantly greater travel distance and speed with our method compared to the fade transition. These findings suggest that our transition method enhances the sense of self-experience and presence and also encourages more physical movement than the portal and fade methods. This study contributes to the understanding of how multi-modal transition methods can effectively enhance user experiences in a VE and create more immersive virtual environments.
Traditional techniques for rendering continuous surfaces from dynamic, noisy point clouds using multi-camera setups often suffer from disruptive artifacts in overlapping areas, similar to z-fighting. We introduce BlendPCR, an advanced rendering technique that effectively addresses these artifacts through a dual approach of point cloud processing and screen space blending. Additionally, we present a UV coordinate encoding scheme to enable high-resolution texture mapping via standard camera SDKs. We demonstrate that our approach offers superior visual rendering quality over traditional splat and mesh-based methods and exhibits no artifacts in those overlapping areas, which still occur in leading-edge NeRF and Gaussian Splat based approaches like Pointersect and P2ENet. In practical tests with seven Microsoft Azure Kinects, processing, including uploading the point clouds to GPU, requires only 13.8 ms (when using one color per point) or 29.2 ms (using high-resolution color textures), and rendering at a resolution of 3580 x 2066 takes just 3.2 ms, proving its suitability for real-time VR applications.
In this paper, we propose a competitive game in which a player wearing an augmented reality (AR) head-mounted display (HMD) and a player not wearing an HMD share not only a virtual environment but also the structure of a physical environment. Through the proposed game, we explore the interaction between players in an online multiplayer game using an AR HMD, which is enjoyable and has a high social presence. For this exploration, we created a game design that actively utilizes a physical environment and the asymmetry between players wearing and not wearing an HMD. We implemented the designed game and conducted a user study (n=14) to evaluate the game using the Game Experience Questionnaire and an our own questionnaire. The results revealed that the players had a highly positive affect toward the game and showed a high social presence. We also obtained insights into how to make the game more interesting and to increase social presence of players.