Natural interaction plays a crucial role in keeping users immersed in virtual reality, particularly when embodied conversational agents (ECAs) serve as expert guides. This becomes especially important during procedural generation, as it often involves disruptive waiting periods. This paper investigates how waiting time in virtual reality can be meaningfully filled by using an ECA for educational purposes. We present a framework that enables the ECA to manipulate procedural generation parameters and initiate the process through natural conversation. During the generation period, the ECA interacts with users to deliver educational content, effectively utilizing the waiting time. We conducted a user study to evaluate three agent behaviors during procedural generation waiting periods: topic-restricted conversation, unrestricted conversation, and topicrestricted conversation with active turn-taking. We examined how the different conversational strategies impact user motivation, time perception, and subjectively-rated learning outcomes during procedural generation waiting periods. Results suggest that topic-restricted conversations significantly enhance users' learning outcomes about circularity in the building lifecycle, while unrestricted conversations lead to a significant decrease in self-reported general learning outcomes. Additionally, active turn-taking alone did not significantly influence motivation or alter users' perception of time. These findings suggest that allowing unrestricted conversations may result in off-topic or uninformative exchanges, reducing the learning benefit. In contrast, keeping conversations focused on a specific topic enhances perceived learning outcomes.
Improper technique is detrimental to a tennis player's performance and increases the risk of injury. Regular training routines are necessary but take effort and motivation, and on-court training requires resources that pose additional barriers. Virtual reality paves the way for engaging self-training applications that address these barriers, but without a coach, motion errors may go uncorrected. We present a complementary way of practicing aspects of proper tennis forehand technique in virtual reality, utilizing automated motion analysis for immediate post-action multimodal feedback. Our VR tennis training utilizes motor learning principles and motion analysis to reinforce proper movement patterns and provide timely corrections. We overcome the problem of complex motion analysis by breaking the motion into distinct phases and utilizing the concept of coaching rules. After each shot, auditory and visual feedback is given, focusing on one aspect at a time. The exclusive use of the Meta Quest's partial motion capture poses technical challenges, restricting the set of applicable coaching rules and feedback due to the limited number of tracked joints. However, it allows us to present a more accessible and flexible alternative to on-court training and existing self-training setups. We conducted a user study $(\mathrm{N} = 26)$ following a within-subjects pretest-posttest design to evaluate short-term effects of our VR tennis training. Results demonstrate significant improvements in motivation, performance metrics, and participants' self-reported confidence in technique from the pre- to posttest, suggesting a potential short-term learning effect. Qualitative insights reveal that participants believe our VR training can complement traditional tennis training to a certain degree.
In this paper, we present OffiStretch, a camera-based system for optimal stretching guidance at home or in the workplace. It consists of a vision-based method for real-time assessment of the user’s body pose to provide visual feedback as interactive guidance during stretching exercises. Our method compares the users’ actual pose with a pre-trained target pose to assess the quality of stretching for a number of different exercises. We utilize angular and spatial pose features to perform this comparison for each individual exercise. The result of this pose assessment is presented to the user as real-time visual feedback on an "augmented mirror" display. As our method relies simply on a single RGB camera, it can be easily utilized in everyday training scenarios. We validate our method in a user study, comparing users’ performance and motivation in stretching when receiving audio-visual guidance on a TV screen both with and without our live feedback. While participants performed equally well in both conditions, feedback boosted their motivation to perform the exercises, highlighting its potential for increasing users’ well-being. Moreover, our results suggest that participants preferred stretching exercises with our live feedback over the condition without the feedback. Finally, an expert evaluation with professional physiotherapists reveals that further work must target improvements of the feedback to ensure correct guidance during stretching.
Analysis of human motion is instrumental in many areas including sports, arts, and rehabilitation. This paper presents a novel method for human motion analysis with the focus on tennis training and forehand technique assessment. We address the problems of automatic motion analysis and incorrect technique identification by a machine learning approach. We utilize the concept of training rules that are used to individually assess specific aspects of a given type of motion. Our method for motion analysis is based on insights from professional trainers and our training rules are co-designed with them. The presented method is evaluated quantitatively using recorded dataset of tennis forehand motions. This evaluation compares two variants of sport technique correctness classification: informed and uninformed learning. Both learning variants fall into the category of supervised learning, but informed learning additionally utilizes motion features and motion phases derived from tennis training methodology. Our experiments suggest that informed learning leads to higher accuracy and faster speed of the algorithm. Finally, we studied our method in a qualitative expert study.
With the emergence of large language models (LLMs), conversational agents have gained significant attention across various domains, including virtual reality (VR). This paper investigates the use of conversational agents as an interface for procedural building design in VR. We propose a voice interface that allows a user to control parameters of procedural generation and gain insights about the building construction metrics through natural conversation. The pipeline introduced for the conversational agent involves utilizing LLMs in two separate API calls for natural language understanding and natural language generation. This separation enables the invocation of various actions in procedural generation as well as meaningful agent responses to building-related questions. Furthermore, we conducted a user study to assess our proposed conversational interface in comparison to a traditional graphical user interface (GUI) in a VR architectural design task focused on circular economy. The study scrutinize the user-reported usability, presence, realism, errors, and effectiveness of both interfaces. Results suggest that while the non-embodied conversational agent enhances effectiveness due to its explanatory capabilities, it surprisingly decreases realism compared to the GUI. Overall, the preference between the conversational agent and the GUI varied greatly among participants, highlighting the need for further research into the evolving shift towards speech interaction in VR.
Temporal alignment is an inherent task in most applications dealing with videos: action recognition, motion transfer, virtual trainers, rehabilitation, etc. In this paper we dive into the understanding of this task from a geometric point of view: in particular, we show that the basic properties that are expected from a temporal alignment procedure imply that the set of aligned motions to a template form a slice to a principal fiber bundle for the group of temporal reparameterizations. A temporal alignment procedure provides a reparameterization invariant projection onto this particular slice. This geometric presentation allows us to elaborate a consistency check for testing the accuracy of any temporal alignment procedure. We apply this consistency check to some alignment procedures from the literature based on dynamic programming for the task of aligning motions of tennis players. The comparison of the obtained results leads us to propose a version of dynamic programming that incorporates keyframe correspondences. The temporal alignment procedures produced are not only more accurate, but also computationally more efficient.
We propose a ray-triangle intersection algorithm with fast-rejection strategies. We intersect the ray with the triangle plane, then transform the intersection problem into 2D by applying a transformation matrix to the ray-plane intersection point. For 2D transformation, we study two different approaches. The first approach uses a transformation matrix which transforms the triangle into a unit triangle. Then, simple 2D tests are performed. The second approach transforms the triangle into a 2D triangle while preserving similarity. This allows us to prune (i.e., to clip away) areas surrounding the triangle, determining whether the transformed intersection point lies within the triangle. We discuss several optimizations for this pruning approach. We implemented both approaches into the CPU-based ray-tracing framework PBRT, version 3, and we performed a time-based comparison against PBRT’s default intersection algorithm and Baldwin and Weber’s algorithm. The results show that our algorithms are faster than the default algorithm. They are comparable to or slightly slower than Baldwin and Weber’s algorithm, however, the pruning approach produces watertight results and may be further optimized. Moreover, additional CPU/GPU experiments outside of PBRT document promising speedup over the standard Möller–Trumbore algorithm in areas like ray-casting or collision detection.
Due to product individualization, customization and rapid technological advances in manufacturing, production systems are faced with frequent reconfiguration and expansion. Industrial buildings that allow changing production scenarios require flexible load-bearing structures and a coherent planning of the production layout and building systems. Yet, current production planning and structural building design are mostly sequential and the data and models lack interoperability. In this paper, a novel parametric evolutionary design method for automated production layout generation and optimization (PLGO) is presented, producing layout scenarios to be respected in structural building design. Results of a state-of-the-art analysis and a case study are combined to develop a novel concept of integrated production cubes and the design space for PLGO as basis for a parametric production layout design method. The integrated production cubes concept is then translated into a parametric PLGO framework, which is tested on a pilotproject of a hygiene production facility to evaluate the framework and validate the defined constraints and objectives. Results suggest that our framework can produce feasible production layout scenarios, which respect flexibility and building requirements. In future research the design process will be extended by the development of a multi-objective evolutionary optimization process for industrial buildings to provide flexible building solutions that can accommodate a selection of several prioritized production layouts.
Integrated industrial building design, incorporating building and production planning, is challenging due to sequential processes and isolated use of discipline-specific tools. Integrated optimizations are seldomly performed and sufficient decision support is lacking at early design stage. This paper presents a parametric framework for semi-automated multi-objective optimization and decision support (PMOOD) for flexible and eco-efficient industrial building structures, integrating production layout planning. The evolutionary algorithm finds best performing design solutions in terms of life cycle costs, life cycle assessment, recycling potential and flexibility. We evaluated the framework in an interdisciplinary user study based on a pilot-project from a hygiene production facility.
Parametric multi-objective optimization tools bear the potential to integrate, optimize, and explore design spaces to support interdisciplinary decision-making. A parametric optimization and decision support tool was developed (POD tool), and an evolutionary multi-objective optimization algorithm implemented (POD MOO tool) to automate design search for flexible integrated industrial building design. Both tools were tested and compared within a user study, simulating an interdisciplinary industrial building design process to evaluate if the MOO creates advanced building options in design search. Evaluations of questionnaires show the preference to search for a design by manipulating parameters instead of automatically generated designs from the algorithm.
Adaptive user interfaces can improve experiences in Extended Reality (XR) applications by adapting interface elements according to the user's context. Although extensive work explores different adaptation policies, XR creators often struggle with their implementation, which involves laborious manual scripting. The few available tools are underdeveloped for realistic XR settings where it is often necessary to consider conflicting aspects that affect an adaptation. We fill this gap by presenting AUIT, a toolkit that facilitates the design of optimization-based adaptation policies. AUIT allows creators to flexibly combine policies that address common objectives in XR applications, such as element reachability, visibility, and consistency. Instead of using rules or scripts, specifying adaptation policies via adaptation objectives simplifies the design process and enables creative exploration of adaptations. After creators decide which adaptation objectives to use, a multi-objective solver finds appropriate adaptations in real-time. A study showed that AUIT allowed creators of XR applications to quickly and easily create high-quality adaptations.
For efficient facility management it is of high importance to monitor building information, such as energy consumption, indoor temperature, occupancy as well as changes in building structure. In this paper we present a novel methodology for monitoring information about building via gamification. In our approach, the employees of a facility record the states of building elements by playing a competitive mobile game. Traditionally, external sensors are used to automatically collect information about the building usage. In contrast to that, our methodology utilizes personal mobile phones of employees as sensors to identify objects of interest and report their state. Moreover, we propose to use crowdsourcing as a tool for data collection. This way the users of the mobile game are collecting points and compete with each other. At the end of the game the winning team gets the reward. We utilized various gamification strategies to increase motivation of users to collect building data. We extended the traditional 3D BIM model with temporal domain to enable tracking of building changes over time. Finally, we run an experiment with real use case building in which the employees used our system for the duration of three months. We studied our approach and our motivation strategies in a post-experiment study. Our results suggest that gamification can be a viable tool for building information monitoring. Additionally, we note that motivation plays a critical role in the data acquisition by gamification.
To exploit the potential of immersive network analytics for engaging and effective exploration, we promote the metaphor of "egocentrism", where data depiction and interaction are adapted to the perspective of the user within a 3D network. Egocentrism has the potential to overcome some of the inherent downsides of virtual environments, e.g., visual clutter and cyber-sickness. To investigate the effect of this metaphor on immersive network exploration, we designed and evaluated interfaces of varying degrees of egocentrism. In a user study, we evaluated the effect of these interfaces on visual search tasks, efficiency of network traversal, spatial orientation, as well as cyber-sickness. Results show that a simple egocentric interface considerably improves visual search efficiency and navigation performance, yet does not decrease spatial orientation or increase cyber-sickness. An occlusion-free Ego-Bubble view of the neighborhood only marginally improves the user's performance. We tie our findings together in an open online tool for egocentric network exploration, providing actionable insights on the benefits of the egocentric network exploration metaphor.
Information visualization techniques play an important role in Virtual Reality (VR) because they improve task performance, support cognitive processes, and eventually increase the feeling of immersion. Deaf and Hard-of-Hearing (DHH) persons have special needs for information presentation because they feel and perceive VR environments differently. Therefore, it is necessary to pay attention to requirements about presenting information in VR for this group of users. Previous research showed that adding special features and using haptic methods helps DHH persons to do VR tasks better. In this paper, we propose a novel Omni-directional particle visualization method and also evaluate multi-modal presentation methods in VR for DHH persons, such as audio, visual, haptic, and a combination of them (AVH). Additionally, we compare the results with the results of persons without hearing problems. The methods for information presentation in our study focus on spatial object localization in VR. Our user studies show that both DHH persons and persons without hearing problems were able to do VR tasks significantly faster using AVH. Also, we found out that DHH persons can do visual-related VR tasks faster than persons without hearing problems by using our new proposed visualization method. Our results suggest that the benefits of using audio among persons without hearing problems and the benefits of using vision among DHH persons cause an interesting balance in the results of AVH between both groups. Finally, our qualitative and quantitative evaluation indicates that both groups of participants preferred and enjoyed AVH modality more than other modalities.
Lenka Lhotska合作论文数Department of Cybernetics , Faculty of Electrical Engineering
Czech Technical University2