Expressive facial animation depends on models that can convey subtle, context-dependent emotions. This paper explores the zero-shot potential of large language models (LLMs) to generate facial animations. Using the cognitively-grounded Ortony-Clore- Collins (OCC) model as a framework, we designed 110 text-based scenarios and evaluated the ability of different LLMs – Gemini 2.5 Pro, GPT-4o, and Llama 3.1-8b – to generate corresponding facial animations using the Facial Action Coding System (FACS). A perceptual user study confirmed that animations from Gemini 2.5 Pro were highly recognizable, with participants successfully matching facial expressions to the correct context at rates significantly above chance. We also layered speech-driven procedural lip synchronization on top of the generated facial expressions to assess the resilience of the emotional content. A second perceptual study revealed that recognition remained robust, suggesting that lip synchronization can be incorporated into synthetic performances without masking intended affective states. Quantitative analysis of the generated Action Units (AUs) and their temporal dynamics indicated both consistency within each OCC emotion category and diversity across different scenarios. Further analysis revealed that the emotional expressions group into six clusters: happiness, sadness, anger/disgust, fear/surprise, shame, and neutral. This work demonstrates a viable, lightweight pipeline connecting textual narrative directly to motion generation without requiring custom training or large-scale motion capture datasets.
Virtual reality relies on tracking to provide users with a compelling and comfortable experience. However, such VR tracking increases the amount of data we generate about ourselves when online. In this work, we investigate the extent to which social gestures can be used to identify us. Using a dataset of 11 speakers, collected with motion capture, we train a fully convolutional network with features that emulate the motion tracking done by virtual reality systems. We find high identification accuracy of up to 99%, even with very short clips (4 seconds).
Expressive facial animation depends on models that can convey subtle, context-dependent emotions. Procedural methods often rely on manually-tuned heuristics, while data-driven techniques are constrained by the diversity of their training data. This paper explores the zero-shot potential of large language models (LLMs) to generate facial animations. Using the cognitively-grounded Ortony-Clore-Collins (OCC) model as a framework, we designed 110 text-based scenarios and evaluated the ability of different LLMs - Gemini-2.5 Pro, GPT-4o, and Llama 3.1-8b - to generate corresponding facial animations using the Facial Action Coding System (FACS). A perceptual user study confirmed that animations from Gemini-2.5 Pro were highly recognizable, with participants successfully matching facial expressions to the correct context at rates significantly above chance. A quantitative analysis of the generated Action Units (AUs) indicated both consistency within each OCC emotion category and diversity across different scenarios. Further analysis revealed that the emotional expressions group into six clusters: happiness, sadness, anger/disgust, fear/surprise, shame, and neutral. This work demonstrates a viable, lightweight pipeline connecting textual narrative directly to motion generation without requiring custom training or large-scale motion capture datasets.
The objective of MASSXR 2025, the 3rd Workshop on Multi-modal Affective and Social Behavior Analysis and Synthesis in Extended Reality, was to bring together researchers and practitioners from fields including cybersecurity, human-computer interaction, computer graphics/animation, multi-modal machine learning, Artificial Intelligence (AI), data privacy, and socio-technical studies to discuss the state of security in Extended Reality (XR), as well as future directions and opportunities. Through this, it aimed to achieve adaptive, context-aware security measures that are both technically robust and aligned with user trust and understanding. The workshop provided an opportunity to foster collaborative research efforts and advance the state of secure social interactions within XR, setting a foundation for future innovations in the field.
Designing games with branching story lines, object annotations, scene details, and dialog can be challenging due to the intensive authoring required. We investigate the potential for authoring open-ended behaviors for point-and-click narrative games using GPT-3.5, a large language model. In our approach, we extend a behavior tree scripting system with nodes that query GPT-3.5 to generate object descriptions, conversations with characters, and responses to player actions. GPT-3.5 is used to generate content when it hasn’t been scripted manually and to update game state by asking questions about whether a player’s input achieves a particular game goal. We demonstrate our approach with puzzles based on scenes from an episode of Star Trek Voyager. Our approach aims to blend a specific plot with open-ended story elements while keeping the authoring work minimal. Based on a pilot study of 16 participants and our own testing, we find that the generated responses have high coherency and show signs of humor and novelty, but that utterances could be improved to be more interesting and better support the designer’s intent.
Animating performances for story-based games is a difficult and labor-intensive task. Although much research in animation and intelligent agents has focused on the problem of generating animation from textual descriptions, this work explores a novel approach through the use of a Large Language Model (LLM) that animates non-player characters with personality and emotions. The proposed approach thus reduces the typical authoring and setup requirements. As a proof of concept, we demonstrate the approach within a point-and-click narrative game.
In this work, we investigate three control strategies - joystick, laser, and tilt - for playing a platform game in mobile augmented reality. We analyze these strategies using both objective game metrics as well as self-reported measures of autonomy, competence, and intuitiveness. We found no significant differences in self-reported autonomy and competence between our three controllers, despite clear differences in duration needed to play, clear differences in the amount of movement of the character, and significant differences in intuitiveness ratings. All controllers were effective for playing the game. However, despite the joystick's faster completion times and higher intuitiveness ratings, half of our players still chose our laser or tilt controller as their favorite. These results were consistent even with different difficulties of game level.
In this work, we describe the Curation Tree framework that has matured as part of our game-based research and teaching. We use this framework to manage the top-level experience of our players via a text-based, authorable behavior tree system that is easy to implement, test, and re-use. We present three case studies based on Unity games built within our framework: the first involves virtual reality games involving grasping objects with the hands; the second, Treatment X, uses the framework to implement a causal learning game; and the last involves narrative games based on a large language model. Based on our experiences with this system, we reflect on its limitations and make several recommendations based on its strengths, namely to centralize game logic with behaviors, to centralize user event handling, to prefer simple behaviors over complex ones, to use text-based scripts for authoring with version control, and to separate behaviors from persistent state.
Detailed hand motions play an important role in face-to-face communication to emphasize points, describe objects, clarify concepts, or replace words altogether. While shared virtual reality (VR) spaces are becoming more popular, these spaces do not, in most cases, capture and display accurate hand motions. In this article, we investigate the consequences of such errors in hand and finger motions on comprehension, character perception, social presence, and user comfort. We conduct three perceptual experiments where participants guess words and movie titles based on motion captured movements. We introduce errors and alterations to the hand movements and apply techniques to synthesize or correct hand motions. We collect data from more than 1000 Amazon Mechanical Turk participants in two large experiments, and conduct a third experiment in VR. As results might differ depending on the virtual character used, we investigate all effects on two virtual characters of different levels of realism. We furthermore investigate the effects of clip length in our experiments. Amongst other results, we show that the absence of finger motion significantly reduces comprehension and negatively affects people’s perception of a virtual character and their social presence. Adding some hand motions, even random ones, does attenuate some of these effects when it comes to the perception of the virtual character or social presence, but it does not necessarily improve comprehension. Slightly inaccurate or erroneous hand motions are sufficient to achieve the same level of comprehension as with accurate hand motions. They might however still affect the viewers’ impression of a character. Finally, jittering hand motions should be avoided as they significantly decrease user comfort.
•The special section contains the best work from Motion, Interaction, and Games, 2022.•These top papers span diverse topics in animation, control, and simulation.•Articles present cutting-edge research in deep learning, perceptual studies, and more.
Curiosity is generally considered to be a large driver of video game players’ motivation and enjoyment. However, it is unclear how much curiosity is driven by intrinsic personality factors versus the game’s design. We explore this question through the lens of the puzzle game, Monument Valley. We create two categories of puzzles. The first category consists of simple puzzles which can be quickly solved. The second category consists of puzzles recreated from the original game. Using these puzzles, we create an online experiment platform that asks players about their innate curiosity for exploration and problem solving and then asks them to play our puzzles. In a small pilot study of this system, we analyzed the time-spent, clicks, ratings, and survey responses of 10 participants. Surprisingly, we found differences in time-spent even with our short puzzles. We also found that our participants spent the largest amount of time with puzzles that could not be solved. These results suggest future directions for research into how curiosity and persistence may be related in the context of puzzle solving.
This project explores how to best generate puzzle games that require path planning to solve. In particular, its focus is narrowed to a subset of puzzle games that challenge the player to find a path from point A to point B. This path may not be readily available to the player and must be found by them activating a series of “switches.” These “switches” modify the game environment and, consequently, the paths available to navigate. The goal of this project is to develop an algorithm which can automatically generate this type of puzzle game. To achieve this, we look to a specific path-puzzle game, called Monument Valley , for a basis to begin experimentation. Specifically, the work for the project requires Unity to build Monument Valley styled puzzle game levels and C# to program the “switches” and conduct the level generation.
Virtual environments have become ubiquitous, expanding beyond games into the domains of architecture, engineering, psychology, education, and archaeology. Furthermore, virtual humans can further enhance these environments when they provide compelling and coherent behaviors. In this paper, we present a scripting language based on simple, plain English commands. Our system assists people without game and animation expertise to populate large environments and complex scenarios. To validate our approach, we develop a prototype using Unreal Engine 4 and author a variety of indoor and outdoor agent simulations. Furthermore, we test our prototype with both experienced and inexperienced users, creating scenarios for a residence, mall, psychology scenario, and archaeological site.
The complexity of game play in online multiplayer games has generated strong interest in modeling the different play styles or strategies used by players for success. We develop a hierarchical Bayesian regression approach for the online multiplayer game Battlefield 3 where performance is modeled as a function of the roles, game type, and map taken on by that player in each of their matches. We use a Dirichlet process prior that enables the clustering of players that have similar player-specific coefficients in our regression model, which allows us to discover common play styles amongst our sample of Battlefield 3 players. This Bayesian semi-parametric clustering approach has several advantages: the number of common play styles do not need to be specified, players can move between multiple clusters, and the resulting groupings often have a straight-forward interpretations. We examine the most common play styles among Battlefield 3 players in detail and find groups of players that exhibit overall high performance, as well as groupings of players that perform particularly well in specific game types, maps and roles. We are also able to differentiate between players that are stable members of a particular play style from hybrid players that exhibit multiple play styles across their matches. Modeling this landscape of different play styles will aid game developers in developing specialized tutorials for new participants as well as improving the construction of complementary teams in their online matching queues.
A primary goal of the Virtual Reality (VR) community is to build fully immersive and presence-inducing environments with seamless and natural interactions. To reach this goal, researchers are investigating how to best directly use our hands to interact with a virtual environment using hand tracking. Most studies in this field require participants to perform repetitive tasks. In this article, we investigate if results of such studies translate into a real application and game-like experience. We designed a virtual escape room in which participants interact with various objects to gather clues and complete puzzles. In a between-subjects study, we examine the effects of two input modalities (controllers vs. hand tracking) and two grasping visualizations (continuously tracked hands vs. virtual hands that disappear when grasping) on ownership, realism, efficiency, enjoyment, and presence. Our results show that ownership, realism, enjoyment, and presence increased when using hand tracking compared to controllers. Visualizing the tracked hands during grasps leads to higher ratings in one of our ownership questions and one of our enjoyment questions compared to having the virtual hands disappear during grasps as is common in many applications. We also confirm some of the main results of two studies that have a repetitive design in a more realistic gaming scenario that might be closer to a typical user experience.
Augmented reality (AR) gaming is becoming widely available thanks to improvements in hand-held devices such as phones and tablets. In this work, we describe our system for generating levels for the AR game, Q*bird. In Q*bird, the player must visit every cell in the level while avoiding bees and cannon balls, similarly to the 1982 arcade game, Q*bert. To create a new level, designers place game elements using virtual cards. The system then generates the remainder of the level, ensuring that it’s navigable. Designers can edit these levels by dragging and dropping the created geometry. To test, the designer can drop a character into the level and play it. This system aids playtesting and level design by allowing levels to be quickly specified and tested in the same environment in which the game is played. Furthermore, this system offers an example of how the design of AR levels can also be performed in AR.
Most commercial virtual reality applications with self avatars provide users with a “one-size fits all” avatar. While the height of this body may be scaled to the user’s height, other body proportions, such as limb length and hand size, are rarely customized to fit an individual user. Prior research has shown that mismatches between users’ avatars and their actual bodies can affect size perception and feelings of body ownership. In this paper, we consider how concepts related to the virtual hand illusion, user experience, and task efficiency are influenced by variations between the size of a user’s actual hand and their avatar’s hand. We also consider how using a tracked controller or tracked gestures affect these concepts. We conducted a 2x3 within-subjects study (n=20), with two levels of input modality: using tracked finger motion vs. a hand-held controller (Glove vs. Controller), and three levels of hand scaling (Small, Fit, and Large). Participants completed 2 block-assembly trials for each condition (for a total of 12 trials). Time, mistakes, and a user experience survey were recorded for each trial. Participants experienced stronger feelings of ownership and realism in the Glove condition. Efficiency was higher in the Controller condition and supported by play data of more time spent, blocks grabbed, and blocks dropped in the Glove condition. We did not find enough evidence for a change in agency and the intensity of the virtual hand illusion depending on hand size. Over half of the *e-mail: lorraine@clemson.edu †e-mail: alinen@savvysine.com ‡e-mail: adkins4@clemson.edu §e-mail: ysun3@g.clemson.edu e-mail: arobb@clemson.edu ||e-mail: yuting.ye@oculus.com **e-mail: max.diluca@oculus.com ††e-mail: sjoerg@clemson.edu participants indicated preferring the Glove condition over the Controller condition, mentioning fun and efficiency as factors in their choices. Preferences on hand scaling were mixed but often attributed to efficiency. Participants liked the appearance of their virtual hand more while using the Fit instead of Large hands. Several interaction effects were observed between input modality and hand scaling, for example, for smaller hands, tracked hands evoked stronger feelings of ownership compared to using a controller. Our results show that the virtual hand illusion is stronger when participants are able to control a hand directly rather than with a hand-held device, and that the virtual reality task must first be considered to determine which modality and hand size are the most applicable.
In this work, we investigate the influence of different visualizations on a manipulation task in virtual reality (VR). Without the haptic feedback of the real world, grasping in VR might result in intersections with virtual objects. As people are highly sensitive when it comes to perceiving collisions, it might look more appealing to avoid intersections and visualize non-colliding hand motions. However, correcting the position of the hand or fingers results in a visual-proprioceptive discrepancy and must be used with caution. Furthermore, the lack of haptic feedback in the virtual world might result in slower actions as a user might not know exactly when a grasp has occurred. This reduced performance could be remediated with adequate visual feedback. In this study, we analyze the performance, level of ownership, and user preference of eight different visual feedback techniques for virtual grasping. Three techniques show the tracked hand (with or without grasping feedback), even if it intersects with the grasped object. Another three techniques display a hand without intersections with the object, called outer hand, simulating the look of a real world interaction. One visualization is a compromise between the two groups, showing both a primary outer hand and a secondary tracked hand. Finally, in the last visualization the hand disappears during the grasping activity. In an experiment, users perform a pick-and-place task for each feedback technique. We use high fidelity marker-based hand tracking to control the virtual hands in real time. We found that the tracked hand visualizations result in better performance, however, the outer hand visualizations were preferred. We also find indications that ownership is higher with the outer hand visualizations.
Alla Safonova合作论文数University of Pennsylvania;Computing and Information Science;School of Engineering and Applied Science5