IntroductionThe emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science research at scale. Building upon prior explorations of LLM agent design, our work introduces a simulated agent society where complex social relationships dynamically form and evolve over time.MethodsAgents are given psychological drives and placed in a sandbox survival environment. We conduct an evaluation of the agent society through the lens of Thomas Hobbes’s seminal Social Contract Theory (SCT), analyzing whether agents seek to escape a brutish “state of nature” by surrendering rights to an absolute sovereign in exchange for order and security.ResultsIn our experiments, agents initially engage in unrestrained conflict, mirroring Hobbes’s depiction of the state of nature. However, as the simulation progresses, social contracts emerge, leading to the authorization of an absolute sovereign and the establishment of a peaceful commonwealth founded on mutual cooperation.DiscussionThe congruence between our LLM agent society’s evolutionary trajectory and Hobbes’s theoretical account indicates the capability of LLM’s to model intricate social dynamics that replicate forces which potentially shape human societies. By enabling insights into group behavior and emergent societal phenomena, LLM-driven multi-agent simulations hold potential for advancing our understanding of social structures, group dynamics, and complex human systems.
Grasp User Interfaces (grasp UIs) enable dual-tasking in XR by allowing interaction with digital content while holding physical objects. However, current grasp UI design practices face a fundamental challenge: existing approaches either capture user preferences through labor-intensive elicitation studies that are difficult to scale or rely on biomechanical models that overlook subjective factors. We introduce GraspR, the first computational model that predicts user preferences for single-finger microgestures in grasp UIs. Our data-driven approach combines the scalability of computational methods with human preference modeling, trained on 1,520 preferences collected via a two-alternative forced choice paradigm across eight participants and four frequently used grasp variations. We demonstrate GraspR's effectiveness through a working prototype that dynamically adjusts interface layouts across four everyday tasks. We release both the dataset and code to support future research in adaptive grasp UIs.
Conversational human-AI interaction (CHAI) have recently driven mainstream adoption of AI. However, CHAI poses two key challenges for designers and researchers: users frequently have ambiguous goals and an incomplete understanding of AI functionalities, and the interactions are brief and transient, limiting opportunities for sustained engagement with users. AI agents can help address these challenges by suggesting contextually relevant prompts, by standing in for users during early design testing, and by helping users better articulate their goals. Guided by research-through-design, we explored agentic AI workflows through the development and testing of a probe over four iterations with 10 users. We present our findings through an annotated portfolio of design artifacts, and through thematic analysis of user experiences, offering solutions to the problems of ambiguity and transient in CHAI. Furthermore, we examine the limitations and possibilities of these AI agent workflows, suggesting that similar collaborative approaches between humans and AI could benefit other areas of design.
Everyday objects, like remote controls or electric toothbrushes, are crafted with hand-accessible interfaces. Expanding on this design principle, extended reality (XR) interfaces for physical tasks could facilitate interaction without necessitating the release of grasped tools, ensuring seamless workflow integration. While established data, such as hand anthropometric measurements, guide the design of handheld objects, XR currently lacks comparable data, regarding reachability, for single-hand interfaces while grasping objects. To address this, we identify critical design factors and a design space representing grasp-proximate interfaces and introduce a simulation tool for generating reachability and displacement cost data for designing these interfaces. Additionally, using the simulation tool, we generate a dataset based on grasp taxonomy and common household objects. Finally, we share insights from a design workshop that emphasizes the significance of reachability and motion cost data, empowering XR creators to develop bespoke interfaces tailored specifically to grasping hands.
Physical skill acquisition, from sports techniques to surgical procedures, requires instruction and feedback. In the absence of a human expert, Physical Task Guidance (PTG) systems can offer a promising alternative. These systems integrate Artificial Intelligence (AI) and Mixed Reality (MR) to provide realtime feedback and guidance as users practice and learn skills using physical tools and objects. However, designing PTG systems presents challenges beyond engineering complexities. The intricate interplay between users, AI, MR interfaces, and the physical environment creates unique interaction design hurdles. To address these challenges, we present an interaction design toolkit derived from our analysis of PTG prototypes developed by eight student teams during a 10-week-long graduate course. The toolkit comprises Design Considerations, Design Patterns, and an Interaction Canvas. Our evaluation suggests that the toolkit can serve as a valuable resource for practitioners designing PTG systems and researchers developing new tools for human-AI interaction design.
Tacit knowledge of hand grip force and pressure is essential for effective tool use and object manipulation. However, this information is challenging to communicate in virtual training scenarios. To address this, we introduce realtime hand grip pressure visualization that provides users with feedback on when to tighten or loosen their grip and which specific parts of the hand require adjustment. By offering task-specific visual pressure cues, users can learn and fine-tune their grasp for specific objects and tasks. Our pilot study tested three different visualization levels of detail and results indicate the medium level to outperform the low and high levels of detail.
Many people often take walking for granted, but for individuals with mobility disabilities, this seemingly simple act can feel out of reach. This reality can foster a sense of disconnect from the world since walking is a fundamental way in which people interact with each other and the environment. Advances in virtual reality and its immersive capabilities have made it possible to enable those who have never walked in their life to “virtually” experience walking. We co-designed a VR walking experience with a person with Spinal Muscular Atrophy who has been a lifelong wheelchair user. Over 9 days, we collected data on this person’s experience through a diary study and analyzed this data to better understand the design elements required. Given that they had only ever seen others walking and had not experienced it first-hand, determining which design parameters must be considered in order to match the virtual experience to their idea of walking was challenging. Generally, we found the experience of walking to be quite positive, providing a perspective from a higher vantage point than what was available in a wheelchair. Our findings provide insights into the emotional complexities and evolving sense of agency accompanying virtual walking. These findings have implications for designing more inclusive and emotionally engaging virtual reality experiences.
Collecting multimodal user data during physical tasks such as cooking, maintenance, or physical rehab is crucial to enable the design of better AI models, interfaces, and applications. However, this is a challenging task with external cameras and sensors due to user movement, self-occlusions and diversity of data streams during task performance. In this work, we present ModBand, a wearable sensor headband with an accompanying software pipeline to collect and visualize data such as facial images, pupillometry, egocentric video, and heart rate during physical task performance. ModBand can be modified, extended, and used both as a standalone device or integrated with existing head-mounted AR devices for AI-based task guidance. Our modular design incorporates cost-effective fabrication methods, such as 3D printing, and enables convenient integration or exclusion of sensors to support custom data collection needs.
With recent computer vision techniques and user-generated content, we can augment the physical world with metadata that describes attributes, such as names, geo-locations, and visual features of physical objects. To assess the benefits of these potentially ubiquitous labels for foreign vocabulary learning, we built a proof-of-concept system that displays bilingual text and sound labels on physical objects outdoors using augmented reality. Established tools for language learning have focused on effective content delivery methods such as books and flashcards. However, recent research and consumer learning tools have begun to focus on how learning can become more mobile, ubiquitous, and desirable. To test whether our system supports vocabulary learning, we conducted a preliminary between-subjects (N=44) study. Our results indicate that participants preferred learning with virtual labels on real-world objects outdoors over learning with flashcards. Our findings motivate further investigation into mobile AR-based learning systems in outdoor settings.
Advances in artificial intelligence have transformed the paradigm of human-computer interaction, with the development of conversational AI systems playing a pivotal role. These systems employ technologies such as natural language processing and machine learning to simulate intelligent and human-like conversations. Driven by the personal experience of an individual with a neuromuscular disease who faces challenges with leaving home and contends with limited hand-motor control when operating digital systems, including conversational AI platforms, we propose a method aimed at enriching their interaction with conversational AI. Our prototype allows the creation of multiple agent personas based on hobbies and interests, to support topic-based conversations. In contrast with existing systems, such as Replika, that offer a 1:1 relation with a virtual agent, our design enables one-to-many relationships, easing the process of interaction for this individual by reducing the need for constant data input. We can imagine our prototype potentially helping others who are in a similar situation with reduced typing/input ability.
Virtual content placement in physical scenes is a crucial aspect of augmented reality (AR). This task is particularly challenging when the virtual elements must adapt to multiple target physical environments unknown during development. AR authors use strategies such as manual placement performed by end-users, automated placement powered by author-defined constraints, and procedural content generation to adapt virtual content to physical spaces. Although effective, these options require human effort or annotated virtual assets. As an alternative, we present ARfy, a pipeline to support the adaptive placement of virtual content from pre-existing 3D scenes in arbitrary physical spaces. ARfy does not require intervention by end-users or asset annotation by AR authors. We demonstrate the pipeline capabilities using simulations on a publicly available indoor space dataset. ARfy makes any generic 3D scene automatically AR-ready and provides evaluation tools to facilitate future research on adaptive virtual content placement.
With the rise of Artificial Intelligence (AI) in the mainstream and the impending need for an AI-trained workforce, we must devise strategies to lower the entry barrier to AI education. Advanced mathematical preparation and computational thinking skills are two major barriers in imparting a rigorous AI course at the high school level. Consequently, many existing AI-focused educational programs for high school students are basic primers and lack technical depth. In this paper, we assess two pedagogical instruments for increasing the self-efficacy of students in learning neural networks at the high school level. The first research question is whether high school students learn the basics of neural network design through scaffolded AI research projects. We also explore whether a dual advising structure with a research mentor and a communication teaching assistant enhances student's self-efficacy in computing. For both of these questions, we define key variables to quantify student mastery and their computational thinking using qualitative student feedback and student reflection using GPT-3. We provide a reproducible blueprint for using large language models in this task to assess student learning in other contexts as well. We also correlate our results with a pre- and post-course Likert survey to find significant factors that affect student self-efficacy and belonging in AI. With our course design and dual advising mentoring model, we find that students showed a significant improvement in their ability to articulate technical aspects within the AI domain and an increase in their confidence in speaking up in the AI field. Two out of the ten research projects applied AI techniques beyond classroom teachings, yielding original research contributions, and another six showcased students' capabilities in building neural networks from scratch. Our study has a strong selection bias since it focuses on top-performing students. However, the exploration of the two pedagogical instruments (scaffolding research projects and dual advising structure) aimed at high school students provides promising insights for future AI curricula design at the high school level.