While human-AI decision-making research has primarily used trust measurements to assess the practical usage of AI systems by their end-users, recent empirical evidence suggests that trust measurements do not inform users' appropriate reliance on AI systems. While examining the human-AI decision-making literature, in this work, we review empirical studies that assess people's appropriate reliance on AI advice, differentiating measurements and constructs of appropriate reliance from trust and mere reliance. Our analysis of literature shows that constructs for human-AI appropriate reliance are still fragmented in research. We present three views on appropriate reliance, namely Traditional, Appropriateness, and Dominance, as discussed in research. Using these views, we evaluate objective metrics reported in studies and argue for their consensus to facilitate the comparison across empirical research. We also discuss how studies employ objective metrics and examine their validity in application contexts. Our work contributes to the critical body of research on exploring objective metrics for assessing humans' appropriate reliance on AI advice.
An interactive vignette is a visual storytelling medium that lets the audience role-play a character and interact with non-player characters (NPCs) and the digital environment. Yet, the authoring complexity of interactive vignettes has obstructed their adoption in everyday storytelling, which builds on immediacy. We introduce DiaryPlay, an AI-assisted authoring system that generates interactive vignettes from text stories. The Authoring Component visually elicits three core elements (environment, characters, events) through automation and author refinement. The Viewing Component delivers an interactive story to the audience using an LLM-powered Controlled Divergence Module, which allows divergent player and NPC behaviors within the boundaries defined by the author’s intended story. A technical evaluation shows that the Controlled Divergence module generates believable NPC activities based on both character persona and storyline. A user study demonstrates that DiaryPlay enables low-effort authoring of interactive vignettes for everyday storytelling while providing engaging viewing experiences and conveying the core story message.
AI systems are increasingly being positioned to assist people in decision-making. However, recent empirical studies show critical concerns that people over-rely on AI advice without analytically engaging with it. While HCI research explores how people rely on AI advice, we argue that it largely overlooks an important aspect: replicating realistic decision-making scenarios. Human-AI interaction factors influence people’s reliance on AI advice. To understand human-AI interaction factors and their interplay, we conducted an analytical review of recent studies in human-AI reliance literature. We analyzed the decision-making tasks in research and their validity in application-grounded contexts. Our findings show that user engagement is a precious commodity for relying on AI advice; however, it comes at a cost. We also discuss factors contributing to “appropriate reliance”, existing research gaps, and recommendations for intervention design for human-AI reliance. Our work contributes to the critical body of research on building appropriate reliance on AI advice.
Concerns around users’ over-reliance on artificial intelligence (AI) systems are increasingly prevalent in human-AI decision-making contexts. This is partly due to users’ insufficient understanding and engagement with data and machine learning (ML) models. In this work, we explore how reliance on AI assistance is impacted through engagement with data and ML models. Using a lightweight intervention that helps steer users in investigating and understanding data and ML models, we conducted a user study with 16 participants on decision-making tasks related to customer classification. We evaluated participants’ performance and subjective experience-based feedback. Engagement with data and models improved participants’ performance and perception of trust and reliance. Participants also experienced varying model uncertainty. We find that evaluating system uncertainty inappropriately can influence self-efficacy in contesting system decisions. Our findings inform important implications for improving human-AI collaboration, for instance, by enhancing users’ interaction to engage with data and models.
In sales and marketing, customer segmentation is an important tool for formulating strategies for customer treatment and supply chain management. Most segmentation implementations rely on limited criteria, such as recency, frequency, and monetary (RFM) modeling, which often fail to capture complex business interactions. In this work, we design and evaluate a dynamic multi-criteria decision-making (MCDM) method in a business-to-business (B2B) manufacturing context by 1) extending RFM to dimensions of stability and growth, 2) integrating an adaptive and analytical hierarchical process to match business objectives, and 3) evaluating multivariate time-series clustering models. We then measure customer stability, tracking between-segment transitions, and volatility over time, and apply a graph-based consensus model to further strengthen the analysis. We test the efficacy of the proposed method using a real-world manufacturing company dataset to segment more than 3,000 B2B customers, showing strong robustness to temporal shifts. The implementation enables domain experts with preferential analytics to devise their strategies, providing effective decision support for B2B customer segmentation.
Although artificial intelligence (AI) systems are expected to support business decision-making, their adoption among non-technical professionals, for instance, in business contexts, remains limited. Prior HCI research in business contexts has mainly explored how technical users provide support to decision-makers, overlooking the challenges of actual business decision-makers themselves. In this work, we present an empirical study surveying 65 participants from business domains and interviewing 18 of them to evaluate their perceptions on analytics and/or AI usage. Our findings reveal that business decision-makers, such as sales professionals, do incorporate data analysis for informed decision-making, which they either delegate to analysts or perform by themselves. Most salespeople still rely on traditional tools, with minimal use of data-driven and AI-based analytics due to adoption barriers that range from skill gaps to an understanding of complex statistical models or tools. These findings inform the need for designing human-centered AI systems that better support business users’ decision-making.
Disability Services Office (DSO) professionals at higher education institutions write alt text for visual content. However, due to the complexity of visual content, such as HCI figures in research publications, DSO professionals can struggle to write high-quality alt text if they lack subject expertise. Generative AI has shown potential in understanding figures and writing their descriptions, yet its support for DSO professionals is underexplored, and limited work evaluates the quality of alt text generated with AI assistance. In this work, we conducted two studies: first, we investigated generative AI support for writing alt text for HCI figures with 12 DSO professionals. Second, we recruited 11 HCI experts to evaluate the alt text written by DSO professionals. Findings show that alt text written solely by DSO professionals has lower quality than alt text written with AI assistance. AI assistance also helped DSO professionals write alt text more quickly and with greater confidence; however, they reported inefficiencies in interactions with the AI. Our work contributes to exploring AI support for non-subject expert accessibility professionals.
Location-based games (LBGs) merge digital play with physical environments, creating hybrid spaces that require players to navigate complex trust dynamics. Despite their global popularity, LBGs introduce unique challenges around fairness, safety, and privacy, spanning interactions among players, game systems, local communities, and non-players in shared public spaces. To examine how trust is perceived, built, and sustained in these environments, we conducted in-depth interviews with 26 players of four major LBGs: Pokémon GO, Monster Hunter Now, Ingress, and Pikmin Bloom. Using reflexive thematic analysis, we identified dynamics of trust across four trustor–trustee relationships: player–system, player–player, player–community, and player–non-player in five key aspects: fair play, location privacy, online vetting, hybrid interaction, and public play. Drawing on our findings, we propose a trust model for analyzing and designing trust in LBGs as hybrid spaces, and we outline design implications aimed at strengthening trust building and sustaining trustworthy interactions across the LBG ecology.
Emerging multimodal conversational search (MCS) tools (e.g., Gemini Live) allow users to search for spatiotemporal information through natural language dialogues as they move through urban space. Despite the growing popularity of these tools, there is limited understanding of how people engage with this technology. To address this gap, we developed UrbanSearch, an MCS technology probe designed to capture the user’s current geolocation, time, and visual surroundings. A contextual inquiry (N=23) revealed that MCS tools provide two core values: requiring low effort in forming queries while offering highly relevant responses, and functioning as a central information gateway. As a promising technology, MCS supports environmental learning, in-situ decision making, and personalized navigation. Participants also revealed unmet needs for spatial reasoning and transparent integration of multi-source information, along with concerns related to peripheral awareness, social context, and personal space. Drawing from the findings, we discuss design implications for future MCS tools in urban spaces.
Measuring how users rely on Artificial Intelligence (AI) systems is important to explore their feasibility in real-world contexts. To this end, human-AI decision-making research uses several constructs, bifurcated between trust and reliance, for measuring human perception and behavior with AI. However, such constructs often do not help capture users’ appropriate reliance on AI advice. In this work, we provide a review of measurement constructs used in research to assess people’s appropriate reliance on AI advice. We first clarify the conceptualized differences between trust, reliance, and appropriate reliance. Then, we discuss and argue about different views of reliance and objective metrics reported in studies. Our analysis shows that measurement constructs for human-AI appropriate reliance are still nascent and require consensus among studies. Our work contributes to research on exploring objective metrics for assessing humans’ appropriate reliance on AI advice.
As mixed reality (MR) technologies become increasingly prevalent in entertainment, particularly in gaming, it is crucial to understand how users with diverse communication needs engage with these environments. However, little is known about the motivational experiences of Deaf and Hard of Hearing (DHH) players in MR games. Our study investigates the factors that influence motivation and reengagement among DHH participants using a ten-day deployment of an MR exergame. Six DHH participants with varying hearing status completed daily gameplay sessions and surveys for ten consecutive days, followed by in-depth interviews. We identified key barriers to sustained engagement, including challenges with closed captions (CCs), insufficient instructional clarity, and the absence of dynamic visual cues. Based on participant feedback, we offer design recommendations for improving accessibility and supporting longterm motivation in MR environments, such as customizable CCs, sign language alternatives, and real-time visual guidance. These findings contribute to a deeper understanding of inclusive design for DHH users in MR gaming contexts.
Tasks in augmented reality (AR), such as 3D interaction and instructional comprehension, are often designed for users with uniform sensory abilities. Such an approach, however, can overlook the more nuanced needs of Deaf and Hard of Hearing (DHH) users who might have reduced auditory perception. To better understand these challenges, our study utilized the single-player AR game Angry Birds AR as a probe to explore how 11 DHH participants and 15 hearing participants experienced AR interactions. Our findings highlight that DHH users prefer interaction based on context, effective haptic cues, audio cue substitutes, and clear instructional design. We, therefore, propose the following design recommendations to enhance the accessibility of AR for DHH users. This includes customizable UI options, modular feedback systems, and virtual avatars for sign language instructions.
An interactive vignette is a popular and immersive visual storytelling approach that invites viewers to role-play a character and influences the narrative in an interactive environment. However, it has not been widely used by everyday storytellers yet due to authoring complexity, which conflicts with the immediacy of everyday storytelling. We introduce DiaryPlay, an AI-assisted authoring system for interactive vignette creation in everyday storytelling. It takes a natural language story as input and extracts the three core elements of an interactive vignette (environment, characters, and events), enabling authors to focus on refining these elements instead of constructing them from scratch. Then, it automatically transforms the single-branch story input into a branch-and-bottleneck structure using an LLM-powered narrative planner, which enables flexible viewer interactions while freeing the author from multi-branching. A technical evaluation (N=16) shows that DiaryPlay-generated character activities are on par with human-authored ones regarding believability. A user study (N=16) shows that DiaryPlay effectively supports authors in creating interactive vignette elements, maintains authorial intent while reacting to viewer interactions, and provides engaging viewing experiences.
AI assistance is increasingly used to improve human-AI collaborative decision-making. However, how domain experts integrate their knowledge with grounded constraints and formulate intent with AI systems remains underexplored. In this position paper, we argue for "cognitively aligned" AI assistance, where users engage interactively with symbolic (logic-based) and sub-symbolic AI to interpret, influence, and co-construct decisions. Through this lens, we believe that users can build effective reliance on AI assistance, iteratively anchoring their domain knowledge to adapt their mental models and AI assistance. We explore the current literature and emphasize the need for cognitive (analytical) engagement with AI assistance to improve semantic alignment and interactive affordances for domain experts. We outline a plan for a research study that explores users' interaction with AI assistance and quantitative reasoning in business decision-making.
AI assistance can be dynamically adapted to persuade users to build reliance on AI systems. Personalizing AI assistance based on users' latent traits and real-time behavior can also improve human-AI collaborative decision-making. However, there is limited exploration in the literature on personalizing AI assistance to user traits and behavior. Understanding how users engage and interact with personalized explanations from the lens of reducing over-reliance is also underexplored. In this position paper, we present a rationale for personalized and persuasive interventions to build appropriate reliance and enhance user engagement with AI assistance. We examine the current literature and argue that user-centric persuasion and engagement improve analytical system evaluation and foster reliance on AI assistance. Considering persuasive and personalized AI assistance, we posit a study design for user-centered engagement to improve appropriate reliance.
In our effort to implement an interactive customer segmentation tool for a global manufacturing company, we identified user experience (UX) challenges with technical implications. The main challenge relates to domain users' effort, in our case sales experts, to interpret the clusters produced by an unsupervised Machine Learning (ML) algorithm, for creating a customer segmentation. An additional challenge is what sort of interactions should such a tool support to enable meaningful interpretations of the output of clustering models. In this case study, we describe what we learned from implementing an Interactive Machine Learning (IML) prototype to address such UX challenges. We leverage a multi-year real-world dataset and domain experts' feedback from a global manufacturing company to evaluate our tool. We report what we found to be effective and wish to inform designers of IML systems in the context of customer segmentation and other related unsupervised ML tools.
Users often over-rely on AI-assisted decisions without analytically engaging with them, even in practical domains. In this work, we explore persuading users to analytically engage with AI assistance to reduce their over-reliance using a complex business case of customer classification. We explore the effect of persuasive cognitive engagement through explanations and communicating system uncertainty to examine the behavior of participants having diverse expertise. We leverage their feedback and objective behavior to understand their perception of the AI performance. Our findings show a contrast in participants' subjective and objective behavior, indicating inappropriate reliance on AI assistance with the perception of system performance. However, we observe the positives of interactive cognitive engagement and identify further directions to get deeper insights into expert domains with personalized AI assistance and behavioral persuasion.