
Most academic institutions have returned to primarily in-person instruction, but e-proctoring software remains a convenient and resource-efficient option to remotely administer and monitor online examinations, despite questions about its effectiveness. Our paper reports on semi-structured interviews with 33 individuals (19 instructors and 14 students) at a large Canadian university. We use the Networked Privacy theoretical framework as a lens to consider the privacy implications of our interview results and how these impact interactions with e-proctoring technology. Our work offers three research contributions: (1) a qualitative comparison of both instructors and students’ perspectives on e-proctoring, and their perspective on the other group's motivations, (2) an examination of using e-proctoring within a university through the lens of networked privacy, enabling reflection on the interconnected aspects and power dynamics involved, and (3) actionable recommendations for institutions and instructors who wish to continue using e-proctoring, as well as design recommendations for technology providers.
Despite skepticism surrounding the integration of Generative AI (GenAI) into creative domains, its potential to transform Computer-Aided Design (CAD) workflows remains largely underexplored. In this study, we investigate how professional 3D CAD designers envision and evaluate GenAI-supported design processes. Using a “design concepts” approach, we produced videos that illustrated futuristic GenAI capabilities, ranging from modifying existing models to generating intricate models from text prompts. Our mixed-methods study, involving an international survey of 106 CAD designers and interviews with eight participants, reveals a surprising openness to adopting GenAI. Contrary to common narratives of creative resistance, participants prioritized efficiency gains and sought AI features that support functionality and creativity. Moreover, we found that no single metaphor fully captures GenAI’s role in CAD: participants alternately positioned it as tool, creative partner, and skilled mentor. This framing advances HCI theory on human-AI collaboration and has implications for designing future GenAI systems for CAD workflows.
The NASA-Task Load Index (NASA-TLX) is a widely used multidimensional scale for measuring subjective workload in Human-Computer Interaction (HCI) research. Despite widespread adoption, diverse reporting methods hinder cross-study comparisons and standardization. We systematically reviewed 522 CHI papers (2006–2024) to identify usage variations and common implementation errors, and collected subscale values from 683 tasks across 185 papers for meta-analysis. Correlation, regression, and factor analyses revealed stronger inter-subscale correlations than the original development study, different beta weights, and weakened discriminability. Factor analysis identified a two-factor structure, with Performance showing notably low communality as a distinct dimension. We propose a Short TLX (STLX) that retains Mental Demand , Physical Demand , and Performance as a condensed instrument, given that many studies already omit subscales. We provide standardized implementation guidelines and contribute an interactive research database and a data collection toolkit to facilitate consistent use of NASA-TLX across the HCI community.
Monitoring patients’ pain is essential in clinical contexts, particularly post-surgery. However, collecting pain data is time-consuming and relies on manual data capture by clinical staff. Few studies assess the efficacy of pain-logging technology in a long-term, ecologically valid manner. We report on a 6-year deployment of a minimalist pain-logging device, the PainPad, in post-surgery orthopaedic wards in a large UK hospital. Our data show that patients (n > 500) tend to submit more pain scores that are at a higher rating compared to nurse-collected data. Interviews with the nursing staff using the device highlight the importance of understanding the complexities of the context of use when trying to integrate novel technology into clinical practice. Finally, we contribute reflections on the lessons learnt from running a significant long-term deployment of a prototype health technology in a naturalistic setting.
Self-tracking apps are a prominent component of today’s wellness and health management strategies. These apps collect a variety of personal data, which is typically stored in a remote central location and utilised by machine learning models. Federated Learning (FL) brings about a fundamental change in this for machine learning. Through interviews and one workshop with 18 university students (majority of whom (16) identified as women), we investigate whether users’ increased awareness of enhanced-privacy protection in exchange for potentially reduced accuracy and fairness afforded by FL can alter users’ perceptions of these apps and the types of data they feel comfortable sharing. Participants’ willingness to pay for the FL application varied depending on their sensitivity towards data privacy. However, participants were still reluctant to share certain types of data even after knowing that FL keeps data on their devices. This was due to a persistent lack of trust in companies accompanied by a lack of awareness of the privacy implications of sharing sensitive data. Our findings suggest the need to further inform and educate users about the privacy advantages of FL. Companies must also prioritise establishing trust with users, as this was found to be a strong factor in increasing users’ acceptability of the apps, irrespective of the privacy-enhanced technology employed.
AI is not only a neutral tool in team settings; it influence the social and cognitive fabric of collaboration. Across two randomized experiments, we demonstrate that AI exposure produces causal spillover into human–human interaction—affecting shared language, collective attention, shared mental models, and social cohesion. These spillover effects occur robustly across settings, modalities, tasks, and AI qualities, suggesting that mere exposure to AI drives the influence. AI functions as an implicit “social forcefield,” influencing not only how people speak, but also how they think, what they attend to, and how they relate to each other. We argue for shifting the design paradigm from optimizing “AI as a tool” to understanding AI as a socially influential actor whose effects extend beyond the human–AI interface.
Everyone needs to somehow manage personal information, and there are numerous choices for technologies that support personal information management (PIM). Meanwhile, many suffer from experiences of information overload related to PIM. We developed a quantitative model of information overload in PIM and tested it against data we collected in a large-scale survey study (N = 1,011). Using structural equation modeling, we found that using many PIM technologies accounts for information overload through practices of keeping information and feeling of desperation, and through organizing and feeling of efficacy, but not directly. Our study deepens the understanding of information overload in the context of PIM and distinguishes different roles that technologies, practices, and affects have in relation to it. The findings highlight how designs and interventions targeted to reduce negative affects related to PIM could be the most effective in countering information overload.
Conditionally automated vehicles require drivers to remain hazard-aware while disengaged from driving and engaged in Non-Driving Related Tasks (NDRTs) Conventional visual warning interfaces do not fully address this attentional conflict. While psychological research demonstrates that social cues are processed more efficiently than visual information, this has never been evaluated for safety–critical warnings in conditionally automated driving. This article presents the first empirical investigation into using a virtual agent to indicate danger and aid situational awareness while driver performs an NDRT. Across two experiments (n = 48), hazard prediction ability was assessed with either a visual cue or a virtual agent cueing danger. When including an attention-capturing element, both visual and social cue types mitigated driver distraction, maintaining hazard prediction comparable with undistracted driving. These findings suggest that virtual agents can be as effective as conventional visual cues in aiding situational awareness in safety–critical automated driving contexts, offering design insights for in-vehicle attentional displays supporting driver awareness.
Applications of large language models (LLMs) in software engineering are soaring quickly. Despite ample evidence supporting LLMs’ capabilities of generating high-quality and creative suggestions, human–LLM interactions in software testing remain under-explored, and there lacks empirical evidence supporting efficacious human–LLM system designs. In software testing, the industry’s pursuit of “autonomous” systems neglects intensive human involvement in prior-design and post-review processes. This article discusses two empirical user studies for a test case brainstorming task: (1) exploring user behaviors in human–LLM interactions compared to web search ( \(N_{1}=16\) ) and (2) investigating three modified interaction strategies—preemptive prompting, buffered response, and guided input ( \(N_{2}=24\) ). We consolidate nine cross-disciplinary metrics to quantitatively evaluate the holistic performance of the human–LLM system, covering three perspectives: test quality, creativity, and attention span. Our findings reveal that users spend 126% more time interacting with LLMs compared to Google search. Moreover, preemptively prompting the LLM system significantly improves test quality and task creativity by over 30%, while simultaneously reducing user idle time by up to 49%. Based on the results, this article discusses three interaction design principles—mixed initiative, acceptability, and appropriation—as guidance for future iterations of an efficacious LLM-assisted software testing system.
Developmental Language Disorder (DLD) is characterized by persistent and significant language impairments. Children with DLD may experience difficulties in grammar, phonology, vocabulary, and specific morphosyntactic structures, such as clitic pronouns and passive constructions. To address these challenges, we propose an innovative treatment approach termed the Acting Argument Paradigm (AAP). Grounded in psycholinguistic theories, the core idea of this approach is to translate the abstract linguistic concept of argument movement into physical actions that actively engage children during language therapy. To operationalize the AAP, we developed a Tangible User Interface, namely Moovy, that enables children to manipulate physical objects and fosters a multisensory learning experience. This article presents the AAP and Moovy and discusses the results of a 6-week experimental study involving children (N = 30) with DLD. Preliminary findings indicate Moovy’s potential as an effective tool for speech therapy programs, offering promise for improving the linguistic skills of children with DLD.
AI decision-support tools typically offer a fixed type of assistance, like AI recommendations and explanations, regardless of the specific decision, individual, or broader context. This fixed design has been shown to hinder both human-AI decision accuracy and human skill improvement in the task. We posit that AI assistance needs to be dynamic, changing in response to contextual factors (e.g., AI uncertainty, task difficulty), individual differences, and specified objectives (e.g., decision accuracy, skill improvement). To enable such adaptive support, we propose Reinforcement Learning (RL) as a general approach for modeling human-AI decision-making to optimize human-AI interaction for diverse objectives. RL enables optimizing various objectives in AI-assisted decision-making by tailoring and adaptively providing decision support to humans—the right type of assistance, to the right person, at the right time. We instantiated our approach with two objectives: human-AI accuracy on the decision-making task and human skill improvement (i.e., learning about the task) and learned decision support policies from previous human-AI interaction data. We compared the optimized policies against several baselines in AI-assisted decision-making. Across two experiments (N = 316 and N = 964), our results consistently demonstrated that people interacting with policies optimized for accuracy achieve significantly higher accuracy—and even human-AI complementarity—compared to those interacting with any other type of AI support. Our results further indicated that human learning was more difficult to optimize than accuracy. While the policies learned the best available actions to optimize learning, participants who interacted with learning-optimized policies showed significant learning improvement only at times. Our research (1) demonstrates offline RL to be a promising approach to model the dynamics of human-AI decision-making, leading to policies that may optimize various objectives and provide novel insights about the AI-assisted decision-making space, and (2) emphasizes the importance of considering skill improvement and other human-centric objectives beyond accuracy in AI-assisted decision-making, opening up the novel research challenge of optimizing human-AI interaction for such objectives.
To address behavioral health shortages, we propose virtual counselor extenders: socially interactive virtual health agents (VHAs) that operationalize evidence-based interventions to expand clinicians’ reach and mitigate barriers to access for stigmatized populations. Current VHAs often lack systematic therapeutic communication and evidence-based environmental design. We present five contributions: (1) the implementation of a complete Brief Motivational Interviewing (BMI) intervention for heavy-drinking adults; (2) the application of a social cue taxonomy to systematically document multimodal active listening behaviors; (3) the design of a therapeutic ”soft room” grounded in environmental psychology; (4) a quantitative evaluation of usability and engagement with the target population; and (5) a qualitative thematic analysis of user acceptance regarding agent demographics and environmental realism. We conclude by extending the taxonomy with therapeutic-specific cues and providing design recommendations for future virtual counseling systems. This work demonstrates that autonomous agents can deliver complex therapeutic interventions while maintaining high user acceptance.
Technology-mediated scams are a pervasive and enduring form of digital harm affecting older adults. While prior research has largely emphasized prevention, awareness, and risk perception, comparatively little attention has been paid to what follows after a scam occurs. Conceptualizing the scam as an unfolding event, we argue that each stage of the scam—before, during, and after—is shaped by distinct sociotechnical arrangements and relational dynamics that demand tailored forms of analysis and intervention. Through a multi-method study, we examine the during and after, paying specific attention to the post-scam, during which older adults navigate the affective, social, and material aftermath of scams. In doing so, we extend HCI’s engagement with digital harm by articulating the post-scam as a critical site of coordination, redress, and justice-seeking, offering new directions for research and intervention.
We introduce Prompt-to-Touch, a proof-of-concept of a multi-step pipeline that generates haptic effects based on the textual description of a haptic experience. Our pipeline first translates the haptic effect description to a sound effect description using our Foley-Interpreter component. It then uses a text-to-audio model to generate a sound effect from the sound description. Afterwards, the sound effect is converted into a perceivable haptic effect using our Dynamic-Audio-Processor . Finally, the haptic effect is post-processed to compensate for actuator-specific frequency response characteristics. We validate our concept in two preliminary human evaluation studies (n = 20, n = 10) and a technical analysis. Our results indicate that our pipeline can generate effects to enhance immersive multimedia experiences, abstract desktop/XR interactions, and social communication applications. They also still reveal significant potential for further improvement. Our pipeline could be used to develop future text-driven haptic design and automation tools. We provide open-source code to support future extensions.
Virtual reality enables immersion in any environment. According to the hue-heat hypothesis, an environment’s color temperature influences thermal perception and regulation, factors important for productivity and energy consumption. Previous work suggests that visual thermal cues, e.g., snowy or desert environments, affect thermal perception and regulation. As color temperature was not controlled, it is unknown if hue or thermal cues caused the effects. In our first study, we derived two virtual environments with cold or warm thermal cues. In the second study, participants experienced these environments with varying hues for 5 minutes. We found that hue influences thermal sensation and thermal cues influence sensation, comfort, and skin temperature. As these effects seem to increase beyond 5 minutes, a third study with longer exposure and more extreme hue differences confirmed and extended these findings. Overall, our results consistently show that hue and visual thermal cues independently affect thermal perception and regulation. This highlights the importance of considering both when designing virtual environments.
In AI decision support, explainability (e.g., disclosing system weaknesses) should help users spot erroneous recommendations, but its efficacy may depend on task difficulty. We ran three experiments in a simulated medical visual detection task. In Experiment 1, we manipulated error difficulty (easy vs. difficult) and explainability (non-XAI vs. XAI). Experiment 2 added a virtually impossible error difficulty. Across both, explainability consistently reduced reliance on incorrect recommendations for difficult errors, showed no benefit for easy errors, and showed a small benefit for impossible errors. Experiment 3 varied error and task difficulty within-subjects and extended these patterns; as task difficulty rose, participants behaved less rationally, exhibiting both under- and overreliance. Notably, these behavioral benefits were generally not accompanied by reduced trust in the AI system. Our findings suggest that disclosing system weaknesses enhances detection of AI errors but is most effective for tasks of moderate difficulty where AI recommendations are still verifiable.
This study explores whether privacy literacy, self-efficacy, and concerns mediate the age-related effects on privacy decision behavior. To study privacy decision behavior, we designed an experiment that integrates both heuristic and cognitive manipulations in the decision scenario. 625 old and younger adults participated in the experiment and used our web-based application, “RecipeDigger.” The application recorded users’ privacy decision behavior in the form of accepting or rejecting cookies, which offered a personalized service. Our findings indicate that some of the differences in privacy decision-making between older and younger adults can be traced to having different levels of privacy literacy. Older and younger adults with higher privacy literacy can better align their privacy preferences with their disclosure behavior. By bridging the gap between psychological theories and privacy research, this study provides a comprehensive understanding of the factors influencing privacy decisions among older and younger adults, offering methodological, theoretical, and policy implications.
Eating is complex, as is people’s relationship to food, which significantly impacts well-being and health. Mindful eating interventions have been shown to be effective in addressing problematic eating, and mindful eating has also gained increasing interest as a research topic in Human-Computer Interaction (HCI), albeit with limited grounding in health research, so that we know little about how HCI work could be informed by health research on mindful eating and its interventions. Such a theoretical foundation grounded in health research is crucial to ensure the design of safe and effective mindful eating technologies and interventions. To address this gap, we present an analysis of health research on mindful eating principles, measurement scales, and therapeutic interventions such as the MB-EAT program, which helped us identify the main aspects of mindful eating. Then, we conducted a scoping review on technologies targeting these aspects, from which we curated 16 design exemplars representing the breadth of these technologies and generated a new SmartPlate conceptual design. Then, we designed the Mindful Eating Design Critique (MEDEC) cards and reported workshops with 36 mindful eating practitioners who used the MEDEC cards to critique the set of design exemplars. Our main contributions include a solid theoretical foundation grounded in health research for the design of mindful eating technologies, the MEDEC cards as a novel critique tool consisting of 29 cards improved based on practitioners’ feedback, and a design framework to further support research and development of mindful eating technologies.
Group recommendation systems (GRS) provide personalized recommendations to multiple users. While GRS has garnered attention as a social catalyst that can foster social interactions among users, how to design GRS to effectively support such experiences remains underexplored. To address this gap, we conducted an empirical study of a YouTube video GRS involving 23 participants across eight groups over 16 days. Using experience prototyping combined with a diary study and follow-up interviews, we investigated users’ perceptions of three design attributes of GRS: algorithmic controllability, algorithmic transparency, and behavioral visibility. Our findings reveal the social experiences enabled by these design attributes, as well as the challenges and concerns that emerged. Based on these findings, we conceptualize five roles of GRS as a social catalyst (i.e., Hinter, Connector, Probe, Homing Pigeon, and Provocateur ) and suggest design considerations for leveraging GRS as a novel design material to facilitate meaningful social interactions.
User interactions with mobile applications (apps) are accompanied by continuous visual changes in the Graphical User Interface (GUI), guiding task completion and feedback. These changes help users complete intended tasks or assess the appropriateness of their actions, typically conveyed through visual cues such as appearance and color. While such visual changes are effective for sighted users, they are inaccessible to blind users, creating substantial barriers to GUI interaction. To address these challenges, we propose VisualDroid , a method based on a multi-modal large language model (LLM) for testing and classifying GUI visual changes using a tailored three-hop reasoning prompting framework. VisualDroid achieved an F1-score of 94.7% in 34 apps from 17 domains, surpassing all baseline methods. When evaluated on five open-source apps from F-Droid, our method enabled developers to resolve three identified issues, with two still under review. In terms of efficiency and cost, our method indicates minimal resource consumption.