AI systems are increasingly used in high-stakes domains such as credit rating, where fairness concerns are critical. Existing fairness assessments are typically conducted by AI experts or regulators using predefined protected attributes and metrics, which often fail to capture the diversity and nuance of fairness notions held by the individuals who are affected by these systems' decisions, such as decision subjects. Recent work has therefore called for involving affected individuals in fairness assessment, yet little empirical evidence exists on how they create their own fairness criteria or what kinds of criteria they produce - knowledge that could not only inform experts' fairness evaluation and mitigation, but also guide the design of AI assessment tools. We address this gap through a qualitative user study with 18 participants in a credit rating scenario. Participants first articulated their fairness notions in their own words. Then, participants turned them into concrete quantified and operationalized fairness criteria, through an interactive prototype we designed. Our findings provide empirical evidence of the process through which people's fairness notions emerge via grounding in model features, and uncover a diverse set of individuals' custom-defined criteria for both outcome and procedural fairness. We provide design implications for processes and tools that support more inclusive and value-sensitive AI fairness assessment.
Explanations from AI systems can illuminate, yet they can misguide. This half-day MIRAGE workshop at IUI 2026 confronts the Explainability Pitfalls and Dark Patterns embedded in AI-generated explanations. Evidence now shows that explanations may inflate unwarranted trust, warp mental models, and obscure power asymmetries—even when designers intend no harm. We convene an interdisciplinary group of researchers and practitioners to define, detect, and defuse these hazards. By shifting the focus from making explanations to making explanations safe, MIRAGE propels the community toward an accountable, human-centered AI future.
Persuasive technologies help users meet various types of behavioural goals, and are becoming increasingly personalised to individual users' needs. Research has been siloed into single domains of behaviour change, such as health or environmental behaviours, leaving open questions on how users across domains respond to personalised interventions. To fill this gap, we reviewed 56 publications between 2013 and 2024, which developed personalised persuasive technologies. We analysed how the technologies were designed and how users evaluated them. We found that persuasive technologies build user profiles that are either static or dynamic, and use this information to deliver three types of personalised interventions: personalised goals, personalised messages, or personalised timing of reminders. Personalised technologies were more effective than one-size-fits-all and utilised a combination of behaviour change techniques. Users not only evaluated personalisation positively, but wanted to know how it was achieved. In addition, users preferred goals that were easy to meet, and appreciated empathetic support from persuasive technology when a goal was not met. Based on these findings, we discuss methodological, theoretical, and practical implications, as well as future research directions for personalised persuasive technologies.
As AI deployment accelerates globally, the urgency for responsible AI(RAI) frameworks has catalysed unprecedented policy initiatives and research investments worldwide. The CHI community, positioned uniquely at the intersection of technology, people, society, and values, has a crucial role in shaping RAI development. This meet-up brings together international HCI researchers working across value-sensitive design, human-centered AI, ethics, explainability, trustworthy AI, sustainability and equitability to foster collaborative dialogue on CHI's perspectives and potential in RAI. Through an interactive town hall format featuring community standups, democratic topic selection, and structured roundtable discussions, we will generate actionable outcomes including network building, knowledge sharing, and concrete collaboration opportunities for the CHI community around RAI. This meetup would serve as a foundational response from the CHI community to establish priorities, forge connections, and catalyse the next generation of RAI research that centers human values, promotes fairness, and advances democratic principles in an era of AI transformation.
People are increasingly turning to generative AI (e.g., ChatGPT, Gemini, Copilot) for emotional support and companionship. While trust is likely to play a central role in enabling these informal and unsupervised interactions, we still lack an understanding of how people develop and experience it in this context. Seeking to fill this gap, we recruited 24 frequent users of generative AI for emotional support and conducted a qualitative study consisting of diary entries about interactions, transcripts of chats with AI, and in-depth interviews. Our results suggest important novel drivers of trust in this context: familiarity emerging from personalisation, nuanced mental models of generative AI, and awareness of people's control over conversations. Notably, generative AI's homogeneous use of personalised, positive, and persuasive language appears to promote some of these trust-building factors. However, this also seems to discourage other trust-related behaviours, such as remembering that generative AI is a machine trained to converse in human language. We present implications for future research that are likely to become critical as the use of generative AI for emotional support increasingly overlaps with therapeutic work.
Human testers-often end-users themselves-can judge system errors in context and reveal failures that automated methods miss. Inspired by software testing practices, we conducted an exploratory study in which 15 participants tested a satellite image classifier using an interactive tool that allowed them to collect data and create test cases. While participants shared a common understanding of the model's behavior and generally adopted a failure-driven approach, we observed significant variability in their testing behaviors, including the number of test cases created, the timing of seeking feedback, the distribution of effort across classes, and the types of failures identified. Although links between specific strategies and outcomes remain unclear, our findings provide a first step toward understanding human testing of ML models and inform future research on human-driven AI auditing.
Artificial intelligence (AI) applications have become ubiquitous in their impact on individuals and society, highlighting a crucial need for their responsible development. Recent research has called for participatory AI auditing, empowering individuals without AI expertise to audit AI applications throughout the entire AI development pipeline. Our work focuses on investigating how to support these kinds of auditors through participatory AI auditing tools and processes. We conducted a series of co-design workshops, using two health-related predictive AI applications as examples. Our results show that participants wanted to be part of AI audits, and were insightful in identifying the potential impacts of applications, but needed to be assisted in conducting audits, especially how to measure impacts. Importantly, participants provided examples of impacts not considered in current risk/harm taxonomies. Our findings provide implications for the design of tools and processes to empower everyone to contribute to responsible AI development in the future.
Assessing fairness in artificial intelligence (AI) typically involves AI experts who select protected features, fairness metrics, and set fairness thresholds to assess outcome fairness. However, little is known about how stakeholders, particularly those affected by AI outcomes but lacking AI expertise, assess fairness. To address this gap, we conducted a qualitative study with 26 stakeholders without AI expertise, representing potential decision subjects in a credit rating scenario, to examine how they assess fairness when placed in the role of deciding on features with priority, metrics, and thresholds. We reveal that stakeholders' fairness decisions are more complex than typical AI expert practices: they considered features far beyond legally protected features, tailored metrics for specific contexts, set diverse yet stricter fairness thresholds, and even preferred designing customized fairness. Our results extend the understanding of how stakeholders can meaningfully contribute to AI fairness governance and mitigation, underscoring the importance of incorporating stakeholders' nuanced fairness judgments.
People are increasingly using generative artificial intelligence (AI) for emotional support, creating trust-based interactions with limited predictability and transparency. We address the fragmented nature of research on trust in AI through a multidisciplinary conceptual review, examining theoretical foundations for understanding trust in the emerging context of emotional support from generative AI. Through an in-depth literature search across human-computer interaction, computer-mediated communication, social psychology, mental health, economics, sociology, philosophy, and science and technology studies, we developed two principal contributions. First, we summarise relevant definitions of trust across disciplines. Second, based on our first contribution, we define trust in the context of emotional support provided by AI and present a categorisation of relevant concepts that recur across well-established research areas. Our work equips researchers with a map for navigating the literature and formulating hypotheses about AI-based mental health support, as well as important theoretical, methodological, and practical implications for advancing research in this area.
Achieving a sustainable future requires behaviour change on a large scale. One possible approach is to improve persuasive technology by mirroring each users' personality, but research on this is very limited. In this work, we explore whether the Big Five personality traits underpin differences in persuasive text for sustainability, both in terms of content and linguistics. We also investigate whether personality scores can be reliably predicted from persuasive text, using machine learning (ML) techniques. Our results show that personality traits appear to influence some aspects of the content, but not linguistics; however, predicting personality is more successful through linguistics than content. We provide a follow-on analysis of which features are most informative to predict personality scores. Based on these results, we provide recommendations for automatic personality recognition and synthesis in persuasive technologies.
While gaze data brings benefits like allowing hands-free interaction, it can also reveal sensitive information about people, such as their gender, age, and geographical origin. Privacy leakage and safeguards have been explored for gaze data collected through headsets and stationary eye trackers, but never for gaze data collected through handheld mobile devices, like smartphones. Eye tracking on handheld mobile devices has the potential to be ubiquitous, but gaze data is typically of lower quality, compounded by additional noise and instability due to less controlled environments and screen size constraints. To address this gap, we provide the first evidence of privacy leakage through gaze data collected on handheld mobile devices. In a user study (N=35), we collected our novel SmartEyePhone dataset of gaze data using a smartphone’s front-facing camera. Second, we present the first evaluation and comparison of 3 Differential Privacy (DP) techniques against our dataset and the TüEyeQ dataset, which was collected in prior work using a stationary remote eye tracker. We found that SmartEyePhone dataset leaks on average 65.5% of private data. DP mechanisms reduce privacy leakage in our data by 22.33% using Laplace mechanism, 10.43% using Exponential mechanism, 22.60% using Gaussian mechanism, and 28.34% using AI model perturbation. However, this also reduces the accuracy of the main task prediction, and is impacted by the choice of privacy parameter values. We present insights on the trade-off between privacy preservation and the practical usefulness of gaze data. Our insights advance the understanding of privacy in mobile settings and pave the way for privacy preserving gaze-enabled handheld mobile devices.
Representation bias is one of the most common types of biases in artificial intelligence (AI) systems, causing AI models to perform poorly on underrepresented data segments. Although AI practitioners use various methods to reduce representation bias, their effectiveness is often constrained by insufficient domain knowledge in the debiasing process. To address this gap, this paper introduces a set of generic design guidelines for effectively involving domain experts in representation debiasing. We instantiated our proposed guidelines in a healthcare-focused application and evaluated them through a comprehensive mixed-methods user study with 35 healthcare experts. Our findings show that involving domain experts can reduce representation bias without compromising model accuracy. Based on our findings, we also offer recommendations for developers to build robust debiasing systems guided by our generic design guidelines, ensuring more effective inclusion of domain experts in the debiasing process.
Many AI technologies are now being integrated into everyday life. However, how can we ensure that this AI is 'responsible'? In this keynote, current efforts at developing responsible AI, focusing on explanations, fairness and accountability, are reviewed. I will offer suggestions at how we can improve engineering approaches in this area.
Numerous fairness metrics have been proposed and employed by artificial intelligence (AI) experts to quantitatively measure bias and define fairness in AI models. Recognizing the need to accommodate stakeholders' diverse fairness understandings, efforts are underway to solicit their input. However, conveying AI fairness metrics to stakeholders without AI expertise, capturing their personal preferences, and seeking a collective consensus remain challenging and underexplored. To bridge this gap, we propose a new framework, EARN ( Explain, Ask, Review, and Negotiate ) Fairness, which facilitates collective metric decisions among stakeholders without requiring AI expertise. The framework features an adaptable interactive system and a stakeholder-centered EARN Fairness process to Explain fairness metrics, Ask stakeholders' personal metric preferences, Review metrics collectively, and Negotiate a consensus on metric selection. To gather empirical results, we applied the framework to a credit rating scenario and conducted a user study involving 18 decision subjects without AI knowledge. We elicited their personal metric preferences and subsequently we studied how they reached metric consensus in team sessions. Our work shows that the EARN Fairness framework supports stakeholders to express and negotiate fairness preferences, and we provide practical guidance for implementing human-centered AI fairness in high-risk contexts. Through this approach, we aim to reach consensus of fairness perspectives, fostering more equitable and inclusive AI fairness.
As Artificial Intelligence (AI) becomes increasingly integrated into high-stakes domains like healthcare, effective collaboration between healthcare experts and AI systems is critical. Data-centric steering, which involves fine-tuning prediction models by improving training data quality, plays a key role in this process. However, little research has explored how varying levels of user control affect healthcare experts during data-centric steering. We address this gap by examining manual and automated steering approaches through a between-subjects, mixed-methods user study with 74 healthcare experts. Our findings show that manual steering, which grants direct control over training data, significantly improves model performance while maintaining trust and system understandability. Based on these findings, we propose design implications for a hybrid steering system that combines manual and automated approaches to increase user involvement during human-AI collaboration.
Explanations in interactive machine-learning systems facilitate debugging and improving prediction models. However, the effectiveness of various global model-centric and data-centric explanations in aiding domain experts to detect and resolve potential data issues for model improvement remains unexplored. This research investigates the influence of data-centric and model-centric global explanations in systems that support healthcare experts in optimising models through automated and manual data configurations. We conducted quantitative (n=70) and qualitative (n=30) studies with healthcare experts to explore the impact of different explanations on trust, understandability and model improvement. Our results reveal the insufficiency of global model-centric explanations for guiding users during data configuration. Although data-centric explanations enhanced understanding of post-configuration system changes, a hybrid fusion of both explanation types demonstrated the highest effectiveness. Based on our study results, we also present design implications for effective explanation-driven interactive machine-learning systems.
With the increasing adoption of Artificial Intelligence (AI) systems in high-stake domains, such as healthcare, effective collaboration between domain experts and AI is imperative. To facilitate effective collaboration between domain experts and AI systems, we introduce an Explanatory Model Steering system that allows domain experts to steer prediction models using their domain knowledge. The system includes an explanation dashboard that combines different types of data-centric and model-centric explanations and allows prediction models to be steered through manual and automated data configuration approaches. It allows domain experts to apply their prior knowledge for configuring the underlying training data and refining prediction models. Additionally, our model steering system has been evaluated for a healthcare-focused scenario with 174 healthcare experts through three extensive user studies. Our findings highlight the importance of involving domain experts during model steering, ultimately leading to improved human-AI collaboration.
Objectives In response to the lack of digital support for older people to plan their lives for quality of life, research was undertaken to co-design and then evaluate a new digital tool that combined interactive guidance for life planning with a computerised model of quality of life. Method First, a workshop-based process for co-designing the SCAMPI tool with older people is reported. A first version of this tool was then evaluated over eight consecutive weeks by nine older people living in their own homes. Four of these people were living with Parkinson's disease, one with early-stage dementia, and four without any diagnosed chronic condition. Regular semi-structured interviews were undertaken with each individual older person and, where wanted, their life partner. A more in-depth exit interview was conducted at the end of the period of tool use. Themes arising from analyses of content from these interviews were combined with first-hand data collected from the tool's use to develop a description of how each older person used the tool over the 8 weeks. Results The findings provided the first evidence that the co-designed tool, and in particular the computerised model, could offer some value to older people. Although some struggled to use the tool as it was designed, which led to limited uptake of the tool's suggestions, the older people reported factoring these suggestions into their longer-term planning, as health and/or circumstances might change. Conclusions The article contributes to the evolving discussion about how to deploy such digital technologies to support quality of life more effectively.
Jon Bird合作论文数Dept. of Computing
The Open University4
Richard Butterworth合作论文数Bridgeman Art Library3