BackgroundGenomic data can advance precision medicine; however, to continue developing more targeted treatments, genomic datasets need to be integrated with health care data and become more disease-focused. This integration, in turn, amplifies existing challenges in health care data management, such as handling large data volumes, adhering to data standards, and protecting sensitive information. Addressing these challenges calls for unified digital ecosystems that combine data collection, standardization, analysis, and governance within a single platform, thereby reducing the technical burden for users. Currently, a clear set of indications about functional and nonfunctional requirements to help designers translate stakeholder needs into actionable design specifications is missing. ObjectiveThis scoping review aimed to identify the functional and nonfunctional requirements most frequently discussed in the literature from the perspective of end users (eg, clinicians and data analysts) to inform the design of a health and genomic data management platform that supports data sharing and analysis in clinical settings by conducting a PRISMA-ScR (Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Review) review. MethodsWe searched for peer-reviewed English studies that focused on platforms for managing genomic data from a user-centered perspective. We considered studies from 2014 to 2024 that were extracted from Scopus, PubMed, Web of Science, and Google Scholar for the scoping review. Insights were extrapolated for a thematic analysis to develop an initial set of requirements. We charted the functional and nonfunctional requirements according to their frequency of occurrence in the literature to provide a structured overview of the most commonly reported requirements. ResultsFrom 410 initial items, 210 items were preliminarily selected, and 53 items were included in the final analysis. Three primary groups of 26 interface functional requirements emerged: (1) general data management (acquisition, standardization, and sharing), (2) data processing and analysis (preprocessing and analysis pipelines), and (3) data visualization and reporting. Twenty nonfunctional requirements were identified and organized in 4 groups: (1) communication and support, (2) platform technical infrastructure, (3) user experience and user interface characteristics, and (4) security and compliance. We also investigated the issues that need to be resolved to develop an ideal platform. ConclusionsWe identified and mapped the most frequently reported functional and nonfunctional requirements of clinical and data professionals when discussing a health and genomic data management platform. The 3 key functional requirements should be supported by nonfunctional requirements such as secure technical infrastructure and governance mechanisms that enable compliant data processing and sharing. Designers may use these insights and mapping to develop standardized data platforms that promote efficient data exchange between institutions and experts while ensuring regulatory compliance and secure access, as proposed by the European Health Data Space.
Artificial intelligence (AI) has been recognized by the World Health Organization for its transformative potential in addressing global reproductive healthcare challenges, including inequitable access to monitoring and treatment and limited diagnostic precision. AI offers significant promise in enhancing diagnostic accuracy enabling data-driven decision-making for personalized preventive and therapeutic interventions. However, its deployment also raises ethical and operational concerns, such as data privacy risks, algorithmic bias, legal complexities, cultural sensitivity, overreliance on AI-generated recommendations, and the potential deskilling of clinicians. Addressing these challenges requires inclusive frameworks for responsible integration. Moreover, AI-driven digital transformation must align with the broader call for sustainable innovation outlined in the United Nations’ 2030 Agenda and European Union regulations defining the requirement for safe and secure integration standards. This research explores pathways for responsible and sustainable AI adoption in reproductive healthcare while mitigating associated risks. It introduces the (H)iCARE framework, which advocates for (1) human-centric hybrid models that integrate AI-driven innovations with clinical expertise to ensure balanced innovations. The framework also (2) embraces a broader, humanity-oriented perspective to ensure inclusivity beyond a limited subset of stakeholders, and (3) fosters a learning-driven approach that prioritizes continuous skill development to prevent cognitive complacency. While developed in the context of reproductive healthcare, its principles extend across the healthcare sector, providing a foundation for AI-integrated information system design and a roadmap for ethical, sustainable advancements, additionally fostering discussion on future research priorities.
Chatbots are becoming increasingly essential to information retrieval and decision support. Yet, it remains unclear how and if the (perceivable) unfairness of chatbot responses affects user experience (UX) with such tools. A pre-experimental phase involved 10 experts and 30 participants in testing a set of fair and unfair chatbot responses related to a fictional Master’s program. Six pairs of highly discriminable fair and unfair answers to a set of six questions about the Master’s program were included for the experiment. The experimental phase involved 75 participants who interacted with chatbots featuring different appearances (male, female, neutral) and programmed to provide responses to the questions about the Master’s program with varying levels of fairness, namely: completely unfair, completely fair, or partially fair (i.e., 50
Artificial Intelligence (AI) has the potential to enhance clinical decision-making and communication by supporting diagnostic reasoning, but it also raises legal, ethical, and decision-related challenges due to its ‘black box’ nature. Explainable AI (XAI) aims to address this challenge by attempting to enhance the transparency of AI decisions, and by providing user-tailored explanations that clarify how these decisions are made. The present work investigated how AI-based diagnostic insights can be presented to satisfy the informational needs of doctors and patients. A total of 58 participants (26 doctors) were asked through open questions to first describe in their own words what kind of information they would like to receive from an AI system. Then, they were presented with scenarios and asked to select which information they would like to receive selecting among four types of explanations identified in literature: contrastive, counterfactual, causal, and mechanistic explanation. The quantitative results showed that doctors were more likely to select contrastive explanations, potentially such explanations can help them distinguish between alternative diagnoses. Contrastingly, patients favored counterfactual explanations that serve them to identify alternative scenarios and possible outcomes, significantly differing from doctors’ preferences. Doctors focused on detailed, data-driven symptom information, whereas patients emphasized prognosis, risks, and treatment options. Despite these differences, both groups valued understanding symptom progression, treatment pathways, and the reasoning behind AI-driven diagnosis. The findings of this preliminary study provide actionable guidance for XAI developers, highlighting exactly which aspects of AI-assisted decision-making should be made explainable to different user groups. Tailoring explanations in this way can improve interpretability, trust, and practical utility of AI-assisted diagnostic tools in healthcare.
Rating scales are widely used in user experience (UX) research to compare and rank design alternatives, yet their development typically relies on psychometric procedures originally created for assessing stable individual traits. We argue that this practice commits the psychometric fallacy: evaluating designometric instruments—tools intended to discriminate between designs—using person × item response matrices that exclude the essential third dimension of design. Because design evaluation inherently forms a design × person × item data structure, valid scale development requires samples of designs large enough to assess how well items rank design alternatives. We show that collapsing the data cube along persons yields a design × item matrix that allows psychometric tools to be applied meaningfully, whereas collapsing along designs yields misleading psychometric results about person sensitivity rather than design discriminability. A simulation study demonstrates that scales can appear highly reliable psychometrically while being effectively unusable for ranking designs. Secondary analyses of eight commonly used UX scales across large design samples reveal systematic distortions in reliability estimates, item performance, and dimensionality when instruments are evaluated under the psychometric fallacy.We assess the impact of the psychometric fallacy for present users and designers of future rating scales and discuss more advanced methods in designometric modelling.
As artificial intelligence becomes increasingly embedded in daily decision-making processes, the need for effective communication between humans and AI systems grows more crucial. The Adaptive XAI (AXAI) workshop, now in its second edition, focuses on developing intelligent interfaces that can adaptively explain AI's decision-making processes. Building on the success of our inaugural event at IUI 2024, this workshop continues to explore the intersection of Explainable AI and adaptive user interfaces, emphasizing the development of interfaces that dynamically adapt to create explanations that resonate with diverse users. In line with the human-centric principles of the Future Artificial Intelligence Research (FAIR) project, we examine how emerging technologies such as conversational agents and Large Language Models can enhance AI explainability while ensuring explanations remain malleable and responsive to users' evolving cognitive states and contextual needs.
In Italy, where 22% of the population lives with a disability, digital accessibility remains below recommended standards. In line with that, a recent review of websites of Italian national and local public administration indicated that fewer than 12% meet basic accessibility requirements. This study aims to explore the user experience of people with disabilities across widely used categories of digital services (i.e. social media and entertainment, messaging and video calling, information search, product purchasing, banking and essential services, and booking services), as well as perceived usefulness of accessibility overlays and tools for reporting accessibility problems. A survey was conducted with 266 users (77.5% with disabilities), who were divided into three groups: people without disabilities, people with disabilities reporting high autonomy, in using digital services, and people with disabilities reporting low autonomy. Findings reveal that people with disabilities use digital services less frequently and report lower satisfaction, with UMUX-Lite scores ranging from 56.5 to 70 for low-autonomy users with disabilities, compared to scores ranging from 76.2 to 81.5 for users without disabilities. Statistical analysis also shows that engagement with essential services like banking and online shopping is significantly lower among users with disabilities. Accessibility overlays were rated as neutrally useful, with no significant differences in perceived utility across user groups, while reporting tools for accessibility issues were found to have limited impact on accessibility improvements. These findings underscore a systemic gap in accessible digital services and suggest a need for the active involvement of people with disabilities in the design process to improve both technical accessibility and overall user satisfaction.
The objective of this study is to adapt and evaluate the Turkish version of the Chatbot Usability Scale (BUS-11) through a confirmatory factor analysis method. The BUS-11 scale has been established in various languages except for Turkish; thus, its validation and dissemination could serve as a means to improve chatbot interaction satisfaction among the Turkish-speaking population and hence foster growth in Turkey’s conversational agent market. To achieve this aim, seven customer-oriented chatbots were rated on pre-designed tasks by participants. Data collection involved using the Turkish-adapted BUS11 (TBUS-11) to assess individuals’ experiences after interacting with Turkish-speaking chatbots, along with the Turkish version of the UMUX-LITE scale. Results show that TBUS-11 has been demonstrated to be highly reliable with a strong convergent validity with the UMUX-LITE already validated in Turkish. Moreover, the analysis demonstrated that the dataset supported the five-factor structure of the original version of the scale, thus confirming the psychometric properties of the TBUS. The study successfully adapted the BUS-11 into Turkish, providing a reliable and valid tool for assessing chatbot usability in the Turkish-speaking market. This can potentially enhance user satisfaction and promote the growth of conversational agents in Turkiye.
Evaluating the perceived sense of immersion is essential in virtual reality (VR), particularly in training scenarios, to ensure effective skill transfer to real-world applications. With the rise of mixed reality (MR) simulations, balancing virtuality and reality is a key design challenge. Often, comparisons between different setups along the MR continuum might be necessary to identify the most viable option, which maximizes the training potential and user experience. However, existing presence assessment models are frequently insufficient for A/B testing, particularly in MR contexts, as they fail to fully account for the entire MR continuum. This limitation was evident in our pilot study, employing a seek-and-reach task under two conditions-VR and MR-to experimentally validate a modified presence questionnaire for MR applications. The results indicated that the questionnaire failed to distinguish between VR and MR experiences, highlighting the need for a more comprehensive assessment tool. To address this gap, we propose an experimental protocol for validating a new questionnaire designed to assess the sense of immersion across the full MR continuum, based on the Congruence and Plausibility (CaP) model proposed by Latoschik and Wienrich.
In an increasingly complex everyday life, algorithms-often learned from data, i.e., machine learning (ML)-are used to make or assist with operational decisions. However, developers and designers usually are not entirely aware of how to reflect on social justice while designing ML algorithms and applications. Algorithmic social justice-i.e., designing algorithms including fairness, transparency, and accountability-aims at helping expose, counterbalance, and remedy bias and exclusion in future ML-based decision-making applications. How might we entice people to engage in more reflective practices that examine the ethical consequences of ML algorithmic bias in society? We developed and tested a design-fiction-driven methodology to enable multidisciplinary teams to perform intense, workshop-like gatherings to let potential ethical issues emerge and mitigate bias through a series of guided steps. With this contribution, we present an original and innovative use of design fiction as a method to reduce algorithmic bias in co-design activities.
This study describes an evidence-based Clinical Pathway Mapping (CPM) visualisation method that can be used to enhance healthcare quality in the face of complex systems and constrained resources. The CPM visualisation method we are proposing highlights the connection among the key components of the healthcare work system – individuals, tasks, tools, technology, physical environment, and organisational conditions. When seeking to innovate a clinical pathway by adding or changing a (technological or procedural) component, practitioners need to consider that changes in one component will affect the others, so being able to map and visualise such the relationship among the components is essential to patient safety and care delivery. We present the CPM method using a case study focused on a new Patient-Controlled Analgesia (PCA) pump in the postoperative sector. Feedback was gained through semi-structured interviews with stakeholders from four NHS Trusts in the UK. By evaluating the device's effect on post-operative procedures, the research produced a thorough picture of the postoperative environment as it exists today. We identified the two most likely scenarios of use and carried out a work system analysis to investigate these scenarios in detail and implications for the healthcare work system’s components. This analysis facilitated the identification of new criteria necessary for the device's effective integration into NHS Hospitals. The study emphasises the benefits of utilizing a human and system centred process for visualisation, identifying areas for development, and improving the security and effective application of emerging medical technologies.
Intelligent systems, such as chatbots, are likely to strike new qualities of UX that are not covered by instruments validated for legacy human–computer interaction systems. A new validated tool to evaluate the interaction quality of chatbots is the chatBot Usability Scale (BUS) composed of 11 items in five subscales. The BUS-11 was developed mainly from a psychometric perspective, focusing on ranking people by their responses and also by comparing designs’ properties (designometric). In this article, 3186 observations (BUS-11) on 44 chatbots are used to re-evaluate the inventory looking at its factorial structure, and reliability from the psychometric and designometric perspectives. We were able to identify a simpler factor structure of the scale, as previously thought. With the new structure, the psychometric and the designometric perspectives coincide, with good to excellent reliability. Moreover, we provided standardized scores to interpret the outcomes of the scale. We conclude that BUS-11 is a reliable and universal scale, meaning that it can be used to rank people and designs, whatever the purpose of the research.
As the integration of Artificial Intelligence into daily decision-making processes intensifies, the need for clear communication between humans and AI systems becomes crucial. The Adaptive XAI (AXAI) workshop focuses on the design and development of intelligent interfaces that can adaptively explain AI's decision-making processes and our engagement with those processes. In line with the human-centric principles of the Future Artificial Intelligence Research (FAIR) project1, this workshop seeks to explore, understand and develop interfaces that dynamically adapt, thereby creating explanations of AI-based systems that both relate to and resonate with a range of users with different explanation-based requirements. As AI's role in our lives becomes ever more embedded, the ways in which such systems explain elements about the system need to be malleable and responsive to the ever-evolving individual's cognitive state, relating to contextual needs/focus and to the social setting. For instance, easy to use and effective interaction modalities like Visual Languages can provide users with intuitive mechanisms to interact with, adjust, and reshape AI narratives. This ensures that a richer, more tailored understanding can be provided, allowing explanations to emerge in line with the users' demands and the ever-shifting contexts they find themselves in, both as individuals and as part of a group. The Adaptive XAI workshop extends an invitation to scholars, designers, and tech-nologists to collaboratively shape the future of human-XAI interplay.
This paper presents a human factors qualitative study on an AI application for managing sepsis in Intensive Care Units (ICUs). The study involved semi-structured interviews with nine ICU clinicians and nurses across three London hospitals. It consisted of two parts: the first applied methods to understand sepsis resuscitation processes and establish opportunities for the AI tool to mitigate gaps in the process. The second part examined adherence to AI recommendations based on factors like shift timing and user seniority, and whether shared risk in team decisions affects adherence. The findings revealed that while acknowledging the AI tool's potential benefits, participants would require a clear rationale explaining the AI results. They preferred AI suggestions that aligned with their views and did not risk patient safety, often seeking the confirmation of a colleague in uncertain situations. Overall, the study emphasised the cautious, context-dependent acceptance of AI recommendations in ICU settings. It also demonstrated the need for human factors studies to evaluate the user response to AI and its implications on decision-making.
Abstract Background Effective teamwork is crucial to providing safe and high-quality patient care, especially in acute care. Crew Resource Management (CRM) principles are often used for training teamwork in these situations, with escape rooms forming a promising new tool. However, little is known about escape room design characteristics and their effect on learning outcomes. We investigated the current status of design characteristics and their effect on learning outcomes for escape room-based CRM/teamwork training for acute care professionals. We also aimed to identify gaps in literature to guide further research. Methods Multiple databases were searched for studies describing the design and effect of escape rooms aimed training CRM/teamwork in acute care professionals and in situations that share characteristics. A standardized process was used for screening and selection. An evidence table that included study characteristics, design characteristics and effect of the escape room on learning outcomes was used to extract data. Learning outcomes were graded according to IPE expanded typology of Kirkpatrick’s levels of learning outcome and Medical Education Research Study Quality Instrument (MERSQI) scores were calculated to assess methodology. Results Fourteen studies were included. Common design characteristics were a team size of 4–6 participants, a 40-minute time limit, linear puzzle organization and use of briefing and structured debriefing. Information on alignment was only available in five studies and reporting on several other educational and escape room design characteristics was low. Twelve studies evaluated the effect of the escape room on teamwork: nine evaluated reaction (Kirkpatrick level 1; n = 9), two evaluated learning (Kirkpatrick level 2) and one evaluated both. Overall effect on teamwork was overtly positive, with little difference between studies. Together with a mean MERSQI score of 7.0, this precluded connecting specific design characteristics to the effect on learning outcomes. Conclusions There is insufficient evidence if and how design characteristics affect learning outcomes in escape rooms aimed at training CRM/teamwork in acute care professionals. Alignment of teamwork with learning goals is insufficiently reported. More complete reporting of escape rooms aimed at training CRM/teamwork in acute care professionals is needed, with a research focus on maximizing learning potential through design.
AI-based conversational agents hold significant promise for transforming educational processes, yet there is a lack of empirical research examining users' perceptions of their benefits, concerns, and the implications for design. This study investigates these perceptions among graduate and under-graduate students and educators at the University of Twente (UTwente) in the Netherlands, focusing on the potential advantages and challenges of integrating AI chatbots into education. The study aims to provide preliminary insights into design considerations for implementing AI-driven chatbots as a "buddy system" to support student learning for enhanced learning process and outcome quality. The pilot study involved 58 participants, including bachelor's and master's students from various disciplines, PhD researchers, and educators. The findings contribute to the co-creation of an AI-based StudyBuddy tool, designed to assist students in their academic journey.
Prior research indicates that chatbots have the capacity to significantly enhance learning performance, student satisfaction, and engagement. Chatbots are employed in various educational contexts, serving as content delivery platforms, facilitating student interaction, fostering collaborative learning, and promoting question-and-answer practice, among other applications. Moreover, integrating chatbots into teaching practices empowers educators to analyze and assess students’ learning abilities and comprehension levels. However, much of the existing research on educational instruments, including chatbots, lacks both theoretical support from recent advancements in the learning sciences and an evidence-informed foundation for selecting appropriate data and information models. As a consequence, educational chatbots run the risk of yielding unintended negative consequences instead of delivering the intended benefits. This study seeks to address this gap by grounding the design of educational chatbots in the principles of learning sciences. We argue that effective communication through educational chatbots necessitates formulating information in the form of feedback dialogues to enhance learners’ comprehension. Additionally, we align the design of educational chatbots with learner-centric and mindful technology concepts, inline with Industry 5.0 digitization strategies.
PurposeThis research aims to explore digital feedback needs/preferences in online education during lockdown and the implications for post-pandemic education.Design/methodology/approachAn empirical study approach was used to explore feedback needs and experiences from educational institutions in the Netherlands and Germany (N = 247) using a survey method.FindingsThe results showed that instruments supporting features for effortless interactivity are among the highly preferred options for giving/receiving feedback in online/hybrid classrooms, which are in addition also opted for post-pandemic education. The analysis also showed that, when communicating feedback digitally, more inclusive formats are preferred, e.g. informing learners about how they perform compared to peers. The increased need for comparative performance-oriented feedback, however, may affect students' goal orientations. In general, the results of this study suggest that while interactivity features of online instruments are key to ensuring social presence when using digital forms of feedback, balancing online with offline approaches should be recommended.Originality/valueThis research contributes to the gap in the scientific literature on feedback digitalization. Most of the existing research are in the domain of automated feedback generated by various learning environments, while literature on digital feedback in online classrooms, e.g. empirical studies on preferences for typology, formats and communication channels for digital feedback, to the best of the authors’ knowledge is largely lacking. The findings and recommendations of this study extend their relevance to post-pandemic education for which hybrid classroom is opted among the highly preferred formats by survey respondents.
Giuseppe Liotta合作论文数Computer Science4