AI is now embedded in healthcare, finance, policy, and many other domains, yet genuine human-AI synergy - combined performance that exceeds what either party achieves alone - is uncommon. Meta-analyses show that AI assistance tends to improve human performance compared to working alone, but studies finding true synergy are scarce. We call this persistent shortfall the synergy gap. Most current work treats human-AI combination as an engineering problem and concentrates on interpretability, trust calibration, or interface design. These matter, but they cover only part of what determines whether combination works. Closing the synergy gap, we argue, requires explicit engagement with a wider design space. We map that space through six interconnected elements: sociotechnical context, decision-making frameworks, human decision participants, AI capabilities, interaction, and holistic evaluation. For each element, we describe what it covers, how it shapes the others in practice, and what it implies for design. The result is a shared vocabulary for practitioners building hybrid systems, an analytical lens for researchers studying combination patterns, and a starting point for evaluators interested in the full quality of human-AI decision-making rather than accuracy alone.
Over-reliance on AI systems in high-stakes decision-making leads to avoidable errors and raises concerns about human oversight, responsibility, and bias. This paper proposes a theoretical framework for understanding over-reliance as a biased form of reasoning under uncertainty. We argue that decision-makers may systematically ignore relevant prior information when assessing AI suggestions, thus conflating the system’s reliability for the probability of reaching a correct decision. This “base rate neglect” fallacy violates sound Bayesian reasoning and is well-documented in empirical studies from cognitive science. By reanalysing experimental data from a study on forensic decision-making, we show that biased reasoning leading to over-reliance may occur on a case-by-case basis rather than as a stable tendency of the decision-maker. An interesting implication of our framework is that improving AI accuracy alone may worsen rather than reduce over-reliance. Looking for alternative solutions, we draw on the cognitive science literature to present an interface prototype that supports Bayesian reasoning by helping users understand uncertainty, properly combine different pieces of evidence, and rationally adjust their opinions in AI-assisted evaluative judgments.
As AI systems increasingly permeate domains where human understanding, trust, and accountability are paramount, the demand for explainable and ethically aligned user interfaces has become urgent. In this paper, we propose applying the Transmeta Design framework—originally developed to enhance usability through transparency and meta-communication in Web services—to the development of AI-driven interactive systems. We argue that Transmeta Design provides a promising methodological basis for engineering explainable AI, user-centred decision support systems, and ethically coherent interfaces. By leveraging its dual dimensions of transparency and meta-communication, we articulate a methodology that aligns with the goals of human-centred AI. We demonstrate its application with two case studies: an AI-assisted clinical decision support system and an intelligent financial planning assistant.
Artificial Intelligence is driving society toward an increasingly algorithmic future, advancing innovation through data-driven predictive analytics. At the core of this transformation lie advanced Machine Learning-based systems, which hold significant societal potential but often remain inaccessible to domain specialists due to their complexity. This reliance on computing experts limits the integration of domain knowledge and raises concerns about transparency and inclusivity. Drawing from research in HCI, Computer-Supported Cooperative Work, and Cognitive Load Theory, this work explores how Visual Programming Languages (VPLs) and touch-based interfaces can support novice collaboration in ML-based system design. We evaluate PyFlowML Touch, a touch-enabled extension of the previously developed PyFlowML system, which enhances node layout and introduces interactive feedback mechanisms tailored to co-located novice collaboration. Designed for domain specialists, PyFlowML guides users through the ML process and incorporates Explainable AI techniques to improve understanding of model behavior [81, 82]. In this exploratory study, a multifaceted evaluation combining cognitive walkthrough and heuristic inspection yields promising insights, suggesting that the touch-enabled interface can reduce intrinsic and extraneous cognitive load while promoting schema construction and collective ML understanding through interactive visualizations.
As artificial intelligence becomes increasingly embedded in daily decision-making processes, the need for effective communication between humans and AI systems grows more crucial. The Adaptive XAI (AXAI) workshop, now in its second edition, focuses on developing intelligent interfaces that can adaptively explain AI's decision-making processes. Building on the success of our inaugural event at IUI 2024, this workshop continues to explore the intersection of Explainable AI and adaptive user interfaces, emphasizing the development of interfaces that dynamically adapt to create explanations that resonate with diverse users. In line with the human-centric principles of the Future Artificial Intelligence Research (FAIR) project, we examine how emerging technologies such as conversational agents and Large Language Models can enhance AI explainability while ensuring explanations remain malleable and responsive to users' evolving cognitive states and contextual needs.
Context and motivation: End-user development focuses on enabling non-professional programmers to create or extend software applications on their own. However, before beginning the development process, software engineering best practices recommend performing requirements engineering (RE) activities, including requirements modelling. Question/problem: There is limited research on how end-users can model system requirements. Principal ideas/results: In this experience report, we investigate the problem of end-user requirements modelling in an EU-funded project about agricultural digitalisation. Specifically, a team of agronomists was directly involved in the creation of UML, iStar, and BPMN diagrams to model the transformation of socio-technical processes in four different concrete scenarios. They followed a formalisation procedure proposed within an RE method designed to help stakeholders evaluate the impact of agricultural digitalisation. Starting from textual reports including a description of the process as-is and the process-to-be, they followed step-by-step guidelines for model creation. Contribution: This paper reports insights from the experience from the viewpoint of the agronomists and software engineers involved. We identify nine key lessons that highlight the added value of end-user requirements modelling for achieving a shared and in-depth understanding of the socio-technical processes under analysis.
Diagrams can be valuable tools in requirements engineering to establish a shared understanding between software engineers and stakeholders. However, interacting with these visual representations can be challenging for some stakeholders who prefer textual descriptions and may need support to interpret notation elements and understand the diagram structure and meaning. To address this need, we explore the use of Large Language Models to effectively assist stakeholders interacting with diagrams by providing automatic textual explanations and contextual guidance. Specifically, we aim to design and evaluate with stakeholders an interactive layer (integrated into an end-user-oriented modelling tool) that provides automatic diagram explanations in natural language. As a first step toward our research objective, this paper investigates the capability of GPT4 to generate appropriate textual descriptions from domain models. We use a test data set consisting of UML class diagrams in various formats, belonging to the domain of digital agriculture, and develop a set of prompts to generate the interactive explanatory layer. We conduct a technical evaluation of the output, focusing on correctness, completeness, and understandability. The results provide valuable insights to inform future design and research, while also revealing potential challenges in real-world applications.
Computational thinking (CT) skills provide structured approaches to problem-solving that are valuable for navigating the increasing complexity of technological environments. CT skills can be assessed through various methods and perspectives. EUDability provides a framework for evaluating end-user development (EUD) tools, with core dimensions directly aligned to CT skills. This paper explores how CT skills manifest in the creation of visual models, an activity that supports the representation and understanding of socio-technical systems. We propose an evaluation method employing ModeLLer, a block-based EUD modelling tool and a user study. We carry out a structured evaluation integrating the EUDability inspection and process-based CT skills evaluation with the assessment of the artefacts produced by end-users. Results highlight the ability of ModeLLer to support the modelling activity while also providing insights into the EUDability of the tool and end-users' CT skills. Furthermore, although preliminary, our results illustrate the relationship between tool capabilities, user skills, and modelling outcomes.
This paper presents a critical perspective on the ecological validity challenges in evaluating AI-assisted decision-making tools for healthcare, illustrated through insights from a case study on oral cancer diagnosis. We argue that current experimental approaches often fail to capture the complexities of clinical environments in three critical dimensions: the temporal dynamics of decision-making, the holistic nature of clinical reasoning, and the multifaceted requirements for performance evaluation. Our case study with ten dental care specialists of varying experience levels revealed significant misalignments between our controlled experimental design and the realities of clinical practice. Participants’ qualitative feedback highlighted how real-world diagnosis involves contextual information beyond images, follows different temporal patterns than rapid experimental tasks, and requires evaluation metrics beyond simple accuracy. Based on these observations, we suggest pathways for enhancing ecological validity in AI healthcare research: incorporating longitudinal evaluation approaches, designing systems that integrate multiple information streams, and developing nuanced performance metrics that reflect clinical priorities. This work contributes to the ongoing dialogue about bridging the gap between AI research and its practical implementation in high-stakes medical settings.
Digital technologies are transforming agriculture, affecting social, institutional, economic, environmental, and technological dimensions. To ensure sustainable development, it is essential to anticipate these impacts and create conditions for sustainable change. Living Labs (LLs) concept facilitates this by involving various stakeholders in co-designing solutions. This paper presents a socio-technical process modelling method using Model-driven requirements engineering (MoDRE) techniques. It employs UML class diagrams, iStar diagrams, and BPMN diagrams to model process structures, goals, and flows. The method, part of the Horizon Europe project CODECS, involves data collection, diagram design, and iterative feedback, tested in a precision irrigation pilot study in Tuscany. Preliminary results demonstrate the method's effectiveness in supporting interdisciplinary teams, fostering better communication, and aiding in the analysis of digitalisation impacts on agricultural processes. Furthermore, the discussion with stakeholders allowed the fine-tuning of the models and enriched the method for co-creating the diagrams with a toolkit composed of a set of guidelines for eliciting process-relevant information from LLs, a checklist and a detailed procedure for graphical representation.
This paper explores the synergistic integration of Artificial Intelligence (AI) into telemedicine systems, emphasizing the critical roles of co-design and explainability aspects. As telemedicine evolves to offer more sophisticated, personalized, and accessible healthcare services, the incorporation of AI presents unique challenges and opportunities for developing future telemedicine solutions that will be effective, transparent, and trusted by both healthcare providers and patients through co-design to increase explainability. Overall, we aim at emphasizing the crucial role of co-design in bridging the gap between AI, explainability, and trustworthiness of AI-based systems.
Rosa Lanzilotti合作论文数Department of Computer Science, University of Bari5
Paolo Bottoni合作论文数Department of Computer Science, Sapienza University of Rome4