Recent successes in artificial intelligence (AI) have ignited debate over its role in human-machine systems—specifically, whether an AI system should be viewed as a tool or a teammate. This article consists of a set of essays that explore this question from a macrocognitive viewpoint. These essays reveal similarities across stances, and divergences regarding interpretations of the teammate metaphor. Discussions include concerns about potential risks and effects of the teammate metaphor on users, in addition to expressions of the value of the metaphor for work system design. The essays highlight the role of metaphors in both enhancing and stifling design creativity, the blending of metaphors in design, and the need for empirical evaluation to guide clear design choices. Moreover, regardless of which approach is adopted, the “technology-first” mindset of system developers, the emphasis on relatively simple human-machine interactions over complex, evolving work systems, and the insidious effects of AI on human expertise all challenge progress in the design of macrocognitive work systems. The essays present some specific research directions.
This article describes the main contributions made by the late Paul J. Feltovich to the fields of cognitive engineering and decision making.
The quest for an adequate model of trust in AI systems has led to rosters of influencing factors and proposals for how trust can be "designed-in." An alternative is to consider trusting as an emergent. While progressive maturation of user trust may be desired, it is the exception. Trust can arise rapidly, but it can diminish in a heartbeat. It can be orthogonal to reliance. Trusting and relying are context-contingent. Work-centered design principles can make trustability achievable, but it can only be realized through user experience. And principles for "designing-in" trust do not span the creativity gap.
Introduction Many Explainable AI (XAI) systems provide explanations that are just clues or hints about the computational models-Such things as feature lists, decision trees, or saliency images. However, a user might want answers to deeper questions such as How does it work?, Why did it do that instead of something else? What things can it get wrong? How might XAI system developers evaluate existing XAI systems with regard to the depth of support they provide for the user's sensemaking? How might XAI system developers shape new XAI systems so as to support the user's sensemaking? What might be a useful conceptual terminology to assist developers in approaching this challenge? Method Based on cognitive theory, a scale was developed reflecting depth of explanation, that is, the degree to which explanations support the user's sensemaking. The seven levels of this scale form the Explanation Scorecard. Results and discussion The Scorecard was utilized in an analysis of recent literature, showing that many systems still present low-level explanations. The Scorecard can be used by developers to conceptualize how they might extend their machine-generated explanations to support the user in developing a mental model that instills appropriate trust and reliance. The article concludes with recommendations for how XAI systems can be improved with regard to the cognitive considerations, and recommendations regarding the manner in which results on the evaluation of XAI systems are reported.
If a user is presented an AI system that portends to explain how it works, how do we know whether the explanation works and the user has achieved a pragmatic understanding of the AI? This question entails some key concepts of measurement such as explanation goodness and trust. We present methods for enabling developers and researchers to: (1) Assess the a priori goodness of explanations, (2) Assess users' satisfaction with explanations, (3) Reveal user's mental model of an AI system, (4) Assess user's curiosity or need for explanations, (5) Assess whether the user's trust and reliance on the AI are appropriate, and finally, (6) Assess how the human-XAI work system performs. The methods we present derive from our integration of extensive research literatures and our own psychometric evaluations. We point to the previous research that led to the measurement scales which we aggregated and tailored specifically for the XAI context. Scales are presented in sufficient detail to enable their use by XAI researchers. For Mental Model assessment and Work System Performance, XAI researchers have choices. We point to a number of methods, expressed in terms of methods' strengths and weaknesses, and pertinent measurement issues.
When people make plausibility judgments about an assertion, an event, or a piece of evidence, they are gauging whether it makes sense that the event could transpire as it did. Therefore, we can treat plausibility judgments as a part of sensemaking. In this paper, we review the research literature, presenting the different ways that plausibility has been defined and measured. Then we describe the naturalistic research that allowed us to model how plausibility judgments are engaged during the sensemaking process. The model is based on an analysis of 23 cases in which people tried to make sense of complex situations. The model describes the user's attempts to construct a narrative as a state transition string, relying on plausibility judgments for each transition point. The model has implications for measurement and for training.
The development of AI systems represents a significant investment of funds and time. Assessment is necessary in order to determine whether that investment has paid off. Empirical evaluation of systems in which humans and AI systems act interdependently to accomplish tasks must provide convincing empirical evidence that the work system is learnable and that the technology is usable and useful. We argue that the assessment of human–AI (HAI) systems must be effective but must also be efficient. Bench testing of a prototype of an HAI system cannot require extensive series of large‐scale experiments with complex designs. Some of the constraints that are imposed in traditional laboratory research just are not appropriate for the empirical evaluation of HAI systems. We present requirements for avoiding “unnecessary rigor.” They cover study design, research methods, statistical analyses, and online experimentation. These should be applicable to all research intended to evaluate the effectiveness of HAI systems.
This paper summarizes the psychological insights and related design challenges that have emerged in the field of Explainable AI (XAI). This summary is organized as a set of principles, some of which have recently been instantiated in XAI research. The primary aspects of implementation to which the principles refer are the design and evaluation stages of XAI system development, that is, principles concerning the design of explanations and the design of experiments for evaluating the performance of XAI systems. The principles can serve as guidance, to ensure that AI systems are human-centered and effectively assist people in solving difficult problems.
This expert panel is the first of a two-panel series marking the 40 th anniversary of “Cognitive Systems Engineering: New Wine in New Bottles” by Hollnagel and Woods (1983) and, arguably, the beginning of Cognitive Systems Engineering (CSE). These experts were there at (or near) the beginning, devising new methods, expanding and creating new theories, and revealing a new perspective on how complex systems sustain performance and fail. They also wrestled and struggled with these new ideas to propose and implement solutions to improve performance in a number of high-consequence industries. Whether in graduate school or as early-career professionals, they saw the surprises that served as signals that the thinking that brought us to that point would not, alone, be the thinking and doing that would take us further. They will each answer the question, “What ideas and perspectives are important about Cognitive Systems Engineering, and why?”
Demands to manage the risks of artificial intelligence (AI) are growing. These demands and the government standards arising from them both call for trustworthy AI. In response, we adopt a convergent approach to review, evaluate, and synthesize research on the trust and trustworthiness of AI in the environmental sciences and propose a research agenda. Evidential and conceptual histories of research on trust and trustworthiness reveal persisting ambiguities and measurement shortcomings related to inconsistent attention to the contextual and social dependencies and dynamics of trust. Potentially underappreciated in the development of trustworthy AI for environmental sciences is the importance of engaging AI users and other stakeholders, which human-AI teaming perspectives on AI development similarly underscore. Co-development strategies may also help reconcile efforts to develop performance-based trustworthiness standards with dynamic and contextual notions of trust. We illustrate the importance of these themes with applied examples and show how insights from research on trust and the communication of risk and uncertainty can help advance the understanding of trust and trustworthiness of AI in the environmental sciences.
Introduction: The purpose of the Stakeholder Playbook is to enable the developers of explainable AI systems to take into account the different ways in which different stakeholders or role-holders need to "look inside" the AI/XAI systems.Method: We conducted structured cognitive interviews with senior and mid-career professionals who had direct experience either developing or using AI and/or autonomous systems.Results: The results show that role-holders need access to others (e.g., trusted engineers and trusted vendors) for them to be able to develop satisfying mental models of AI systems. They need to know how it fails and misleads as much as they need to know how it works. Some stakeholders need to develop an understanding that enables them to explain the AI to someone else and not just satisfy their own sense-making requirements. Only about half of our interviewees said they always wanted explanations or even needed better explanations than the ones that were provided. Based on our empirical evidence, we created a "Playbook" that lists explanation desires, explanation challenges, and explanation cautions for a variety of stakeholder groups and roles.Discussion: This and other findings seem surprising, if not paradoxical, but they can be resolved by acknowledging that different role-holders have differing skill sets and have different sense-making desires. Individuals often serve in multiple roles and, therefore, can have different immediate goals. The goal of the Playbook is to help XAI developers by guiding the development process and creating explanations that support the different roles.
The special issue on Explainable Artificial Intelligence (XAI) provides a representative snapshot of the state of the art in the 2020-2021 time-frame and highlights future research directions. The scope of the special issue is intentionally broad, ranging from technical contributions to human-centered studies, from surveys to philosophical perspectives. A total of 97 papers were submitted for the special issue, of which 27 were finally accepted for publication after thorough peer review.
A challenge in building useful artificial intelligence (AI) systems is that people need to understand how they work in order to achieve appropriate trust and reliance. This has become a topic of considerable interest, manifested as a surge of research on Explainable AI (XAI). Much of the research assumes a model in which the AI automatically generates an explanation and presents it to the user, whose understanding of the explanation leads to better performance. Psychological research on explanatory reasoning shows that this is a limited model. The design of XAI systems must be fully informed by a model of cognition and a model of pedagogy, based on empirical evidence of what happens when people try to explain complex systems to other people and what happens as people try to reason out how a complex system works. In this article we discuss how and why C. S. Peirce's notion of abduction is a best model for XAI. Peirce's notion of abduction as an exploratory activity can be regarded as supported by virtue of its concordance with models of expert reasoning that have been developed by modern applied cognitive psychologists.
We reflect on the progress in the area of Explainable AI (XAI) Program relative to previous work in the area of intelligent tutoring systems (ITS). A great deal was learned about explanation—and many challenges uncovered—in research that is directly relevant to XAI. We suggest opportunities for future XAI research deriving from ITS methods, as well as the challenges shared by both ITS and XAI in using AI to assist people in solving difficult problems effectively and efficiently.
The process of explaining something to another person is more than offering a statement. Explaining means taking the perspective and knowledge of the Learner into account and determining whether the Learner is satisfied. While the nature of explanation-conceived of as a set of statements-has been explored philosophically and empirically, the process of explaining, as an activity, has received less attention. We conducted an archival study, looking at 73 cases of explaining. We were particularly interested in cases in which the explanations focused on the workings of complex systems or technologies. The results generated two models: local explaining to address why a device (such an intelligent system) acted in a surprising way, and global explaining about how a device works. The examination of the processes of explaining as it occurs in natural settings revealed a number of mistaken beliefs about how explaining happens, and what constitutes an explanation that encourages learning.
The Cognitive Tutorial concept is based on the view that the genuine cognitive challenges to forming functional and accurate mental models of AI systems can be formalized, documented, and "trained in." Its purpose is to serve as a means of global explanation of an AI or machine learning system. A Cognitive Tutorial is created specifically to accelerate proficiency at learning to use intelligent software tools. Therefore, it would be a valuable addition to any "toolkit" for ensuring that intelligent systems are explainable, are adequately explained to users, and the users are satisfied with their understanding of the system. This Report describes the procedures for creating a Cognitive Tutorial, the modules that comprise a Cognitive Tutorial, and example Cognitive Tutorials applied to two AI systems.
This report describes a Self-Explaining Scorecard for appraising the self-explanatory support capabilities of XAI systems. The Scorecard might be useful in conceptualizing the various ways in which XAI system developers are supporting users, and might also help in comparing and contrasting the various approaches.
The purpose of the Stakeholder Playbook is to enable system developers to take into account the different ways in which stakeholders need to "look inside" of the AI/XAI systems. Recent work on Explainable AI has mapped stakeholder categories onto explanation requirements. While most of these mappings seem reasonable, they have been largely speculative. We investigated these matters empirically. We conducted interviews with senior and mid-career professionals possessing post-graduate degrees who had experience with AI and/ or autonomous systems, and who had served in a number of roles including former military, civilian scientists working for the government, scientists working in the private sector, and scientists working as independent consultants. The results show that stakeholders need access to others (e.g., trusted engineers, trusted vendors) to develop satisfying mental models of AI systems. and they need to know "how it fails" and "how it misleads" and not just "how it works." In addition, explanations need to support end-users in performing troubleshooting and maintenance activities, especially as operational situations and input data change. End-users need to be able to anticipate when the AI is approaching an edge case. Stakeholders often need to develop an understanding that enables them to explain the AI to someone else and not just satisfy their own sensemaking. We were surprised that only about half of our Interviewees said they always needed better explanations. This and other findings that are apparently paradoxical can be resolved by acknowledging that different stakeholders have different capabilities, different sensemaking requirements, and different immediate goals. In fact, the concept of “stakeholder” is misleading because the people we interviewed served in a variety of roles simultaneously — we recommend referring to these roles rather than trying to pigeonhole people into unitary categories. Different cognitive styles re another formative factor, as suggested by participant comments to the effect that they preferred to dive in and play with the system rather than being spoon-fed an explanation of how it works. These factors combine to determine what, for each given end-user, constitutes satisfactory and actionable understanding. exp
Explainable Artificial Intelligence (XAI) has re-emerged in response to the development of modern AI and ML systems. These systems are complex and sometimes biased, but they nevertheless make decisions that impact our lives. XAI systems are frequently algorithm-focused; starting and ending with an algorithm that implements a basic untested idea about explainability. These systems are often not tested to determine whether the algorithm helps users accomplish any goals, and so their explainability remains unproven. We propose an alternative: to start with human-focused principles for the design, testing, and implementation of XAI systems, and implement algorithms to serve that purpose. In this paper, we review some of the basic concepts that have been used for user-centered XAI systems over the past 40 years of research. Based on these, we describe the"Self-Explanation Scorecard", which can help developers understand how they can empower users by enabling self-explanation. Finally, we present a set of empirically-grounded, user-centered design principles that may guide developers to create successful explainable systems.
The field of Explainable AI (XAI) has focused primarily on algorithms that can help explain decisions and classification and help understand whether a particular action of an AI system is justified. These \emph{XAI algorithms} provide a variety of means for answering a number of questions human users might have about an AI. However, explanation is also supported by \emph{non-algorithms}: methods, tools, interfaces, and evaluations that might help develop or provide explanations for users, either on their own or in company with algorithmic explanations. In this article, we introduce and describe a small number of non-algorithms we have developed. These include several sets of guidelines for methodological guidance about evaluating systems, including both formative and summative evaluation (such as the self-explanation scorecard and stakeholder playbook) and several concepts for generating explanations that can augment or replace algorithmic XAI (such as the Discovery platform, Collaborative XAI, and the Cognitive Tutorial). We will introduce and review several of these example systems, and discuss how they might be useful in developing or improving algorithmic explanations, or even providing complete and useful non-algorithmic explanations of AI and ML systems.
Alberto J. Cañas合作论文数Institute for Human and Machine Cognition.3
Larry Bunch合作论文数IHMC 2