
As artificial intelligence (AI), machine learning (ML), and other forms of advanced automation are increasingly considered for deployment in safety-critical industries, there is an urgent need for evaluation methods which reliably identify risks of deployment prior to people being harmed. In this narrative review, we discuss the benefits and drawbacks of 11 major methodological decisions underpinning evaluations of AI-infused technologies from the perspective of cognitive systems engineering (CSE) and naturalistic decision making (NDM). These methodological decisions are organized around four aspirations central to the perspective of CSE and NDM: evaluations of AI-infused technologies should be (1) integrated, (2) naturalistic, (3) grounded, and (4) pattern-centered. We use these aspirations to interpret common human-AI evaluation methods and discuss new evaluation challenges for emerging AI-infused technologies. This narrative review is meant to guide both current methods and future research toward safe and effective strategies for evaluating AI-infused technologies, especially in safety-critical settings.
Humans are often the main source of resilience in work systems, responding to surprise events and variety in demands with resourcefulness, initiative, and flexibility that technologies generally lack. Meanwhile, artificial intelligence and automation technologies increasingly are developed to perform macrocognitive work such as planning and decision making in healthcare, transportation, national security, and other high-consequence domains. When developed without consideration for ways this work is performed to assure work system resilience, new technology risks compromising that resilience. We propose a conceptual framework of work system resilience, Transform with Resilience during Upgrades to Socio-Technical Systems (TRUSTS), that translates resilience concepts into work system requirements for technology development teams to consider. We developed TRUSTS in three phases: baseline framework development, specialization for technology development, and iterative refinement. We position the proposed framework relative to four influential bodies of research, describe early applications in technology development, and identify challenges and research needs.
This study examines the information pilots require to maintain situation awareness (SA) in surface trajectory-based operations (STBO), a future airport navigation concept that adds complexity by imposing speed constraints. Systems supporting pilots in STBO have been proposed, but formal analysis of pilot SA requirements have not been documented to guide their design. Such analysis may be conducted with a goal-directed task analysis (GDTA), which identifies SA requirements. However, GDTA relies on operator experience, limiting its use for future operations such as STBO. To answer this challenge, we introduce a future-oriented GDTA (foGDTA), which proposes additions to the interview step of the GDTA. We apply the foGDTA with five professional pilots. Fourteen STBO-specific SA requirements are identified, six of which are not supported by existing displays. From these findings, we propose five design guidelines for STBO support displays to enhance SA and performance. We also highlight broader implications for STBO operations, including the need to factor pilot workload into trajectory computation and to provide more flexibility to flight crews.
This study investigated whether humans can be efficiently trained in human-agent teams (HATs) teamwork competencies to improve HAT collaboration. In HATs, humans and artificially intelligent (AI) agents collaborate on shared tasks, which requires teamwork. However, human and agent approaches to teamwork differ, posing challenges in HATs. These challenges raise the need to train humans to develop teamwork competencies that they can effectively apply in HAT settings. The cooperative video game, KeyWe, was used as a testbed, in which human participants completed tasks with a scripted agent. A HAT training intervention that took less than 30 minutes was developed to train humans on seven teamwork competencies. The training was not associated with the KeyWe game task itself. Half of the participants received the training, and half did not. Participants who received the training delegated a higher percentage of tasks to the agent and more often assigned tasks to the agent by defining strategies than participants who did not receive teamwork training. Trained teams demonstrated resilience by achieving higher task performance when the game difficulty increased. This study demonstrated that training humans to develop teamwork competencies, independent from task training, can enhance collaboration and performance in HATs.
AI systems are rapidly being placed in positions where they can directly influence decisions that require a strong consideration of ethics. This article reports on an experiment in which humans worked with a teammate acting as an expert advisor with expert-level knowledge, who informed and influenced their ethical decision-making. The identity of this expert advisor was manipulated to be either human or AI, and the influence exerted by the expert advisor was manipulated to be either low or high. Further, participants completed four scenarios, each exploring a different ethically charged decision. Results indicated that working with an AI teammate expert advisor led participants to perceive significantly lower levels of stress and responsibility when making ethical decisions, but AI expert advisors were also perceived as performing worse than human expert advisors. Additionally, when expert advisors exerted high influence levels, participants felt significantly less stress during the decision-making process. Finally, the scenario and ethical decisions made by participants had pervasive effects on trust, trustworthiness, perceived performance, and perceived power for both human and AI expert advisors. Future research efforts must ensure that the use of AI expert advisors in human-AI teams does not reduce the responsibility humans bear when making various ethical decisions.
Defence often uses rationalistic approaches for the adoption of AI, but we question whether this is appropriate. By studying adoption using naturalistic approaches in the work environment where AI is actually employed, we better can capture situated issues around trust and adoption. We conducted a longitudinal study of expert intelligence analysts adopting a machine-vision Object Detection and Recognition Overlay (ODRO) within their mission teams and tooling, premised on improving analyst functions. We employ and adapt naturalistic decision-making models to identify key situational challenges, showing that naturalistic approaches can better inform Defence's AI adoption. Specifically, we find that organisational rewards for accuracy influence analysts' willingness to adopt and rely on imperfect tools, and that analysts develop a non-uniform relationship with ODRO over time, which we translate into design implications. Crucially, our results support a model of interdependence over prevalent task-allocation paradigms that assume human and AI as functionally divided, rational entities.
Decision-making during the entrepreneurial process requires adapting to dynamic situational shifts while navigating uncertain and ill-structured environments that impact the economic success of new ventures This study applies a cognitive task analysis approach with nine active, serial entrepreneurs, that have launched 49 startups between them, to examine the critical decisions involved in noticing and evaluating entrepreneurial opportunities in their daily practice when there is too much or limited information. We identify seven cognitive task situations embedded in the value-creating schema utilized by the entrepreneurs: (1) Observing the status quo, (2) Assessing personal interest level, (3) Exploring the problem space, (4) Assessing the ROI of time/effort required, (5) Estimating the value for customers, (6) Identifying others to de-risk the opportunity, and (7) Taking action to build momentum. Each of these cognitive tasks has a set of heuristics for rapid decision-making used as part of the procedural knowledge of the value-creating schema to determine when the composition of cues merits further opportunity development. This set of relation structures provides insight into the abstract schematic mechanisms that facilitate the use of adaptive expertise to reactively and proactively interact with an ill-structured and dynamic environment to accomplish the task.
Situation awareness (SA) is crucial for jobs across the healthcare industry, especially in fast-paced emergency environments. Currently, measures of SA are subjective or require the person to stop and answer questions about the situation. A method to continuously and quantitatively detect levels of SA could potentially serve to assess training effectiveness or to intervene during care when SA is inadequate before mistakes occur. The objective of this study was to evaluate whether electroencephalography (EEG) and/or functional near-infrared spectroscopy (fNIRS) measured at a single location provide additional explanatory value for SA beyond behavioral and experiential characteristics in a simulated emergency care task. Participants with varying levels of medical training completed the task; the Situation Awareness Global Awareness Technique (SAGAT) was used as ground truth for SA assessment. SAGAT scores correlated with both task performance and experience level; however, neither EEG nor fNIRS were correlated with SA.
Air traffic complexity demands high situational awareness (SA) among air traffic controllers. SA comprises three levels: perception, comprehension, and projection. This study examined the development of SA in inexperienced participants. We wanted to study whether mere exposure to SA levels through probes specifically designed for each level could generate a scaffold for the development of SA. Participants were assigned to five groups: three SA level questionnaires, a complete questionnaire, and a control group. Following the SPAM methodology, we measured response times, accuracy, and ATC task performance. We found that the SPAM questionnaire produced reactivity effects in the ongoing scenario, and that these effects were modulated by the SA level group. However, these effects did not systematically transfer to a subsequent scenario for all the SA level groups. Specifically, only the projection group continued to use the same strategy developed in the first scenario. In addition, our results support that the conception of SA structure is not entirely hierarchical since participants could answer higher-level probes without first having to answer lower-level probes. Implications for designing SA training interventions are discussed. These findings indicate that real-time SA questionnaires elicit short-term reactivity but fail to produce lasting impacts, limiting their utility for SA training.
Given the same set of training, qualifications, and information, experts in a reliable and trustworthy system should theoretically make similar decisions. However, undesired decision variability ("noise") has been observed in many high-stakes domains, which is both concerning and unsurprising given the complexity of decisions and known influence of several factors. Moreover, not all variability is undesirable, and the study of intuitive decision-making in naturalistic settings demonstrates the skill of the human expert when evaluating complex cases. The current paper describes a "noise audit" where we evaluated individual risk-informed decision making for regulatory action at U.S. commercial nuclear power plants. In a scenario study, individuals exhibited high consistency in inspection type decisions, although they varied substantially when assigning a significance code, resulting in many decisions that were different than what would be expected by the agency. Together, the results indicate the presence of decision "noise" in one of the two nuclear regulatory processes examined, though some of that variability is desirable given plant and case-specific differences.
In the treatment of cancer, doctors in resource-constrained settings are routinely required to make high-stake decisions under uncertain conditions along with rising burden of cancer, substantial patient load on doctors, and poor healthcare infrastructure that further complicate the decision process. However, very few studies examine the real-world decision process given the prevalence of such factors. Using the naturalistic decision-making framework, this study aimed to understand the critical decisions that doctors face in the treatment of cancer and how they take such decisions. Twenty-five doctors working in a government hospital in India were interviewed, and the data was subjected to thematic analysis. Three central concepts were identified: communicating the illness, different shades of paternalism during the treatment process, and the complexity of the decision process. The critical decisions that doctors faced were found to be treatment-related decisions, resource constraint decisions, and difficulties in communicating the illness. The study provides deep insights into how the doctors 'tailor' the decisions according to the unique contextual realities of individual patients. It further elucidates the intersecting nature of these contextual factors and how they influence the decision process. The study makes significant theoretical and practical implications, such as designing support systems and training programs that enhance decision-making in cancer care within resource-limited healthcare settings.
Incorporating advanced autonomous systems, such as intelligent tactical autopilots, onboard manned military fighter aircraft will allow pilots to focus on managing additional air combat tasks while their autonomous counterparts maneuver the aircraft. One key challenge of maximizing the mission's overall success is that the pilots must optimally divide their attention between monitoring the autonomy, to ensure safety, and other necessary tasks. The research conducted here explores methods for calibrating attention and affecting human trust in an autonomy-aided fighter cockpit while minimizing pilot workload. An experiment was performed to measure pilot eye-gaze, head position, workload, task performance, and trust while flying with autonomous agents that were simulated to be in control of the aircraft during live aerial combat maneuvers. During the combat maneuvers, the participants completed non-flying tasks directly related to the combat scenario, with and without a novel attention calibration system in the loop. With the addition of an attention calibration system in the loop, overall pilot task performance and trust increased, and cognitive workload decreased.
Future planetary exploration extravehicular activity (EVA) will require crew to perform field geology to accomplish mission science goals. However, science operations will be constrained by the EVA environment and may conflict with how field geology has evolved on Earth. In this study, semi-structured interviews were conducted with field-geology experts (n = 8) who specialize in domains that align with NASA's science goals to (1) characterize the constraints, goals, and current operations of Earth-based field geology, and (2) characterize the mental models geologists use for operational decision-making. A qualitative analysis was conducted on interview transcripts to synthesize themes across participant responses. Results were compared to the current EVA work domain to identify changes due to the introduction of science objectives and moving EVAs to natural surfaces with gravity, for example, an increased need for replanning and documentation capabilities. Additionally, tasks during exploration EVAs that may be the most affected by the risk of Earth independent human-system operations were identified. Novel physical and cognitive decision-making factors were identified that influence field geology operational planning. The findings from this study can be used to inform creating work aids that leverage the unique strengths of human geologists while still maintaining crew safety in harsh planetary environments.
The advancement of automated driving technology has led to the enhanced performance of various subsystems integrated into automated vehicles, with their design conforming to international standards and regulatory guidelines. However, the corresponding performance evaluations often overlook human perspectives, such as those of drivers or passengers, potentially resulting in a lack of system safety or customer satisfaction. Therefore, these factors must be considered during testing to ensure the secure and satisfactory operation of automated vehicles. To evaluate the performance of systems embedded in automated vehicles from a human perspective, this paper introduces a framework based on the system-theoretic process analysis, which is a systems approach to analyze hazards attributable to human factors. The Driver Availability Recognition System in the SAE Level 3 automated vehicle was selected as a verification system, and a full-scale driving simulator was used to create the experimental environment. Forty volunteers participated in the performance verification experiment, and data obtained from the human-in-the-loop experiment were incorporated into the proposed framework. The results demonstrated that the framework can ensure effective performance evaluation from a human perspective and suggest safety requirements. The development and application of the proposed framework are anticipated to facilitate the successful rollout of automated vehicles.
Recent successes in artificial intelligence (AI) have ignited debate over its role in human-machine systems—specifically, whether an AI system should be viewed as a tool or a teammate. This article consists of a set of essays that explore this question from a macrocognitive viewpoint. These essays reveal similarities across stances, and divergences regarding interpretations of the teammate metaphor. Discussions include concerns about potential risks and effects of the teammate metaphor on users, in addition to expressions of the value of the metaphor for work system design. The essays highlight the role of metaphors in both enhancing and stifling design creativity, the blending of metaphors in design, and the need for empirical evaluation to guide clear design choices. Moreover, regardless of which approach is adopted, the “technology-first” mindset of system developers, the emphasis on relatively simple human-machine interactions over complex, evolving work systems, and the insidious effects of AI on human expertise all challenge progress in the design of macrocognitive work systems. The essays present some specific research directions.
Despite their intentioned focus on process, macrocognitive models run the risk of advancing a perspective on human behavior that can be called “composite/constraints”—that is, the various, long-prevailing views that people are nothing but composites of factors and variables who behave according to their particular makeup and/or their circumstances. The perspective neglects a key feature of the human experience, primarily because it is mostly invisible to both the researcher and their participants. This paper introduces this invisible feature and explores its integration with a core macrocognitive model, Klein’s Recognition-Primed Decision model. It also suggests implications of this analysis for the theoretical and methodological concerns of the NDM paradigm, applications of NDM models, and for an understanding of artificial general intelligence.
The healthcare community has, in recent years, paid increased attention to errors in the diagnostic process, especially in emergency departments (EDs). However, the expertise of frontline professionals could be an important factor in mitigating risks, preventing diagnostic errors, and improving performance. The current research focused on identifying expertise in ED personnel, complementing efforts to examine errors in reasoning or diagnosis. Data were collected from a three-year, multi-site examination of various EDs. Through the analysis of 43 Critical Decision Method interviews (Klein, Calderwood, & MacGregor, 1989), we identified six aspects of ED expertise: technical, perceptual, organizational, teamwork, emotional, and sensemaking. We also identified barriers to the development of expertise in the ED, such as limited follow-up and feedback on cases. Our findings were consistent with other naturalistic examinations of expertise. Finally, we discuss implications for ED practice, diagnostic safety, and training.
This study was designed to validate the factor structure of the Startle and Surprise Inventories using multilevel confirmatory factor analysis in an ecologically valid flightdeck setting. The Startle and Surprise Inventories were developed to assess self-report startle and surprise to target stimuli. As their use expands in operational settings, construct validity should be further examined in contexts with ecological validity. 208 observations were collected from 26 professional pilots exposed to eight scenarios with varied levels of startle and surprise in a motion-based simulator. After each scenario, pilots completed the Startle and Surprise Inventories. A two-factor model, comprising the constructs Startle and Surprise, demonstrated superior and acceptable fit over a one-factor model. All items demonstrated significant factor loadings at both within- and between-scenario levels in the two-factor solution. McDonald’s ω ranged from ω = 0.88 to ω = 0.96 for the Startle Inventory, and ω = 0.77 to ω = 0.96 for the Surprise Inventory, indicating acceptable to excellent internal consistency. The findings offer empirical support for the construct validity and reliability of the Startle and Surprise Inventories in a highly ecologically-valid setting. The validated and reliable measures can inform evidence-based safety training protocols and interventions in aviation and other safety-critical domains.
This study contributes to the domain of explainable AI (XAI) by examining how uncertainty in an AI agent's diagnosis affects human decision-making in high-stakes, time-critical environments. We conducted a randomized controlled experiment in which participants diagnosed simulated spacecraft anomalies with the help of an AI agent named Daphne. Each participant completed two sessions with different levels of explainability: one in which Daphne provided Basic explanations (qualitative likelihood scores) and another in which it offered Advanced explanations (quantitative likelihood scores with detailed justifications). Each session included a balanced set of scenarios featuring high and low diagnostic uncertainty. We evaluated the participants' task performance, trust in the AI agent, reliance on its recommendations, satisfaction, and self-reported confidence. Across all measures, the participants' performance declined significantly under high uncertainty conditions. However, Advanced explanations were particularly effective in improving performance when uncertainty was high, suggesting that explanation depth plays a key role in mitigating the negative effects uncertainty. These results highlight the importance of transparency in AI systems designed for decision-making. Our findings offer practical insights into the design of XAI tools that can better support human operators, especially in safety-critical domains like spaceflight, where managing uncertainty is a key challenge in effective human-AI collaboration.
High-consequence domains such as medical triage for critical casualties often involve decisions for which there is no single right answer, and for which expert decision makers will disagree. These decisions thus represent a particular challenge for trustworthy and ethical AI in these domains, as simply training AI to make "the right decision" is not possible. In the current research, we offer a framework for aligning AI systems to the key, domain-specific attributes of decision-makers in order to produce trust in human experts. In Study 1, we investigate six possible attributes and find that the degree of alignment predicts rated trust ratings and delegation using short vignettes for assessment and delegation. In Study 2, we extend these results to demonstrate that alignment predicts trust in decisions from two different AI systems across two distinct attributes. The results thus offer a dimensionality-reducing approach to AI trustworthiness in high-consequence domains.