The integration of machine learning (ML) into practical, real-world applications across different sectors often faces limitations. One barrier to ML integration may be a lack of trust in ML tools, as they can make unpredictable errors and lack transparency. Developers, recognizing this potential lack of trust, have developed confidence scores to accompany ML decisions. However, these confidence scores also suffer from accuracy, transparency, and domain-relevancy issues. Here, we demonstrate a new type of confidence score—the Expert-Derived Confidence (EDC) score—based on human experts who have studied the performance of a classifier and can predict the accuracy of its classification decisions. Confidence scores generated by a human (possibly a peer) may be a more reliable, more transparent, and more relevant guide for relying on ML within specific use cases. This study demonstrates the generation and utility of EDC scores to support reliance on an ML classifier for automated particle analysis.
To address the increasing need for efficient data analysis and decision-making in defense, the U.S. Navy is prioritizing AI/ML systems capable of handling multi-source data and suggesting courses of actions. Historically, many such systems failed due to technical issues and lack of usability or mission relevance. Research on human-AI collaboration aims to create AI systems that better integrate with frontline operators. A recent National Academies of Sciences, Engineering, and Medicine Report (NASEM) outlined 57 research goals, but the U.S. Navy requires a more-focused set of priorities. A workshop with 23 experts from various fields was held at the Naval Information Warfare Center Pacific, resulting in five key research priorities spanning different time frames. This panel discusses the workshop’s findings, highlighting critical questions going into this workshop and those coming out of it. Panel participants come from government, academia, and industry, providing unique perspectives on big questions for human-AI teaming.
With the increasing use and adoption of artificial intelligence (AI), the reliability of modern data systems will be driven by a tighter teaming between human experts and intelligent machine teammates. As in the case of human-human teams, the success of human-machine teams will also rely on clear communication about mutual goals and actions. In this paper, we combine related literature from cognitive psychology, human-machine teaming, uncertainty in data analysis, and multi-agent systems to propose a new form of uncertainty: interaction uncertainty for characterizing bidirectional communication in human-machine teams. We map the causes and effects of interaction uncertainty and outline potential ways to mitigate uncertainty for mutual trust in a high-consequence real-world scenario.
As the field of deep learning has emerged in recent years, the amount of knowledge and expertise that data scientists are expected to absorb and maintain has correspondingly increased. One of the challenges experienced by data scientists working with deep learning models is developing confidence in the accuracy of their approach and the resulting findings. In this study, we conducted semi-structured interviews with data scientists at a national laboratory to understand the processes that data scientists use when attempting to develop their models and the ways that they gain confidence that the results they obtained were accurate. These interviews were analysed to provide an overview of the techniques currently used when working with machine learning (ML) models. Opportunities for collaboration with human factors researchers to develop new tools are identified.
IntroductionThe field of machine learning and its subfield of deep learning have grown rapidly in recent years. With the speed of advancement, it is nearly impossible for data scientists to maintain expert knowledge of cutting-edge techniques. This study applies human factors methods to the field of machine learning to address these difficulties.MethodsUsing semi-structured interviews with data scientists at a National Laboratory, we sought to understand the process used when working with machine learning models, the challenges encountered, and the ways that human factors might contribute to addressing those challenges.ResultsResults of the interviews were analyzed to create a generalization of the process of working with machine learning models. Issues encountered during each process step are described.DiscussionRecommendations and areas for collaboration between data scientists and human factors experts are provided, with the goal of creating better tools, knowledge, and guidance for machine learning scientists.
This paper describes a methodology for developing a new confidence metric to improve power grid operator reliance on ML event classifiers. Unlike traditional confidence scores that are generated by the ML, this confidence metric is generated by humans who have spent time studying the performance boundaries of the ML classifier. We refer to this metric as an Expert Derived Confidence (EDC) score. As an initial test of our methodology four participants (3 Subject Matter Experts, 1 Novice) learned the boundaries of an ML’s performance by studying a subset of events in the ML’s training data. Next, the participants rated their confidence in the ML’s ability to classify similar events. The researchers found that all participants’ EDC scores were correlated with the ML’s own uncertainty quantification score and on average EDC scores showed greater confidence in the ML’s ability to correctly classify events when compared to the ML’s own confidence scores. In addition, averaging EDC scores across all participants was the strongest predictor of model performance and predicted performance even after controlling for the ML’s own confidence.
A novel set of system-state and control-action penalty functions are introduced as an alternative to traditional performance index contingency ranking. The novel system state penalty metrics are formulated based on piecewise linear functions of the system voltage and branch flow, guided by Weber’s Law of human cognition. Novel continuous and discrete control action metrics are also developed to measure the inherent cost and risk associated with every action taken by human power system operator to resolve violations on a pre-contingent basis. These new metrics are combined with traditional human factors indices for measuring human-machine trust and cognitive workload to create a systematic framework for measuring and evaluating operator trust and reliance on artificial intelligence (AI) algorithms for control room use. An existing AI-based contingency analysis recommender tool using a semi-supervised action algorithm is selected for a series of experiments with operations engineering staff using the IEEE 118 Bus System. The penalty metrics presented are demonstrated for both steady-state contingency analysis and transient stability studies, with the operations participants able to reduce the total system penalty in 85% of scenarios through remedial actions. A human-machine team was able to achieve equal or lower continuous control action penalty scores than the participant without availability of the recommender in 57% of experiment scenarios and lower continuous control action penalty scores than the AI tool alone in 83% of scenarios.
Introducing machine learning (ML) assistance into any established process comes with adoption barriers, including entrenched procedures, technological and human readiness levels, human-machine trust, and work culture resistance to change. These barriers are even greater in critical operations such as operating a national or regional power grid, in which both regulatory frameworks and the importance of maintaining reliability levels causes additional resistance to the adoption of new computational support. Developers of future systems and job aides must consider not only technical aspects, but also whether new systems are usable by power system operators. This work presents the methodology and results of a study to evaluate the usability and readiness of a prototype recommender system for power grid contingency analysis. We explore operator cognitive load and evaluate operator performance when solving a collection of scenarios both with and without recommender assistance. We also examine operator trust in the system. We report insights gained on the readiness of the system using a collection of evaluation techniques.
This document presents a sample operations manual for the IEEE 118 Bus Model, which is a synthetic test case developed in 1962 from a section of the transmission grid operated by American Electric Power (AEP). The model is one of the most used synthetic test cases for development of power system applications and is referenced by over 9000 papers. However, the model lacks any context for use with real-time energy management system (EMS) applications, including common considerations, such as generator ramp rates, reactive capabilities, operating limits, and other information typically used by power system operators for real-time decision making. This manual divides the IEEE 118 Bus Model into three operating areas, defines various operating limits, and sets recommended operating procedures for responding to a few types of emergency operating conditions. The manual can be used in support of a wide variety of human-in-the-loop evaluation methodologies for new advanced power applications.
Although a wide array of tools and technologies have been developed over the last decade to support power grid operators, deployment of these tools has been less successful. One reason for unsuccessful deployment may be a focus on error reduction without an adequate understanding of the factors that contribute to operator error in the control room. An analysis of these factors (i.e., vulnerabilities) may provide the baseline understanding needed to inform new technology integration. In an attempt to learn more about these vulnerabilities and their perceived impact on human error we collected and analyzed survey data from 20 electric grid control room operators. We asked survey respondents to consider the various operator, technology and interaction vulnerabilities that may arise during work in the control room and record their attitudes and experiences toward each. Results suggest operator inexperience, high mental workload and fatigue are the most common vulnerabilities experienced during a shift. Technology solutions should set operators up for success by addressing these factors. Survey results were analyzed to explore these vulnerabilities in greater depth.
This work presents the application of a methodology to measure domain expert trust and workload, elicit feedback, and understand the technological usability and impact when a machine learning assistant is introduced into contingency analysis for real-time power grid simulation. The goal of this framework is to rapidly collect and analyze a broad variety of human factors data in order to accelerate the development and evaluation loop for deploying machine learning applications. We describe our methodology and analysis, and we discuss insights gained from a pilot participant about the current usability state of an early technology readiness level (TRL) artificial neural network (ANN) recommender.
When new technology, such as artificial intelligence (AI), is introduced into an existing workflow it may impact risk by mitigating some vulnerabilities and threats in the workflow while introducing others. We present a versatile methodology for assessing the vulnerabilities and threats that impact overall risk in a workflow to inform technology integration. Our method involves a four step assessment of risk including a qualitative expert knowledge elicitation, identification of risk components, quantitative data collection and analysis based on a formula for generating a risk score. The quantification of risk can be used to guide technology integration. We describe our methodology and demonstrate its utility by applying it to the Derivative Classification review process. This work was funded by the Department of Energy (DOE).
Trust calibration for a human–machine team is the process by which a human adjusts their expectations of the automation’s reliability and trustworthiness; adaptive support for trust calibration is needed to engender appropriate reliance on automation. Herein, we leverage an instance-based learning ACT-R cognitive model of decisions to obtain and rely on an automated assistant for visual search in a UAV interface. This cognitive model matches well with the human predictive power statistics measuring reliance decisions; we obtain from the model an internal estimate of automation reliability that mirrors human subjective ratings. The model is able to predict the effect of various potential disruptions, such as environmental changes or particular classes of adversarial intrusions on human trust in automation. Finally, we consider the use of model predictions to improve automation transparency that account for human cognitive biases in order to optimize the bidirectional interaction between human and machine through supporting trust calibration. The implications of our findings for the design of reliable and trustworthy automation are discussed.
Detect the expected, discover the unexpected was the founding principle of the field of visual analytics. This mantra implies that human stakeholders, like a domain expert or data analyst, could leverage visual analytics techniques to seek answers to known unknowns and discover unknown unknowns in the course of the data sense-making process. We argue that in the era of AI-driven automation, we need to recalibrate the roles of humans and machines (e.g., a machine learning model) as teammates. We posit that by realizing human-machine teams as a stakeholder unit, we can better achieve the best of both worlds: automation transparency and human reasoning efficacy. However, this also increases the burden on analysts and domain experts towards performing more cognitively demanding tasks than what they are used to. In this paper, we reflect on the complementary roles in a human-machine team through the lens of cognitive psychology and map them to existing and emerging research in the visual analytics community. We discuss open questions and challenges around the nature of human agency and analyze the shared responsibilities in human-machine teams.
We propose common ground and autonomy are the two critical dimensions necessary for intelligent machine agents to make the transition from tool to teammate. Existing models delineate a number of teammate characteristics. We explore how these teammate characteristics can be distilled into common ground and autonomy and suggest research steps to test our proposal.
Many of the challenges associated with cybersecurity operations are also ripe opportunities for the application of human-machine teaming. Advances in cognitive science, artificial intelligence, and machine learning promise the technology to achieve the level of intelligence and sophistication necessary to create a true intelligent teammate. This panel gathers experts from across the community to discuss the challenges and opportunities for human-machine teaming in cybersecurity operations.
Automation can be unreliable. This makes appropriate trust and reliance difficult to calibrate. One solution to building appropriate trust is to increase automation transparency by displaying information to the operator about the technology’s underlying analytical principles. However, displaying this additional information may increase operator workload. The research and development community must balance the competing demands of providing adequate transparency and keeping operator workload low. To investigate the complex effects of transparency on workload, a modeling approach can be used by computing a measure of processing efficiency called the capacity coefficient. We conducted a study to examine the impact of increasing transparency on operator workload using the capacity coefficient. We present the data from one participant with the goal of demonstrating the utility of the capacity coefficient. We discuss how this participant’s data highlights the inferences possible from capacity analysis for measuring the impact of display design and increased transparency on operator workload.
A variety of factors can affect one’s reliance on an automated aid. Some of these factors include one’s perception of the system’s trustworthiness, such as perceived reliability of the system or one’s ability to understand the system’s underlying reasoning. A mismatch between the operator’s perception and the true capabilities and characteristics of the system can lead to inappropriate reliance on the tool. This improper use of the system can manifest as either underutilization of the technology or complacency resulting from over-trusting the system. Increasing an automated tool’s transparency is one approach that enables the operator to more appropriately rely on the technology. Transparent automated systems provide additional information that allows the user to see the system’s intent and understand its underlying processes and capabilities. Several researchers have developed frameworks to support the design of more transparent automation. However, these frameworks may not fully consider the particular challenges to transparency design introduced by automation that leverages machine learning. Like all automation, these systems can benefit from transparency. However, artificial intelligence poses new challenges that must be considered when designing for transparency. Unique considerations must be made in terms of the type, and amount or level of transparency information conveyed to the user.