As the field of deep learning has emerged in recent years, the amount of knowledge and expertise that data scientists are expected to absorb and maintain has correspondingly increased. One of the challenges experienced by data scientists working with deep learning models is developing confidence in the accuracy of their approach and the resulting findings. In this study, we conducted semi-structured interviews with data scientists at a national laboratory to understand the processes that data scientists use when attempting to develop their models and the ways that they gain confidence that the results they obtained were accurate. These interviews were analysed to provide an overview of the techniques currently used when working with machine learning (ML) models. Opportunities for collaboration with human factors researchers to develop new tools are identified.
IntroductionThe field of machine learning and its subfield of deep learning have grown rapidly in recent years. With the speed of advancement, it is nearly impossible for data scientists to maintain expert knowledge of cutting-edge techniques. This study applies human factors methods to the field of machine learning to address these difficulties.MethodsUsing semi-structured interviews with data scientists at a National Laboratory, we sought to understand the process used when working with machine learning models, the challenges encountered, and the ways that human factors might contribute to addressing those challenges.ResultsResults of the interviews were analyzed to create a generalization of the process of working with machine learning models. Issues encountered during each process step are described.DiscussionRecommendations and areas for collaboration between data scientists and human factors experts are provided, with the goal of creating better tools, knowledge, and guidance for machine learning scientists.
The author aims to encourage mathematicians to learn about and conduct research in the area of cognitive science by walking the reader through the process of utilizing tools and formalism from mathematics to address challenges in perception.
The mapper graph is a popular tool from topological data analysis that provides a graphical summary of point cloud data. It has been used to study data from cancer research, sports analytics, neurosciences, and machine learning. In particular, mapper graphs have been used recently to visualize the topology of high-dimensional artificial neural activations from convolutional neural networks and large language models. However, a key question that arises from using mapper graphs across applications is how to compare mapper graphs to study their structural differences. In this paper, we introduce a distance between mapper graphs using tools from optimal transport. We demonstrate the utility of such a distance by studying the topological changes of neural activations across convolutional layers in deep learning, as well as by capturing the loss of structural information for a multiscale mapper.
A novel set of system-state and control-action penalty functions are introduced as an alternative to traditional performance index contingency ranking. The novel system state penalty metrics are formulated based on piecewise linear functions of the system voltage and branch flow, guided by Weber’s Law of human cognition. Novel continuous and discrete control action metrics are also developed to measure the inherent cost and risk associated with every action taken by human power system operator to resolve violations on a pre-contingent basis. These new metrics are combined with traditional human factors indices for measuring human-machine trust and cognitive workload to create a systematic framework for measuring and evaluating operator trust and reliance on artificial intelligence (AI) algorithms for control room use. An existing AI-based contingency analysis recommender tool using a semi-supervised action algorithm is selected for a series of experiments with operations engineering staff using the IEEE 118 Bus System. The penalty metrics presented are demonstrated for both steady-state contingency analysis and transient stability studies, with the operations participants able to reduce the total system penalty in 85% of scenarios through remedial actions. A human-machine team was able to achieve equal or lower continuous control action penalty scores than the participant without availability of the recommender in 57% of experiment scenarios and lower continuous control action penalty scores than the AI tool alone in 83% of scenarios.
Topological data analysis (TDA) is a branch of computational mathematics, bridging algebraic topology and data science, that provides compact, noise-robust representations of complex structures. Deep neural networks (DNNs) learn millions of parameters associated with a series of transformations defined by the model architecture, resulting in high-dimensional, difficult-to-interpret internal representations of input data. As DNNs become more ubiquitous across multiple sectors of our society, there is increasing recognition that mathematical methods are needed to aid analysts, researchers, and practitioners in understanding and interpreting how these models' internal representations relate to the final classification. In this paper, we apply cutting edge techniques from TDA with the goal of gaining insight into the interpretability of convolutional neural networks used for image classification. We use two common TDA approaches to explore several methods for modeling hidden-layer activations as high-dimensional point clouds, and provide experimental evidence that these point clouds capture valuable structural information about the model's process. First, we demonstrate that a distance metric based on persistent homology can be used to quantify meaningful differences between layers, and we discuss these distances in the broader context of existing representational similarity metrics for neural network interpretability. Second, we show that a mapper graph can provide semantic insight into how these models organize hierarchical class knowledge at each layer. These observations demonstrate that TDA is a useful tool to help deep learning practitioners unlock the hidden structures of their models.
With rapid growth in technology, there has been a corresponding growth in research focused on the ways that human-machine interactions can be improved. As part of that work, researchers have explored how human expertise can inform technology design and evaluation. For example, interaction with subject matter experts (SMEs) or end users can help to design and enhance a machine. The human factors of technology release can be divided into five steps: discovery, planning, development, evaluation, and deployment. This framework is a higher-level abstraction of the Human Readiness Levels for technology use and adoption (See et al. 2018). In this exposition, we discuss how human factors methodologies, principles, and practices can be realized in the first phase, Discovery, of the technology development process.
Introducing machine learning (ML) assistance into any established process comes with adoption barriers, including entrenched procedures, technological and human readiness levels, human-machine trust, and work culture resistance to change. These barriers are even greater in critical operations such as operating a national or regional power grid, in which both regulatory frameworks and the importance of maintaining reliability levels causes additional resistance to the adoption of new computational support. Developers of future systems and job aides must consider not only technical aspects, but also whether new systems are usable by power system operators. This work presents the methodology and results of a study to evaluate the usability and readiness of a prototype recommender system for power grid contingency analysis. We explore operator cognitive load and evaluate operator performance when solving a collection of scenarios both with and without recommender assistance. We also examine operator trust in the system. We report insights gained on the readiness of the system using a collection of evaluation techniques.
This document presents a sample operations manual for the IEEE 118 Bus Model, which is a synthetic test case developed in 1962 from a section of the transmission grid operated by American Electric Power (AEP). The model is one of the most used synthetic test cases for development of power system applications and is referenced by over 9000 papers. However, the model lacks any context for use with real-time energy management system (EMS) applications, including common considerations, such as generator ramp rates, reactive capabilities, operating limits, and other information typically used by power system operators for real-time decision making. This manual divides the IEEE 118 Bus Model into three operating areas, defines various operating limits, and sets recommended operating procedures for responding to a few types of emergency operating conditions. The manual can be used in support of a wide variety of human-in-the-loop evaluation methodologies for new advanced power applications.
Although a wide array of tools and technologies have been developed over the last decade to support power grid operators, deployment of these tools has been less successful. One reason for unsuccessful deployment may be a focus on error reduction without an adequate understanding of the factors that contribute to operator error in the control room. An analysis of these factors (i.e., vulnerabilities) may provide the baseline understanding needed to inform new technology integration. In an attempt to learn more about these vulnerabilities and their perceived impact on human error we collected and analyzed survey data from 20 electric grid control room operators. We asked survey respondents to consider the various operator, technology and interaction vulnerabilities that may arise during work in the control room and record their attitudes and experiences toward each. Results suggest operator inexperience, high mental workload and fatigue are the most common vulnerabilities experienced during a shift. Technology solutions should set operators up for success by addressing these factors. Survey results were analyzed to explore these vulnerabilities in greater depth.
This work presents the application of a methodology to measure domain expert trust and workload, elicit feedback, and understand the technological usability and impact when a machine learning assistant is introduced into contingency analysis for real-time power grid simulation. The goal of this framework is to rapidly collect and analyze a broad variety of human factors data in order to accelerate the development and evaluation loop for deploying machine learning applications. We describe our methodology and analysis, and we discuss insights gained from a pilot participant about the current usability state of an early technology readiness level (TRL) artificial neural network (ANN) recommender.
As data grows in size and complexity, finding frameworks which aid in interpretation and analysis has become critical. This is particularly true when data comes from complex systems where extensive structure is available, but must be drawn from peripheral sources. In this paper we argue that in such situations, sheaves can provide a natural framework to analyze how well a statistical model fits at the local level (that is, on subsets of related datapoints) vs the global level (on all the data). The sheaf-based approach that we propose is suitably general enough to be useful in a range of applications, from analyzing sensor networks to understanding the feature space of a deep learning model.
This paper presents an interface and analysis technique for quickly conducting expert elicitation with the goal of determining entity importance. Our interface deploys a two-alternative choice experiment that is capable of representing knowledge graphs in an easy to interpret fashion for users with limited experience with knowledge graphs. Our analysis methodology takes advantage of conjoint analysis techniques and provides entity weights for many SMEs simultaneously. The results largely align with individual participant fits.
Trust calibration for a human–machine team is the process by which a human adjusts their expectations of the automation’s reliability and trustworthiness; adaptive support for trust calibration is needed to engender appropriate reliance on automation. Herein, we leverage an instance-based learning ACT-R cognitive model of decisions to obtain and rely on an automated assistant for visual search in a UAV interface. This cognitive model matches well with the human predictive power statistics measuring reliance decisions; we obtain from the model an internal estimate of automation reliability that mirrors human subjective ratings. The model is able to predict the effect of various potential disruptions, such as environmental changes or particular classes of adversarial intrusions on human trust in automation. Finally, we consider the use of model predictions to improve automation transparency that account for human cognitive biases in order to optimize the bidirectional interaction between human and machine through supporting trust calibration. The implications of our findings for the design of reliable and trustworthy automation are discussed.
Background Representing biological networks as graphs is a powerful approach to reveal underlying patterns, signatures, and critical components from high-throughput biomolecular data. However, graphs do not natively capture the multi-way relationships present among genes and proteins in biological systems. Hypergraphs are generalizations of graphs that naturally model multi-way relationships and have shown promise in modeling systems such as protein complexes and metabolic reactions. In this paper we seek to understand how hypergraphs can more faithfully identify, and potentially predict, important genes based on complex relationships inferred from genomic expression data sets. Results We compiled a novel data set of transcriptional host response to pathogenic viral infections and formulated relationships between genes as a hypergraph where hyperedges represent significantly perturbed genes, and vertices represent individual biological samples with specific experimental conditions. We find that hypergraph betweenness centrality is a superior method for identification of genes important to viral response when compared with graph centrality. Conclusions Our results demonstrate the utility of using hypergraphs to represent complex biological systems and highlight central important responses in common to a variety of highly pathogenic viruses.
As data structures and mathematical objects used for complex systems modeling, hypergraphs sit nicely poised between on the one hand the world of network models, and on the other that of higher-order mathematical abstractions from algebra, lattice theory, and topology. They are able to represent complex systems interactions more faithfully than graphs and networks, while also being some of the simplest classes of systems representing topological structures as collections of multidimensional objects connected in a particular pattern. In this paper we discuss the role of (undirected) hypergraphs in the science of complex networks, and provide a mathematical overview of the core concepts needed for hypernetwork modeling, including duality and the relationship to bicolored graphs, quantitative adjacency and incidence, the nature of walks in hypergraphs, and available topological relationships and properties. We close with a brief discussion of two example applications: biomedical databases for disease analysis, and domain-name system (DNS) analysis of cyber data.
Ground Truth program was designed to evaluate social science modeling approaches using simulation test beds with ground truth intentionally and systematically embedded to understand and model complex Human Domain systems and their dynamics Lazer et al. (Science 369:1060–1062, 2020). Our multidisciplinary team of data scientists, statisticians, experts in Artificial Intelligence (AI) and visual analytics had a unique role on the program to investigate accuracy, reproducibility, generalizability, and robustness of the state-of-the-art (SOTA) causal structure learning approaches applied to fully observed and sampled simulated data across virtual worlds. In addition, we analyzed the feasibility of using machine learning models to predict future social behavior with and without causal knowledge explicitly embedded. In this paper, we first present our causal modeling approach to discover the causal structure of four virtual worlds produced by the simulation teams—Urban Life, Financial Governance, Disaster and Geopolitical Conflict. Our approach adapts the state-of-the-art causal discovery (including ensemble models), machine learning, data analytics, and visualization techniques to allow a human-machine team to reverse-engineer the true causal relations from sampled and fully observed data. We next present our reproducibility analysis of two research methods team’s performance using a range of causal discovery models applied to both sampled and fully observed data, and analyze their effectiveness and limitations. We further investigate the generalizability and robustness to sampling of the SOTA causal discovery approaches on additional simulated datasets with known ground truth. Our results reveal the limitations of existing causal modeling approaches when applied to large-scale, noisy, high-dimensional data with unobserved variables and unknown relationships between them. We show that the SOTA causal models explored in our experiments are not designed to take advantage from vasts amounts of data and have difficulty recovering ground truth when latent confounders are present; they do not generalize well across simulation scenarios and are not robust to sampling; they are vulnerable to data and modeling assumptions, and therefore, the results are hard to reproduce. Finally, when we outline lessons learned and provide recommendations to improve models for causal discovery and prediction of human social behavior from observational data, we highlight the importance of learning data to knowledge representations or transformations to improve causal discovery and describe the benefit of causal feature selection for predictive and prescriptive modeling.
We explore rigorous, systematic, and controlled experimental evaluation of adversarial examples in the real world and propose a testing regimen for evaluation of real-world adversarial objects. We show that for small scene/ environmental perturbations, large adversarial performance differences exist. Current state of adversarial reporting exists largely as a frequency count over a dynamic collections of scenes. Our work underscores the need for either a more complete report or a score that incorporates scene changes and baseline performance for models and environments tested by adversarial developers. We put forth a score that attempts to address the above issues in a straightforward exemplar application for multiple generated adversary examples. We contribute the following: 1. a testbed for adversarial assessment, 2. a score for adversarial examples, and 3. a collection of additional evaluations on testbed data.