This paper describes a generalizable model evaluation method that can be adapted to evaluate AI/ML models across multiple criteria including core scientific principles and more practical outcomes. Emerging from prediction competitions in Psychology and Decision Science, the method evaluates a group of candidate models of varying type and structure across multiple scientific, theoretic, and practical criteria. Ordinal ranking of criteria scores are evaluated using voting rules from the field of computational social choice and allow the comparison of divergent measures and types of models in a holistic evaluation. Additional advantages and applications are discussed.
Many behavioral science studies result in large amounts of unstructured data sets that are costly to code and analyze, requiring multiple reviewers to agree on systematically chosen concepts and themes to categorize responses. Large language models (LLMs) have potential to support this work, demonstrating capabilities for categorizing, summarizing, and otherwise organizing unstructured data. In this paper, we consider that although LLMs have the potential to save time and resources performing coding on qualitative data, the implications for behavioral science research are not yet well understood. Model bias and inaccuracies, reliability, and lack of domain knowledge all necessitate continued human guidance. New methods and interfaces must be developed to enable behavioral science researchers to efficiently and systematically categorize unstructured data together with LLMs. We propose a framework for incorporating human feedback into an annotation workflow, leveraging interactive machine learning to provide oversight while improving a language model's predictions over time.
Objectives Digital map applications produce maps on the fly, using label placement algorithms to optimize the map layout for various factors. Previous work has made great strides in quickly generating layout placement algorithms that minimize collisions or reduce clutter, but little has been done to explore optimizing layouts while considering the visual preferences of map users. This dataset was collected as part of a study to better understand people’s visual preferences for various label positions in map layouts and to support the development of new label placement algorithms that can generate map layouts that consider those preferences. Data description This dataset of label placement preferences includes two parts. The first part contains 956 binary choice responses that asked participants to choose their favorite of two maps. Labels on each map differed in either their alignment to the point of interest or distance from it. The dataset also includes 1912 ranked choice responses where participants were asked to rank 3 maps that differed in both alignment and distance. The alignment and distance positions considered are the same across both parts of the study.
Though an accurate measurement of entropy, or more generally uncertainty, is critical to the success of human–machine teams, the evaluation of the accuracy of such metrics as a probability of machine correctness is often aggregated and not assessed as an iterative control process. The entropy of the decisions made by human–machine teams may not be accurately measured under cold start or at times of data drift unless disagreements between the human and machine are immediately fed back to the classifier iteratively. In this study, we present a stochastic framework by which an uncertainty model may be evaluated iteratively as a probability of machine correctness. We target a novel problem, referred to as the threshold selection problem, which involves a user subjectively selecting the point at which a signal transitions to a low state. This problem is designed to be simple and replicable for human–machine experimentation while exhibiting properties of more complex applications. Finally, we explore the potential of incorporating feedback of machine correctness into a baseline naïve Bayes uncertainty model with a novel reinforcement learning approach. The approach refines a baseline uncertainty model by incorporating machine correctness at every iteration. Experiments are conducted over a large number of realizations to properly evaluate uncertainty at each iteration of the human–machine team. Results show that our novel approach, called closed-loop uncertainty, outperforms the baseline in every case, yielding about 45% improvement on average.
Digital maps are important for many decision-making tasks that require situational awareness, navigation, or location-specific data. Often, digital mapping tools must generate a map that displays labels near associated features in a visually appealing manner, without occluding important information. Automated label placement systems generally accomplish this nontrivial task through a combination of heuristic algorithms and cartography rules, but the resulting maps often do not reflect the preferences and needs of the map user. To achieve higher quality map views, research is needed to identify cognitive and computational approaches for generating high-quality maps that meet user needs and expectations. In this paper, we present a study that explores the visual preferences of map users and supports the development of a preference model for digital map displays. In particular, we found that participants demonstrated consistent preferences for how labels are placed near their point of interest, and that they were more likely to choose positions that prioritized alignment over distance when ranking labels that made trade-offs between them.
Machine learning (ML) algorithms are often assumed to be the most accurate way of producing predictive models despite problems with explainability and adverse impact. The 3rd annual Society for Industrial and Organizational Psychology Machine Learning Competition sought to find ML models for personnel selection that could balance the best of ML prediction with the constraint of minimizing selection bias based on race and gender. To test the possible advantages of simple rules over ML algorithms, we entered a simple and explainable rule-based model inspired by recent advances in model comparison. This simple model outperformed most ML models entered and was comparable to the top performers while retaining positive qualities such as explainability and transparency.
The 3rd annual SIOP Machine Learning (ML) Competition sought to find ML models for personnel selection that could balance the best of ML prediction balanced with the constraint of not increasing adverse impact. To test the possible advantages of simple rules over ML algorithms, we entered a simple and explainable rule based model inspired by recent advances in model comparison. This simple model outperformed most ML models entered and was comparable to the top performers.
In many real world situations, collective decisions are made using voting and, in scenarios such as committee or board elections, employing voting rules that return multiple winners. In multi-winner approval voting (AV), an agent submits a ballot consisting of approvals for as many candidates as they wish, and winners are chosen by tallying up the votes and choosing the top-k candidates receiving the most approvals. In many scenarios, an agent may manipulate the ballot they submit in order to achieve a better outcome by voting in a way that does not reflect their true preferences. In complex and uncertain situations, agents may use heuristics instead of incurring the additional effort required to compute the manipulation which most favors them. In this paper, we examine voting behavior in single-winner and multi-winner approval voting scenarios with varying degrees of uncertainty using behavioral data obtained from Mechanical Turk. We find that people generally manipulate their vote to obtain a better outcome, but often do not identify the optimal manipulation. There are a number of predictive models of agent behavior in the social choice and psychology literature that are based on cognitively plausible heuristic strategies. We show that the existing approaches do not adequately model our real-world data. We propose a novel model that takes into account the size of the winning set and human cognitive constraints; and demonstrate that this model is more effective at capturing real-world behaviors in multi-winner approval voting scenarios.
Geospatial information systems (GIS) support decision making and situational awareness in a wide variety of applications. These systems often require large amounts of labeled data to be displayed in a way that is easy to use and understand. Manually editing these information displays can be extremely time-consuming for an analyst. Algorithms have been designed to alleviate some of this work by automatically generating map displays or digitizing features. However, these systems regularly make mistakes, requiring analysts to verify and correct their output. This human-in-the-loop process of validating the algorithm's labels can provide a means to continuously improve a model over time by using interactive machine learning (IML). This process allows for systems that can function with little or no training data and as the features continue to evolve. Such systems must also account for the strengths and limitations of both the analysts and underlying algorithms to avoid unnecessary frustration, encourage adoption, and increase productivity of the human-machine team. In this chapter, we introduce three examples of how IML has been used in GIS systems for airfield change detection, geographic region digitization and digital map editing. We also describe several considerations for designing IML workflows to ensure that the analyst and system complement one another, resulting in increased productivity and quality of the GIS output. Finally, we will consider new challenges that arise when applying IML to the complex task of automatic map labeling.
Selection and effort are central to attention, yet it is unclear whether they draw on a common pool of cognitive resources, and if so, whether there are differences for early versus later stages of cognitive processing. This study assessed effort by quantifying the vigilance decrement, and spatial processing at early and later stages as a function of time-on-task. Participants performed an auditory spatial attention task, with occasional "catch" trials requiring no response. Psychophysiological measures included bilateral cerebral blood flow (transcranial Doppler), pupil dilation, and blink rate. The shape of attention gradients using reaction time indexed early processing, and did not significantly vary over time. Later stimulus-response conflict was comparable over time, except for a reduction to left hemispace stimuli. Target and catch trial accuracy decreased with time, with a more abrupt decrease for catch versus target trials. Diffusion decision modeling found progressive decreases in information accumulation rate and non-decision time, and the adoption of more liberal response criteria. Cerebral blood flow increased from baseline and then decreased over time, particularly in the left hemisphere. Blink rate steadily increased over time, while pupil dilation increased only at the beginning and then returned towards baseline. The findings suggest dissociations between resources for selectivity and effort. Measures of high subjective effort and temporal declines in catch trial accuracy and cerebral blood flow velocity suggest a standard vigilance decrement was evident in parallel with preserved selection. Different attentional systems and classes of computations that may account for dissociations between selectivity versus effort are discussed.
I present an overview of my research which investigates how models of human behavior can inform the design of new algorithms and interfaces. Specifically, I show how precise, testable computational methods and behavioral experiments can be used to simulate heuristics and bias in human attention and decision making.
In order to increase productivity, capability, and data exploitation, numerous defense applications are experiencing an integration of state-of-the-art machine learning and AI into their architectures. Especially for defense applications, having a human analyst in the loop is of high interest due to quality control, accountability, and complex subject matter expertise not readily automated or replicated by AI. However, many applications are suffering from a very slow transition. This may be in large part due to lack of trust, usability, and productivity, especially when adapting to unforeseen classes and changes in mission context. Interactive machine learning is a newly emerging field in which machine learning implementations are trained, optimized, evaluated, and exploited through an intuitive human-computer interface. In this paper, we introduce interactive machine learning and explain its advantages and limitations within the context of defense applications. Furthermore, we address several of the shortcomings of interactive machine learning by discussing how cognitive feedback may inform features, data, and results in the state of the art. We define the three techniques by which cognitive feedback may be employed: self reporting, implicit cognitive feedback, and modeled cognitive feedback. The advantages and disadvantages of each technique are discussed.
Reports on the experiences of doctoral consortia during the COVID-19 pandemic.
In many real world situations, collective decisions are made using voting. Moreover, scenarios such as committee or board elections require voting rules that return multiple winners. In multi-winner approval voting (AV), an agent may vote for as many candidates as they wish. Winners are chosen by tallying up the votes and choosing the top-$k$ candidates receiving the most votes. An agent may manipulate the vote to achieve a better outcome by voting in a way that does not reflect their true preferences. In complex and uncertain situations, agents may use heuristics to strategize, instead of incurring the additional effort required to compute the manipulation which most favors them. In this paper, we examine voting behavior in multi-winner approval voting scenarios with complete information. We show that people generally manipulate their vote to obtain a better outcome, but often do not identify the optimal manipulation. Instead, voters tend to prioritize the candidates with the highest utilities. Using simulations, we demonstrate the effectiveness of these heuristics in situations where agents only have access to partial information.
In many collective decision making situations, agents vote to choose an alternative that best represents the preferences of the group. Agents may manipulate the vote to achieve a better outcome by voting in a way that does not reflect their true preferences. In real world voting scenarios, people often do not have complete information about other voter preferences and it can be computationally complex to identify a strategy that will maximize their expected utility. In such situations, it is often assumed that voters will vote truthfully rather than expending the effort to strategize. However, being truthful is just one possible heuristic that may be used. In this paper, we examine the effectiveness of heuristics in single winner and multi-winner approval voting scenarios with missing votes. In particular, we look at heuristics where a voter ignores information about other voting profiles and makes their decisions based solely on how much they like each candidate. In a behavioral experiment, we show that people vote truthfully in some situations and prioritize high utility candidates in others. We examine when these behaviors maximize expected utility and show how the structure of the voting environment affects both how well each heuristic performs and how humans employ these heuristics.
Attention plays a fundamental role in higher-level cognition. In this paper we develop a computational model for how auditory spatial attention is distributed in space. Our model builds on the assumption that attentional bias has bottom-up and top-down components. We represent each component and their synthesis as a map, associating a level of attentional bias to locations in space. The maps and their interaction are modeled using an artificial intelligence approach based on constraints. We describe the behavioral task we have designed to measure the attentional bias and discuss the results. We then test different hypotheses on the shape and interaction modalities of the maps in terms of how well they fit our behavioral data. The findings showed that combining top-down and bottom-up spatial attention gradients that differ in their spatial properties produced the best fit to behavioral data, and suggested several novel mechanisms for future testing.
Costly mistakes can occur when decision makers rely on intuition or learned biases to make decisions. To better understand the cognitive processes that lead to bias and develop strategies to combat it, we developed an intelligent agent using the cognitive architecture, ACT-R 7.0. The agent simulates a human participating in a decision making task designed to assess the effectiveness of bias reduction strategies. The agent's performance is compared to that of human participants completing a similar task. Similar results support the underlying cognitive theories and reveal limitations of reducing bias in human decision making. This should provide insights for designing intelligent agents that can reason about bias while supporting decision makers.
Attention has been the focus of a considerable amount of research in cognitive models. Yet, most of the work has been devoted to studying visual attention. In this paper we focus, instead, on auditory attention and on a model for how it is distributed in space following basic ideas of top-down and bottom-up attentional control from verbal models. In particular, we extend a previous computational model [Golob et al., 2016; 2017] which is organized around three main components: a goal map, a saliency map, and a priority map. The goal map models the distribution of attention which is allocated by choice (top-down component). The saliency map, as the name suggests, models attention related to the saliency of auditory stimuli (bottom-up component) and the priority map synthesizes the other two maps in an overall distribution of the attentional bias. This model was shown to be successful in modeling behavioral data of experiments where there is a single attended location. We relax this assumption and extend the framework to encompass scenarios where there can be multiple attended locations. Most importantly, we leverage the parameters learned by fitting the behavioral data with single attended location to make predictions for the case in which sounds are presented at multiple locations with equal probability. Our predictions feature a very small error with respect the new behavioral data and are shown to leave very small room for improvement. This is an important step in the, still largely unexplored, field of auditory attention modeling as it provides a first example of how the computational model can be used as a predictor.
It is well-established that spatial attention can be allocated as a gradient that diminishes from a central focus. In this paper we consider auditory attention and we develop a model for how it is distributed in space following basic ideas of top-down and bottom-up attentional control from verbal models [6, 12]. There are three main components of our model: a goal map, a saliency map, and a priority map. The goal map models the distribution of attention which is allocated by choice (top-down component). The saliency map, as the name suggests, models attention related to the saliency of auditory stimuli (bottom-up component) and the priority map synthesizes the other two maps in an overall distribution of the attentional bias. We model the three maps and their interaction using the well established AI framework of constraint satisfaction problems. We study several hypotheses on the maps and we contrast the results in terms of data obtained running different kinds of experiments. Our computational model, is to the best of our knowledge is the first which targets specifically the auditory system. Our constraint-based approach is very flexible in terms of embedding and testing different hypotheses on the components and constraint propagation techniques allow both to focus on single components as well as to consider the system dynamics as whole. The predictions arising from our model well fit the experimental data, are cognitive plausible and provide new interesting insights to the mechanism of attention control.