Efficient restoration of electrical grids from a complete shutdown (i.e. ‘blackstart’) is nontrivial and critical to the safety and resiliency of communities. The problem is well studied from an optimization perspective; grids are analyzed and a set of blackstart plans are created so they are ready for an operator to follow them in an effort to restore the system. However, most existing approaches don’t explicitly account for time, and assume a known network state and full observability and control by grid operators which is not always the case. Here, we introduce a technique for the integrated modeling of optimized restoration plans with a cognitive simulation of operator actions. To this end, we built the CogTasks simulation library that allows for the cognitively sound simulation of hierarchical tasking and problem solving to serve as the basis of the grid operator model. We utilize a power flow-informed restoration framework to generate optimal restoration plans that are then enacted by our cognitive agent on a a dynamic representation of the grid. We show that utilizing this cognitive agent introduces extra noise and stability considerations into the implementation of the restoration plan; while the optimization reward function is still reasonably well correlated with the observed restoration time in our tests, this integrated model reveals value to including operator limitations and actions in restoration planning and lends explainability to the optimized schedules.
Success in extreme environments comes with a cost of subtle performance decrements that if not mitigated properly can lead to lifethreatening consequences. Identification and prediction of performance decline could alleviate deleterious consequences and enhance success in challenging and high-risk operations. The Rim-to-Rim Wearables at the Canyon for Health (R2R WATCH) project was designed to examine the cognitive, physiological, and biological markers of performance decline in the extreme environment of the Grand Canyon Rim-to-Rim (R2R) hike. The study utilized commercial off-the-shelf cognitive and physiological monitoring techniques, along with subjective self-assessments and hematologic measurements to determine subject performance and changes across the hike. The multiyear effort collected these multiple data streams in parallel on a large sample of participants hiking the R2R, leading to a rich and complex data set. This article describes the methodology and its evolution as devices and measurements were assessed after each data collection event. It also highlights a subset of the patterns of results found across the data streams. Subsequent work will draw on this data set to focus on building more sophisticated, predictive statistical models and dive deeper into specific analyses (such as the physiological and biological profiles of hikers who were left behind by their hiking partners).
The Tularosa study was designed to understand how defensive deception-including both cyber and psychological-affects cyber attackers. Over 130 red teamers participated in a network penetration task over two days in which we controlled both the presence of and explicit mention of deceptive defensive techniques. To our knowledge, this represents the largest study of its kind ever conducted on a professional red team population. The design was conducted with a battery of questionnaires (e.g., experience, personality, etc.) and cognitive tasks (e.g., fluid intelligence, working memory, etc.), allowing for the characterization of a "typical" red teamer, as well as physiological measures (e.g., galvanic skin response, heart rate, etc.) to be correlated with the cyber events. This paper focuses on the design, implementation, data, population characteristics, and begins to examine preliminary results.
Commercial off-the-shelf (COTS) wearable devices are used to quantify physiology during physical activities to monitor levels of fitness and to prevent overexertion. We argue that there are limitations and challenges to measuring physiological data with current state-of-the-art wearable devices, both with the hardware as well as the data itself. These limitations and challenges are exacerbated when wearable devices are used in extreme climate environments We discuss these through empirical findings from our study where hikers are suited with wearable technologies as they cross the Grand Canyon. We discuss the performance of various wearable technologies in the extreme environment of the canyon as well as the concerns with downloaded data. These findings highlight the needs and opportunities for the wearable devices market, specifically how wearable technologies could mature to quantify performance and fatigue through real-time data collection and analysis.
Exposure to extreme environments is both mentally and physically taxing, leading to suboptimal performance and even life-threatening emergencies. Physiological and cognitive monitoring could provide the earliest indicator of performance decline and inform appropriate therapeutic intervention, yet little research has explored the relationship between these markers in strenuous settings. The Rim-to-Rim Wearables at the Canyon for Health (R2RWATCH) study is a research project at Sandia National Laboratories funded by the Defense Threat Reduction Agency to identify which physiological and cognitive phenomena collected by non-invasive wearable devices are the most related to performance in extreme environments. In a pilot study, data were collected from civilians and military warfighters hiking the Rim-to-Rim trail at the Grand Canyon. Each participant wore a set of devices collecting physiological, cognitive, and environmental data such as heart rate, memory, ambient temperature, etc. Promising preliminary results found correlates between physiological markers recorded by the wearable devices and decline in cognitive abilities, although further work is required to refine those measurements. Planned follow-up studies will validate these findings and further explore outstanding questions.
In many settings, multi-tasking and interruption are commonplace. Multi-tasking has been a popular subject of recent research, but a multitasking paradigm normally allows the subject some control over the timing of the task switch. In this paper we focus on interruptions-situations in which the subject has no control over the timing of task switches. We consider three types of task: verbal (reading comprehension), visual search, and monitoring/situation awareness. Using interruptions from 30 s to 2 min in duration, we found a significant effect in each case, but with different effect sizes. For the situation awareness task, we experimented with interruptions of varying duration and found a non-linear relation between the duration of the interruption and its after-effect on performance, which may correspond to a task-dependent interruption threshold, which is lower for more dynamic tasks.
The Rim-to-Rim Wearables At The Canyon for Health (R2R WATCH) study examines metrics recordable on commercial off the shelf (COTS) devices that are most relevant and reliable for the earliest possible indication of a health or performance decline. This is accomplished through collaboration between Sandia National Laboratories (SNL) and The University of New Mexico (UNM) where the two organizations team up to collect physiological, cognitive, and biological markers from volunteer hikers who attempt the Rim-to-Rim (R2R) hike at the Grand Canyon. Three forms of data are collected as hikers travel from rim to rim: physiological data through wearable devices, cognitive data through a cognitive task taken every 3 hours, and blood samples obtained before and after completing the hike. Data is collected from both civilian and warfighter hikers. Once the data is obtained, it is analyzed to understand the effectiveness of each COTS device and the validity of the data collected. We also aim to identify which physiological and cognitive phenomena collected by wearable devices are the most relatable to overall health and task performance in extreme environments, and of these ascertain which markers provide the earliest yet reliable indication of health decline. Finally, we analyze the data for significant differences between civilians' and warfighters' markers and the relationship to performance. This is a study funded by the Defense Threat Reduction Agency (DTRA, Project CB10359) and the University of New Mexico (The main portion of the R2R WATCH study is funded by DTRA. UNM is currently funding all activities related to bloodwork. DTRA, Project CB10359; SAND2017-1872 C). This paper describes the experimental design and methodology for the first year of the R2R WATCH project.
The Fluid Events Model is aimed at predicting changes in the actions people take on a moment-by-moment basis. In contrast with other research on action selection, this work does not investigate why some course of action was selected, but rather the likelihood of discontinuing the current course of action and selecting another in the near future. This is done using both task-based and experience-based factors. Prior work evaluated this model in the context of trial-by-trial, independent, interactive events, such as choosing how to copy a figure of a line drawing. In this paper, we extend this model to more covert event experiences, such as reading narratives, as well as to continuous interactive events, such as playing a video game. To this end, the model was applied to existing data sets of reading time and event segmentation for written and picture stories. It was also applied to existing data sets of performance in a strategy board game, an aerial combat game, and a first person shooter game in which a participant’s current state was dependent on prior events. The results revealed that the model predicted behavior changes well, taking into account both the theoretically defined structure of the described events, as well as a person’s prior experience. Thus, theories of event cognition can benefit from efforts that take into account not only how events in the world are structured, but also how people experience those events.
Cyber security is a pervasive issue that impacts public and private organizations. While several published accounts describe the task demands of cyber security analysts, it is only recently that research has begun to investigate the cognitive and performance factors that distinguish novice from expert cyber security analysts. Research in this area is motivated by the need to understand how to better structure the education and training of cyber security professionals, a desire to identify selection factors that are predictive of professional success in cyber security and questions related to the development of software tools to augment human performance of cyber security tasks. However, a common hurdle faced by researchers involves gaining access to cyber security professionals for data collection activities, whether controlled experiments or semi-naturalistic observations. An often readily available and potentially valuable source of data may be found in the records generated through cyber security training exercises. These events frequently entail semi-realistic challenges that may be modeled on real-world occurrences, and occur outside normal operational settings, freeing participants from the sensitivities regarding information disclosure within operational environments. This paper describes an infrastructure tailored for the collection of human performance data within the context of cyber security training exercises. Techniques are described for mining the resulting data logs for relevant human performance variables. The results provide insights that go beyond current descriptive accounts of the cognitive processes and demands associated with cyber security job performance, providing quantitative characterizations of the activities undertaken in solving problems within this domain.
Human performance has become a pertinent issue within cyber security. However, this research has been stymied by the limited availability of expert cyber security professionals. This is partly attributable to the ongoing workload faced by cyber security professionals, which is compounded by the limited number of qualified personnel and turnover of personnel across organizations. Additionally, it is difficult to conduct research, and particularly, openly published research, due to the sensitivity inherent to cyber operations at most organizations. As an alternative, the current research has focused on data collection during cyber security training exercises. These events draw individuals with a range of knowledge and experience extending from seasoned professionals to recent college graduates to college students. The current paper describes research involving data collection at two separate cyber security exercises. This data collection involved multiple measures which included behavioral performance based on human-machine transactions and a questionnaire-based assessments of cyber security experience. It was found that participants reporting more experience with cyber security topics and cyber security software tools made greater use of general purpose software tools, combining the use of general purpose tools with specialized cyber security software applications. Given that organizations make substantial investments in cyber security software tools, it is important to recognize that while these tools enable specialized analyses that would not be possible otherwise, they are not sufficient. Instead, effective cyber security operations involve a range of activities that extends from the highly general (e.g., taking notes, Internet search) to domain specific (e.g., disk forensics) and the accompanying work environment should accommodates this range of activities.
In the context of military training simulation, “semi-automated forces” are software agents that serve as role players. The term implies a degree of shared control – increased automation allows one operator to control a larger number of agents, but too much automation removes control from the instructor. The desired amount of control depends on the situation, so there is no single “best” level of automation. This paper describes the rationale and design for Trainable Automated Forces (TAF), which is based on training by example in order to reduce the development time for automated agents. A central issue is how TAF interprets demonstrated behaviors either as an example to follow specifically, or as contingencies to be executed as the situation permits. We describe the behavior recognizers that allow TAF to produce a high-level model of behaviors. We assess the accuracy of a recognizer for a simple airplane maneuver, showing that it can accurately recognize the maneuver from just a few examples.
The Fluid Events Model is aimed at predicting changes in the actions people take on a moment-by-moment basis. In contrast with other research on action selection, this work does not investigate why some course of action was selected, but rather the likelihood of discontinuing the current course of action and selecting another in the near future. This is done using both task-based and experience-based factors. Prior work evaluated this model in the context of trial-by-trial, independent, interactive events, such as choosing how to copy a figure of a line drawing. In this paper, we extend this model to more covert event experiences, such as reading narratives, as well as to continuous interactive events, such as playing a video game. To this end, the model was applied to existing data sets of reading time and event segmentation for written and picture stories. It was also applied to existing data sets of performance in a strategy board game, an aerial combat game, and a first person shooter game in which a participant's current state was dependent on prior events. The results revealed that the model predicted behavior changes well, taking into account both the theoretically defined structure of the described events, as well as a person's prior experience. Thus, theories of event cognition can benefit from efforts that take into account not only how events in the world are structured, but also how people experience those events.
Within large organizations, the defense of cyber assets generally involves the use of various mechanisms, such as intrusion detection systems, to alert cyber security personnel to suspicious network activity. Resulting alerts are reviewed by the organization’s cyber security personnel to investigate and assess the threat and initiate appropriate actions to defend the organization’s network assets. While automated software routines are essential to cope with the massive volumes of data transmitted across data networks, the ultimate success of an organization’s efforts to resist adversarial attacks upon their cyber assets relies on the effectiveness of individuals and teams. This paper reports research to understand the factors that impact the effectiveness of Cyber Security Incidence Response Teams (CSIRTs). Specifically, a simulation is described that captures the workflow within a CSIRT. The simulation is then demonstrated in a study comparing the differential response time to threats that vary with respect to key characteristics (attack trajectory, targeted asset and perpetrator). It is shown that the results of the simulation correlate with data from the actual incident response times of a professional CSIRT.
Virtual agents are used by the military for extensive training of pilots. However, creating virtual agents with the appropriate behaviors is a lengthy and specialized process. We describe a new tool (TAF) to allow rapid creation of virtual adversarial agents by non-specialized personnel. In TAF, users can draw diagrams of agent behavior which are matched within a pre-existing knowledge base. By combining several instances of similar behaviors a behavior model for a new agent is created that matches the diagram provided by the user.
Objective: The aim of this study was to identify the cognitive factors that predictability and adaptability during multitasking with a flight simulator. Background: Multitasking has become increasingly prevalent as most professions require individuals to perform multiple tasks simultaneously. Considerable research has been undertaken to identify the characteristics of people (i.e., individual differences) that predict multitasking ability. Although working memory is a reliable predictor of general multitasking ability (i.e., performance in normal conditions), there is the question of whether different cognitive faculties are needed to rapidly respond to changing task demands ( adaptability). Method: Participants first completed a battery of cognitive individual differences tests followed by multitasking sessions with a flight simulator. After a baseline condition, difficulty of the flight simulator was incrementally increased via four experimental manipulations, and performance metrics were collected to assess multitasking ability and adaptability. Results: Scholastic aptitude and working memory predicted general multitasking ability (i.e., performance at baseline difficulty), but spatial manipulation (in conjunction with working memory) was a major predictor of adaptability (performance in difficult conditions after accounting for baseline performance). Conclusion: Multitasking ability and adaptability may be overlapping but separate constructs that draw on overlapping (but not identical) sets of cognitive abilities. Application: The results of this study are applicable to practitioners and researchers in human factors to assess multitasking performance in real-world contexts and with realistic task constraints. We also present a framework for conceptualizing multitasking adaptability on the basis of five adaptability profiles derived from performance on tasks with consistent versus increased difficulty.
Assessing a person’s ability to multitask is a topic that is gaining increased attention. However, task constraints and difficulty rarely remain constant in real-world environments; when task constraints change, people must adapt to avoid diminished task performance or failure. But can we identify and predict differences in multitasking adaptability? This question was assessed in an experiment wherein participants multitasked in a flight simulator. Task difficulty was incrementally increased across three experimental manipulations. We measured participants' performance on tasks with baseline versus increased difficulty. Cluster analyses on performance identified three distinct adaptability groups in each condition, irrespective of performance at baseline. Furthermore, individual membership in each cluster was quite consistent across different difficulty conditions. Cluster membership in this task was predicted by spatial ability, which is a cognitive ability not related to general multitasking ability.
Modeling agent behaviors in complex task environments requires the agent to be sensitive to complex stimuli such as the positions and actions of varying numbers of other entities. Entity state updates may be received asynchronously rather than on a coordinated clock signal, so the world state must be estimated based on the most recent information available for each entity. The simulation environment is likely to be distributed across several computers over a network. This paper presents the Relational Blackboard (RBB), which is a framework developed to address these needs with clarity and efficiency. The purpose of this paper is to explain the concepts used to represent and process spatio-temporal data in the RBB framework so researchers in related areas can apply the concepts and software to their own problems of interest; detailed description of our own research will be found in other papers. The software is freely available under the BSD open-source license at http://rbb.sandia.gov.