Integrating artificial intelligence (AI) into decision-support systems (DSS) for aviation offers real-time decision support but complicates trust calibration between human operators and AI. This study examined how feedback style from such a DSS, the Cognitive Shadow, influences trust during a simulated weather-avoidance task. Forty-four participants completed 150 knowledge-elicitation trials, followed by 20 test trials where the DSS generated predictions. When participant decisions diverged from the DSS suggestion, it issued explicit recommendations; matching human-DSS decisions prompted no feedback, representing implicit agreement. Trust was measured using the 12-item Checklist for Trust between People and Automation. Rejection of explicit recommendations, as a proportion of all such explicit cues, was negatively correlated with trust (r(41) = −0.62, p < 0.001), while acceptance was positively correlated (r(41) = 0.47, p = 0.001). The proportion of silent agreements showed no association with trust (r(41) = −0.02, p = 0.895). These results suggest that explicit feedback—both confirming and corrective—acts as a key cue for calibrating trust, while implicit agreement carries little weight. Trust appears more sensitive to how the system communicates than to whether its decisions align with those of the user. This aligns with recent findings that transparency, not just accuracy, drives trust in AI. Designing DSS that strategically balance explicit feedback with minimal intrusiveness may enhance operator trust and performance. Future research will manipulate feedback valence and visibility in a between-group design to further disentangle how communication style shapes trust in high-stakes human–AI collaboration.
ObjectiveThis narrative review examines the cognitive, metacognitive, and team competency requirements that may contribute to productive and reliable collaboration between human and AI to address two questions: What capabilities make AI a competent collaborator? What makes humans ready for AI collaboration?BackgroundAs AI systems are increasingly integrated into workplaces and framed as teammates rather than tools, humans face challenges that include maintaining situation awareness, calibrating trust, and working with systems that may surpass them cognitively. We analyzed Human-Agent Teaming (HAT) readiness around two complementary levels: operational team competencies (communication, coordination, and adaptability) and regulatory capacities (trust calibration and metacognitive awareness).MethodWe conducted a structured narrative review of literature from 2010 through January 2026, searching Google Scholar, Scopus, PsycINFO, IEEE Xplore, ACM Digital Library, and Semantic Scholar, complemented by forward citation tracking. After screening 572 records, 192 articles were included for synthesis.ResultsCommunication inflexibility, limited shared understanding, and trust miscalibration emerge as recurring barriers to HAT, while regulatory capacities (trust calibration and metacognitive awareness) represent particularly critical dimensions of HAT readiness that remain to be fully operationalized.ConclusionHAT requires mutual readiness, with both humans and AI developing metacognitive and adaptive capabilities. Despite methodological heterogeneity limiting clear conclusions, cross-training and co-learning methods offer a promising avenue for building shared understanding and calibrated collaboration.ApplicationThis review provides practical principles for designing AI systems that support calibrated collaboration and for preparing humans to work adaptively with AI, thereby enhancing team effectiveness, reliability, and resilience in collaborative work environments.
This narrative review integrates evidence from cognitive science and AI research to challenge commonly accepted dichotomies between human and artificial cognition, such as the assumed divide between genuine human understanding and mere machine pattern matching. Instead, we propose a view that recognises similarities in their cognitive architectures and processes. Human and artificial cognition seem to operate through comparable mechanisms, as both rely on statistical processing, associative pattern recognition and approximation rather than perfect logic. Through a systematic comparison of core cognitive domains across 363 articles, we highlight parallels in capabilities and limitations, including shared vulnerabilities to biases, memory distortions and decision-making opacity. We critically examine popular narratives such as the stochastic parrot argument and the myth of human rationality. These positions often rely on idealised views of human cognition that are contradicted by cognitive and neuroscientific evidence. This review calibrates expectations of both human and artificial systems by moving beyond both AI alarmism and human exceptionalism towards a more empirically grounded perspective on cognition. Our comparative review acknowledges both the shared statistical foundations of intelligence and differences in embodiment, intentionality and phenomenological aspects of cognition. This perspective has implications for human–AI collaboration, cognitive performance benchmarking and research on AI transparency.
Simulator-based training is increasingly used to maintain marine safety and compliance, yet timely assessment of trainee proficiency remains resource-intensive and difficult to scale. This work investigates AI-based trainee proficiency assessment in small-vessel training simulations. Policy capturing is applied to learn how instructors evaluate performance from simulator measures, enabling automated estimation of trainee proficiency aligned with expert judgment and suitable for near-real-time feedback and after-action review. Policy capturing sessions aimed to collect expert judgments on representative samples of synthetic training data logs. This data collection aimed to produce trainee proficiency evaluation models for 9 distinct tasks in the training program. Expert annotations were provided on a continuous proficiency scale and discretized into three levels using fixed thresholds to obtain both categorical (entry/intermediate/proficient) and continuous (percentage) proficiency estimations. Results were analyzed with and without the application of a data denoising procedure aimed at correcting judgmental inconsistencies. Supervised models were trained and selected using repeated 10-fold cross-validation. The best regression model achieved strong performance (mean R^2=0.90 , mean RMSE=2.43 on a 50–100 proficiency scale), compared to a mean R^2=0.70 and mean RMSE=4.76 without denoising. Categorical classification reached a mean accuracy of 94
Security surveillance is characterised by substantial cognitive challenges to operators. Scantracker is a mixed-reality gaze-aware support tool that alerts surveillance operators to neglected cameras, attentional tunnelling, and vigilance decrements. Initial research efforts were conducted in simulated environments to examine the effects of Scantracker on surveillance performance; however, this tool has yet to be deployed and tested in a real-world operational environment. In the current study, we tested Scantracker in an airport operations centre to assess the feasibility of its integration and to collect expert feedback regarding its operational relevance. Operators used Scantracker voluntarily during their work shift, while gaze data enabled system notifications. They provided ratings on perceived utility, workload and ergonomic quality along with qualitative feedback on their experience. The pattern of results highlights the potential of Scantracker to support surveillance operators and demonstrates the value of user-centred field testing for developing intelligent monitoring assistants.
Integrating AI into decision-support systems (DSS) for safety-critical domains like aviation requires aligning system behavior with pilot mental models to provide relevant information. Using the Cognitive Shadow—a DSS that models operator decisions and notifies discrepancies—we evaluated a novel knowledge-elicitation technique: the inverse counterfactual. After selecting their preferred option, users modified a single factor to make their second-best option preferable, creating paired cases across their decision boundary. In a simulated adverse-weather avoidance task, 44 participants completed 130 baseline trials and generated counterfactuals for 20 additional cases. Contrary to expectations, the current implementation of the technique did not enhance human-AI model similarity, as measured by the degree of agreement in a 20-case test phase. However, when counterfactuals involved minimal edits—remaining near the decision boundary—predictive accuracy improved and DSS recommendations were more often accepted. Larger edits degraded performance. These findings demonstrate the feasibility of counterfactual elicitation for improving model alignment with user mental models.
We investigate a predict-then-optimize method for ship refit project scheduling, integrating machine learning (ML) task duration predictions. Ship refit operations encompass various tasks such as renovation and repair in shipyards, which have become increasingly important recently. Efficient scheduling of these tasks relies heavily on accurate task duration estimates from domain experts. Our study focuses on assessing the impact of using ML algorithms for these estimates on ship refit project schedules, evaluated using a periodic re-optimization approach based on actuals, leveraging a constraint programming model in the optimization process. We compare three methods: expert estimates, historical data forecasting, and augmented human-AI forecasting. Through experimental analysis, we evaluate their performance in predicting task duration and improving the robustness of ship refit project schedules. The results demonstrate the benefit of the augmented human-AI model in providing task duration predictions, yielding more robust ship refit project schedules and limiting the disruptions in resource allocation over the planning horizon. Additionally, the experimental results indicate that the predict-then-optimize approach enhances the robustness of ship refit project schedules, suggesting the approach's potential to be readily generalized for optimization in other project scheduling domains, such as aviation maintenance.
Artificial intelligence (AI) systems need to adapt to changing circumstances to maintain relevance in dynamic environments. Inspired by the adaptive advantages of human forgetting, this study investigates the integration of a forgetting function into an AI system. We implemented this mechanism as a training window within the Cognitive Shadow (CS) system, an AI designed to learn and emulate human decision models. This training window hyperparameter—applicable to supervised machine learning algorithms—aims to address the issue of concept drift by prioritizing recent information. The effectiveness of this addition was tested with a simple strategy game similar in dynamics to rock-paper-scissors. Participants played individually against an AI opponent for three 60-round sessions. CS was trained during Session 1 to learn the decision patterns of the player and actively predicted and countered human decisions in Sessions 2 and 3. Analyses showed that including the training window significantly improved prediction accuracy in both Sessions 2 and 3 by emphasizing recent, relevant data. These findings highlight the potential of incorporating human-inspired forgetting mechanisms to enhance AI performance in interactive and dynamic environments, with implications for future decision support systems.
Complex problem-solving (CPS) skills – the ability to comprehend, manage, and adapt to complex, evolving situations – is essential in the 21st-century workplace. However, empirical evidence shows individuals are limited, computationally and cognitively, in managing complex systems (e.g., delayed feedback, nonlinearity, and conflicting goals). Traditional cognitive tasks are deemed too simple and often fail to capture these properties, whereas field studies lack control over conditions and measurement. Microworlds offer controlled complexity: they can reproduce properties of complex systems but remain tractable for systematic manipulation and data collection. We wish to present and demonstrate CODEM (COmplex DEcision Making), a microworld platform designed to simulate complex dynamic systems at different levels and trace the cognitive processes underlying CPS and decision-making. CODEM serves as a testbed where one can build environments with customizable variables, feedback loops, semantics, and opacity. The platform provides performance (e.g., comparing human goal attainment with random simulations), cognitive process-tracing (e.g., use of heuristics) and behavioural (e.g., structural information seeking) logs, intelligent tutor extensions (e.g., for system thinking training), and multiplayer options for collaborative problem solving. CODEM has three key applications. First, as a research tool, it enables systematic study of CPS and decision heuristics under complexity. Second, as a training and awareness tool, it highlights pitfalls in reasoning (e.g., most individuals assume linearity and neglect delayed effects) while promoting system thinking and metacognitive strategies. Third, as a personnel selection tool, it holds promise as a measure of the capacity to manage complexity beyond general intelligence. Results from a series of experiments showed that participants perform poorly on CODEM complex scenarios, often close to or even below chance levels, despite the presence of system thinking behaviour. These results are in line with the view that complexity poses a dire cognitive challenge, highlighting the need for tools to assess, train, and support CPS.
Personnel selection and agent readiness are critical priorities for defense organizations, particularly in ensuring staff are adequately prepared for high-stakes environments such as public safety and military operations. Readiness assessments evaluate cognitive abilities, mental resilience, physical fitness, decision-making skills, and the ability to manage stress and workload. Traditional methods rely on performance indices from psychometric tests and physical evaluations. However, research suggests that incorporating behavioral and neurophysiological data, like cerebral blood flow oxygenation, enhances predictive capabilities for agent readiness. Methods such as the Revised Multi-Attribute Task Battery (MATB-II) have been developed to assess human decision-making and complex problem-solving capabilities, which are skills often considered when assessing readiness. The SHAD (Sensing Humans for Augmented Debrief) system is a training-support interface developed for such a multimodal approach that integrates wearable technology with psychometric profiling. SHAD captures neurophysiological signals-such as heart rate variability, brain oxygenation, and respiratory data-alongside psychometric assessments to predict and evaluate agents' mental and physical states in real time. The experimental design used MATB-II as a task to induce stress and workload in 41 participants from two public safety organizations. A machine learning model enabled readiness predictions, offering insights for personnel selection, mission preparation, and adaptive training. The model was trained on both psychometric data and neurophysiological data collected during the MATB-II task. This solution could significantly enhance public safety and defense missions by optimizing agent performance.
The modernization of Command and Control (C2) for North American Aerospace Defense (NORAD) entails supporting operators with trustworthy AI-based solutions that complement, rather than replace, human abilities. New forms of threats requires more than ever the ability to quickly derive actionable situational awareness from a set of heterogeneous sensors. In this paper, we investigate the use of Human-Automation Teaming (HAT) for achieving accurate, timely continental surveillance in the context of NORAD critical infrastructure protection. We developed a collaborative AI agent solution with awareness, anticipation and decision capabilities, augmented here with the ability to learn human decision policies for improved collaborative situation assessment. The study employs a simulated threat evaluation task to validate the effectiveness of the augmented AI-agent using a multi-model approach combining seven supervised machine learning algorithms. Results show that the policy capturing method classified threat levels with a predictive accuracy of 95% while considering three different types of targets (UAV, drone swarm, small aircraft). We conclude that integrating the policy capturing capability into a collaborative AI-agent constitutes a key step toward enabling a novel human-AI co-learning process for adjustable human-autonomy teaming.
Maintenance task duration estimations help manage shipyard resource usage and allow planners to decide on maintenance priorities within a limited time frame. Better estimated task durations help produce more robust resource schedules, perform more tasks in facilities such as shipyards, reduce resource idling time and increase ship operational availability. However, task duration estimations have until now been historically performed by human experts with essentially no artificial intelligence-based forecasting for shipyard operations. The analysis of historical data is also not a common practice to complement any expert-driven forecasting. To explore opportunities for using AI in this work domain, and to improve on human estimations for task durations, we propose a novel hybrid Human-AI approach that involves integrating human forecasts with data-driven models. Our empirical data comes from two fleet maintenance facilities in Canada, containing more than 13,000 anonymized historical ship work orders (WO) ranging from 2017 to 2022. We used supervised learning algorithms to forecast the preventive maintenance task duration on this data, with and without expert task duration estimates and the results demonstrate that hybrid models perform better than both human expert model and historical data alone. An average of 8.6% improvement from Hybrid human-AI model over human expert model is observed based on R2 evaluation metric. Results suggest that human forecasts, which tend to rely on a broader contextual knowledge than the inputs captured in a historical database, remain key for effective task duration estimation, yet can be fine-tuned by the pattern recognition capabilities of machine learning algorithms.
We present a modular architecture that enables advanced surveillance functions exploiting data collected from heterogeneous sensors dispersed over multiple, often mobile platforms in the field. Examples of such functions are red forces tracking with surveillance gaps, detection of different types of anomalies, search and rescue operation monitoring, and threat alerting. This novel approach combines a distributed fusion engine, an intelligent process manager, and a system of ruggedized computers, enabling information processing in the tactical domain. The hybrid AI-based heterogeneous fusion engine consists of different algorithms, including various detectors and classifiers, represented as services in a light-weight information management and interoperability layer. This architecture layer enables context-dependent discovery of the right sensing and processing services at runtime that are combined using a robust Bayesian fusion layer exploiting complex correlations in the data. The discovered services are distributed over a network of computing nodes by an intelligent process manager, which optimizes network resource allocation according to communication and processing capacities. The fusion engine and the process manager are delivered to the tactical domain using the ruggedized SOTAS computing and communication infrastructure, achieving efficient, actionable, timely, and consistent situation awareness in constrained domains, such as military vehicles. Citation: G. Pavlin, R. Boudreault, A. Penders, M. de Graaf, D. Lafond, A. Swiebel, “Extracting Actionable Information From Heterogeneous Sensors in the Field: A Distributed Hybrid AI Approach in Constrained Domains,” In Proceedings of the Ground Vehicle Systems Engineering and Technology Symposium (GVSETS), NDIA, Novi, MI, Aug. 15-17, 2023.
Sonar sensors use sound waves to sense objects underwater, and are critical to anti-submarine warfare by detecting, locating, and tracking enemy submarines. The performance of sonar systems is measured as the maximum acoustic range – the distance at which a sonar system can detect the sound waves emitted by a submarine. Knowing the predicted acoustic range is essential for maintaining adequate awareness and understanding of the underwater environment, which is critical in anti-submarine warfare. However, predicting acoustic range accurately is challenging due to the complex and dynamic underwater environment. In this study, we investigate how experts in anti-submarine warfare assess the credibility of acoustic range predictions. Four Royal Canadian Navy experts judged the credibility of acoustic range predictions in 480 unique cases. A software tool for modeling human judgment patterns ("Cognitive Shadow") was used to infer how experts' decide whether predicted acoustic range is credible. Analysis of the resulting decision model revealed that experts consistently prioritized specific factors when assessing acoustic range prediction credibility. This study supports the integration of the Cognitive Shadow as a decision support tool for sonar performance assessment in anti-submarine warfare.
Human-autonomy teaming (HAT) is becoming a subject of high interest in the human factors literature. It has several applications, including the collaboration between a human and an autonomous unmanned aerial vehicle (UAV) for security and defence use cases (e.g., for search and rescue tasks). This work is focused on methods for task-allocation between human and autonomous UAV agents. The proposed approach is human-centred, using a coactive design framework which relies on enabling adaptive team dynamics where different agents might act as key players for specific tasks based on an interdependent relationship. This method helps solve complex issues in understanding and adjusting to complementary team dynamics where agents might have different skill levels, experiences, roles, and helps understand which agent is more competent to perform a task. Additionally, such a framework promotes transparency towards the control and task-allocation strategies. To demonstrate this task-allocation strategy, this study looked at the use of neurophysiological features as indicators of task-specific capacities in UAV operations, more specifically electroencephalogram (EEG) signals, which opens up for the development of task-allocation adaptive systems, dependent upon variations in brain activity. Results found that EEG spectral power bands have potential to help determine different task-based abilities across groups (i.e., obstacle avoidance vs. target identification), hence contributing to pinpointing variations in the type of autonomous support needed. Overall, this research explores how task-dependencies can be observed through EEG signals for better transparency and explainability of adaptive control in pilot-AI teaming.
The development of autonomous vehicles such as Unmanned Aerial Systems (UAS) are becoming a major tool for air and information superiority in defence and threat anticipation. Human-Autonomy Teaming (HAT) systems are crucial for the success of team collaboration, trust and mission execution in command and control (C2) operations such as Intelligence, Surveillance and Reconnaissance (ISR). The present work focuses on understanding when and how control should be allocated to autonomous agents. Intelligent adaptive control methods between human and autonomous agents are introduced. The proposed method consists in developing a HAT system based on the co-active design framework. A prototype system was implemented and tested in a human-in-the-loop experiment involving a simulated ISR mission. Findings helped assess individual capacities and detect behavioural patterns for adaptive control. More specifically, it helped recommend the level of support and control to be allocated for each agent. Results focused on the human operator and helped gain insight about how task-allocation strategies could be implemented within a complex ISR operation by breaking it down into sub-tasks. A key outcome of this work to help augment team interoperability in C2 operations. This paper was originally presented at the NATO Science and Technology Organization Symposium (ICMCIS) organised by the Information Systems Technology (IST) Panel, IST-205-RSY - the ICMCIS, held in Koblenz, Germany, 23–24 April 2024.
Video surveillance can be cognitively very demanding as it imposes operators to stay focus for a long time and to provide the right response for relevant stimuli among a lot of information. In this challenging activity, it is relevant to consider the use of decision-aid techniques to improve operators’ alertness. The purpose of this study was to examine the impact of a real-time gaze-based tool named Scantracker—which can identify instances of neglect, over-focus and vigilance decrement using eye tracking and display visual notifications to mitigate such situations—on surveillance performance measures during a surveillance simulation. Augmented reality glasses were used to monitor eye movements in real time for all non-expert participants, but notifications presentation to support attention was visually active for only half of them (Scantracker group), as opposed to the control group without support from the Scantracker. No significant differences were observed across those two groups. However, a within-group comparison contrasting trials with active notifications vs. a silent condition showed a reliable improvement in task accuracy and a reduction in screen neglect duration. Results are discussed in light of potential applications of Scantracker with augmented reality.
The Cognitive Shadow is a decision-support system that uses policy capturing to model human operators' judgment policies and provide online predictions of their decisions. The system can provide support in reaction to a decision mismatch (shadowing mode) or proactively (recommendation mode). The goal of this study was to compare these two modes of operation in their ability to effectively model and support decision-making and to examine impacts on information processing, workload, and trust. Participants took part in an aircraft threat evaluation simulation without decision support or with the Cognitive Shadow (either shadowing or recommendation mode). Dwell time was collected over different areas of the user interface. While the recommendation mode had no advantage over the control group, the shadowing mode resulted in greater human and model accuracy. This mode led to longer dwell time over the parameters zone presenting key information for decision-making. These benefits were maintained even after the tool was removed. Workload was unaffected by the mode, and while trust was initially higher in the recommendation mode, it quickly became equivalent between both modes, overall supporting shadowing as the better configuration for cognitive assistance. Results are discussed in terms of decision processes, operators support, and automation bias.
Task duration estimation is an important element of scheduling and optimization in various task domains. However, work completion times may vary based on endogenous and exogenous factors. Such estimations are typically made by human experts and often entail assumptions and uncertainty. Such imprecision causes either longer or shorter than expected task durations leading to scheduling conflicts or inefficiencies. As an example, in the domain of ship refits, certain tasks do not take place very often. When there is little historical data available, task scheduling and planning can be very challenging and require ongoing replanning effort to compensate for estimation errors. While human experts can provide synthetic cases to train forecasting tools, it is not obvious how one can integrate human forecast with historical cases for decision support purposes. A new forecasting method is introduced in this paper that integrates human experts’ inputs with historical data to create a hybrid model for forecasting and scheduling purposes. We demonstrate through experiments that the proposed hybrid model increases prediction accuracy by 5 to 10% compared to forecasting with only historical data.
Abstract: Single-pilot operations are cognitively challenging for pilots and could benefit from decision-support tools to mitigate risk-prone situations. The Cognitive Shadow is a prototype tool that employs policy capturing, a data-driven technique used to model decisions, to learn users’ judgement policies and alert decision discrepancies from one’s decision pattern. This proof-of-concept study investigates the potential of policy capturing to model pilots’ policies facing unstable approaches. Pilots were presented simulated cases and asked whether to continue descent or to go-around while the policy-capturing tool learned their decision pattern and provided feedback. Individual models reached mean predictive accuracy of ~ 89% while the group model reached 100%. These results speak to the potential of extracting pilots’ knowledge using policy capturing to create decision aids.