This study investigates the development of monotony in air traffic control (ATC) using a multi-modal measurement setup applied in a field study at two operational towers of differing traffic intensity: Tartu (low-traffic) and Tallinn (medium-traffic). The results provide a first step toward establishing a reference baseline for assessing monotony in uneventful low-traffic tower environments using an evidence-based approach. Monotony is conceptualized as a state of reduced physiological and cognitive activation resulting from low task stimulation, repetitive and uneventful operational contexts. We operationalized this concept using eye-tracking indicators, including blink duration, blink frequency, percentage eye closure (PERCLOS), and saccade velocity. Other measures included cardiac indicators derived from electrocardiography (ECG), the Karolinska Sleepiness Scale (KSS), and the Psychomotor Vigilance Task (PVT). These measures were collected from seven air traffic controllers across multiple shifts over an eight-day period.Linear mixed-effects models were fitted to each indicator, revealing consistent time-on-task effects at Tartu across several modalities, including blink duration, PERCLOS, HR (HR), root mean square of successive differences (RMSSD), and standard deviation of normal-to-normal intervals (SDNN). These findings support a pattern of growing parasympathetic activation and reduced arousal associated with monotony. In contrast, Tallinn showed signs of active physiological regulation under higher workload conditions, such as elevated baseline HR and decreasing SDNN, indicative of sustained engagement. The sympathovagal balance (LF/HF) ratio increased over time at both sites, but in the context of stable or increasing RMSSD/SDNN at Tartu, this may reflect compensatory physical activity rather than sympathetic dominance. Monotony indicators differed significant between towers, with up to a 7.0-fold stronger temporal change in PERCLOS and a 1.4-fold difference in SDNN at Tartu compared to Tallinn.The study further discusses methodological feasibility in sterile ATC work environments and highlights how operational differences, such as task rotation and self-directed off-position activities, impact data collection and interpretation. Findings underscore the need for tailored monotony mitigation strategies, particularly in low-traffic towers, and advocate for continued research in operational environments to capture complex real-world phenomena.
Fatigue in air traffic control (ATC) is a critical safety concern, yet the influence of socio-demographic factors such as age and gender remains underexamined. This study investigated how age and gender shape both subjective and objective fatigue in air traffic controllers (ATCOs) across operational shifts. Data were collected from 73 participants at a Scandinavian Area Control Center. Subjective sleepiness was assessed using the Stanford Sleepiness Scale (SSS), and objective fatigue was measured using a 3-minute Psychomotor Vigilance Task (PVT-B), administered at three timepoints per shift (pre, mid, post). Linear mixed models were used to analyze the effects of age (dichotomized at the median: ≤ 45 vs. ≥ 46 years) and gender on fatigue outcomes. Results showed that older ATCOs exhibited slower reaction times (RT), indicating greater objective fatigue, while reporting lower self-rated sleepiness (particularly during night shifts) than their younger counterparts. Younger participants, in contrast, showed faster RTs and longer on-shift sleep, yet reported higher self-rated sleepiness at specific timepoints. Gender differences were less consistent but evident in sustained attention: males showed systematically faster RTs and longer on-shift sleep during night shifts. No overall gender effect was found in self-rated sleepiness, although isolated timepoint differences emerged. Interactions between age, gender, shift type, and timepoint revealed context-specific fatigue patterns. These findings highlight the multidimensionality of fatigue and suggest that age and gender shape fatigue in nuanced, time-dependent ways. Integrating socio-demographic variability into fatigue risk management may support the development of more adaptive and equitable shift designs and countermeasure strategies in safety-critical settings.
In this paper, we examine the feasibility of assessing air traffic controller (ATCO) workload (WL) using non-intrusive eye-tracking measures and machine learning (ML) algorithms. We concurrently acquire electroencephalography (EEG) data from a workloadoptimized wearable device and subjective WL assessments through self-reported Cooper-Harper scale (CHS) workload-rating scores, employing both as label variables. A sample of n = 18 ATCOs participate in simulated work sessions encompassing tasks designed to induce three distinct task-load levels: light, moderate, and heavy. We evaluate the performance of five classical ML models. Focusing on the best-performing models, we apply feature selection techniques to identify reduced sets of eye-tracking features. Starting with 58 features, we use a recursive elimination method based on permutation importance, aiming to determine the minimal feature set while also striving for improved performance. The outcomes yielded promising results in the realm of workloadlevel estimation, achieving 96% accuracy (f1-score=0.87) with 34 features for high workload prediction and 88% accuracy (f1-score=0.82) using 57 features in predicting 3 different levels of workload. We further reduced the feature sets to 6-13 features for different tasks with minimal impact on performance. We identified a \x93knee point\x94 as the optimal balance between model performance and dimensionality. Adding more features beyond this point did little to improve performance, but increased model complexity. These results indicate that even a small number (less than 10) of features can be sufficient for WL prediction.
Detailed reviews and responses can be found in the PDF and HTML versions of this document. The DOI for the original paper is \url{https://doi.org/10.59490/joas.2025.8034}
In this paper, we examine the feasibility of assessing air traffic controller (ATCO) workload using non-intrusive eye-tracking measures and machine learning algorithms. A total of N = 18 ATCOs participated in simulator runs with tasks inducing three task-load levels: light, moderate, and heavy. Task load was modulated through traffic load and the associated increase in complexity. We collected eye-tracking data (statistical summaries of which serve as features) and obtained subjective workload assessments using self-reported Cooper-Harper Scale scores, which act as label variables. We evaluate the performance of eight classical machine learning models, with the k-nearest neighbors and support vector classifier models emerging as the most promising. To optimize performance, we apply feature selection techniques, focusing on these best-performing models. Feature selection via recursive feature elimination (RFE) based on permutation importance reduces the original 42 features while maintaining or improving performance. The outcomes yield promising results in workload-level estimation, achieving an F1 score of 0.870 for low/high workload prediction and an F1 score of 0.788 for predicting three different levels of workload. The RFE process identifies optimal feature sets ranging from 7 to 13 features for different tasks, with minimal impact on performance. A “knee point” is observed, representing the optimal balance between model performance and dimensionality. Adding more features beyond this point contributes little to performance improvement while increasing model complexity. These findings indicate that even a few features can be sufficient for accurate workload prediction. We show that head-movement features provide valuable information. Comparable performance is achieved using only ocular features, but this requires more features. Asymmetry in left and right eye metrics holds workload-related information but transforming them into averages and differences reduces performance. Retaining the original features separately is the most effective approach, incorporating their absolute differences may provide slight benefits in certain models.
Fatigue is a longstanding issue in air traffic control (ATC), closely associated with shift work and time-related factors. However, the dynamics of fatigue across morning, evening, and night shifts in an area control center (ACC) remain largely underexplored. This study examined sleep duration and fatigue progression across different shift types. Both objective (three-minute Psychomotor Vigilance Task, PVT-B) and subjective (Stanford Sleepiness Scale, SSS) measures were conducted at the beginning, middle, and end of each shift. Results indicated that pre-shift sleep duration was shortest before night shifts, likely increasing sleep pressure and reducing alertness during the window of circadian low (WOCL). Subjective fatigue remained stable throughout morning shifts but increased towards the end of evening shifts, reflecting circadian influences. Night shifts exhibited peak fatigue during the WOCL, driven primarily by circadian rhythms rather than task load. Objective measures revealed a mid-shift decline in performance, with only partial recovery in the latter half of night shifts. Compared to day shifts, night shifts resulted in significantly higher fatigue levels, underscoring the critical role of circadian rhythms in fatigue dynamics. These findings highlight the need for targeted fatigue mitigation strategies that address circadian vulnerabilities and irregular sleep patterns in ATC shift systems.
This study explores automation surprise in low-level air traffic control automation by comparing it to flight decks, which have higher automation levels. Automation surprise occurs when an operator's expectations do not match the actual behavior of the system. This phenomenon is recognized in aviation but not in air traffic control due to its lower automation levels. The study surveyed 47 en-route air traffic controllers and found that automation surprise occurred approximately once every 12 shifts. Causes included System Malfunctions (42%), False Display (21%), Unclear Display (23%), and complex tools, such as medium-term conflict detection (48%). Unlike pilots, who experience action-related automation surprise, controllers experienced information-related automation surprise from false or misinterpreted data. This posed lower risks, as discrepancies in expectation were often quickly detected. The study proposes a safety model based on Reason's "Trajectory of Accident Opportunity," in which detecting the "surprise" event serves as a critical barrier for further recovery actions. These insights can inform AI-supported decision-making in ATC and highlight the need for tailored safety measures across automation levels.
In this paper, we validate non-intrusive eye-tracking and head-movement indicators for predicting workload (WL) to support an air traffic controller (ATCO) self-evaluation. This will allow, e.g., to open or close sectors under more accurate consideration of the individual's state, identifying over- and underload risk. We investigate n = 18 ATCOs during simulated working sessions with three varying traffic-load scenarios (light, moderate, heavy task load), adhering to a counterbalanced, within-subjects design. We apply non-intrusive eye tracking and the Cooper-Harper WL Rating Scale (CHS). Further, we employed a wearable electroencephalography (EEG) device optimized for monitoring WL. We evaluate the performance of five classical machine learning models across two distinct labeling tasks: CHS and EEG. For CHS labeling, the models achieve an accuracy of 84% (F1-score 72%) when classifying WL levels into three categories (low, medium, high), and 93% (F1-score 83%) when categorizing WL into two classes (low/medium, high). Similarly, for EEG labeling, the models achieve an accuracy of 86% (F1-score 77%) in the three-level WL classification and 96% (F1-score 84%) in the binary WL classification. With this, we display the potential of machine learning techniques in predicting ATCO WL solely based on eye-tracking and head-movement measures.
Abstract: The study investigated sleepiness and fatigue in a single split nightshift arrangement in air traffic control. The arrangement included two mirrored nightshift types, with sleep and operational phases alternating once mid-shift. Sleep duration, subjective sleepiness, and fatigue as well as sustained attention, which have not been thoroughly studied to date, were examined. The findings suggest different sleep strategies. Working during the second part of the night resulted in reduced sleep duration before the shift. Subjective fatigue exhibited varying patterns depending on the shift type and time, with elevated fatigue observed at the start and middle of the shift. However, subjective sleepiness and sustained attention were primarily influenced by the passage of time.
We propose a new approach to evaluate ocular measurements of air traffic controllers (ATCOs) as potential workload and fatigue indicators. We employ the Fast Fourier transform (FFT) to test our assumption that humans respond to increasing fatigue with harmonic oscillations in the eye movement, while they respond to increasingly high workload with disruptions to these harmonic oscillations. The FFT yields the frequency spectrum and we suggest to use the center of gravity of this spectrum to capture the variations. We give a proof-of concept study to evaluate our approach and we were able to verify our hypotheses in some cases, in particular, we identify the fixation duration as a promising indicator of changes in workload.
Automation in Air Traffic Control (ATC) is gaining an increasing interest. Possible relevant applications are in automated decision support tools leveraging the performance of the Air Traffic Controller (ATCO) when performing tasks such as Conflict Detection and Resolution (CD&R). Another important area of application is in ATCOs’ training by aiding instructors to assess the trainees’ strategies. From this perspective, models that capture the cognitive processes and reveal ATCOs’ work strategies need to be built. In this work, we investigated a novel approach based on topic modelling to learn controllers’ work patterns from temporal event sequences obtained by merging eye movement data with data from simulation logs. A comparison of the work phases exhibited by the topic models and the Conflict Life Cycle (CLC) reference model, derived from post-simulation interviews with the ATCOs, indicated that there was a correspondence between the phases captured by the proposed method and the CLC framework. Another contribution of this work is a method to assess similarities between ATCOs’ work strategies. A first proof-of-concept application targeting the CD&R task is also presented.
Real-world event sequence data, such as activity logs, eye-tracking data, simulation data, and electronic health records, often share characteristics such as a large alphabet of events, fragmentation, noise, and high complexity which makes them difficult to analyze in their raw form. Because of this, simplification and preprocessing through various data transformations are commonly required before the data can be effectively visualized and analyzed. Existing methods for such data transformation are either manually applied and rely heavily on user expertise, or use algorithmic approaches to apply bulk operations which can imply the loss of potentially important information without users being aware. To bridge this gap, we propose a visual analytics approach that aims to successively increase the quality of noisy event sequences by supporting an interactive, context-aware application of data transformations. This is achieved by providing cues concerning the potential loss of information that transformation operations may imply and allowing users to explore, and visually assess their impact on the data. Therefore, a central feature of the approach is that users can tune the data transformation process so that important identified data characteristics are preserved. We motivate the proposed approach in the domain of air traffic control and illustrate it through a usage example, using event sequences derived by merging eye-tracking and simulator data from a human-in-the-loop simulation experiment with 14 air traffic controllers.
The analysis of operator’s work patterns in safety-critical domains is increasingly assisted by eye-tracking technologies. A reason is a growing demand for empirically justified assurance to ascertain an equal level of safety after the implementation of new techniques and higher levels of automation. In the field and simulation studies, head-mounted eye-tracking devices are a preferred choice because of the easy, reliable and fast startup. The downside is time-consuming work to map gaze points to the world/workplace coordinates which makes head-mounted devices applicable to a small number of episodes to analyze. This paper presents a solution called AoI Mapping and Analysis Tool (AMAT) that relies on visual features provided by the integrated scene video camera. AoI-templates can be defined from the video recording and are constantly matched for an AoI-analysis. AMAT was evaluated with a field study example from Air Traffic Control (ATC) in the tower at Linköping City airport and a training simulator study example from Vessel Traffic Service (VTS) in Gothenburg. Example situations demonstrate the capabilities of AMAT to define AoIs, evaluate the validity of the chosen set of AoIs and export to an example analysis program Eloquence for the visualization of results. The AoI-sequence of 4 vessels encounter in the Gothenburg archipelago was shown as well as the comparison of four tower controller’s sequences during the final approach. The final discussion highlights the capabilities and limitations of the implemented techniques that rely on the AKAZE point feature detection and linear homographic transformations. A major source of disturbance lies in the varying light conditions evoking under-exposure of the video camera and fast head movements.
Current practices of risk analysis of novel socio-technical systems rely on the subjective judgment of experts. With a view on the complex interactions between human operators and the environment in ATM, a method is needed for gaining empiric evidence directly from operations. Risk analysis that bases on Human-In-The-Loop-Simulations offer a promising approach by providing an environment in which the novel system can be applied safely. An inherent disadvantage is the effort needed to cope with the strict safety targets in ATM, e.g. 1.88E-8 accidents per operating hour in which safety metrics are subject to the statistic problem of Right Censoring. This paper presents our novel concept to modify conditions of the simulation for gaining a calibrated acceleration effect by which the probability of safety metrics can be estimated from a shorter experimental period. This is motivated by the methodologies of Accelerated Life Testing, in which the Mean-Time-To-Failure of products is forwarded into the experimental period by applying calibrated steps of stress-load. We developed an experimental design that applies a procedure for the induction of a calibrated time-pressure for the stimulation of human error. The results of the proof-of-concept-study show controllable stress-reactions of the test persons.
The purpose of this study was to investigate the effect of shifting among and updating of mental models in multi remote tower operations. Background. Within the development of the future workplace for air traffic controllers (ATCO), innovation shifts to multi remote tower operations. Multi remote tower means, that an ATCO serves air traffic control services for more than one airport at a time from one physically remote located workplace. This change requires adjustments in the way how the ATCO works, as multiple mental models need to be updated constantly. Frequent shifting between airports and mental models respectively is necessary. Updating and shifting are cognitive cost sensitive, as they affect workload and situational awareness negatively. In contrast, higher workload could be beneficial for alertness which could act as a countermeasure in turn. Method. Four 90-min lasting remote tower Human-In-The-Loop simulation runs with traffic and weather were performed by eight conventional tower experienced ATCOs. Each ATCO completed two runs in multi- and two runs in single-mode. Situational awareness, workload and alertness was measured via self-reports before, during and after each run. Results. No differences between both modes with respect to situational awareness and alertness could be found. However, significant workload differences could be found during the simulation runs at two times, due to a simulated snowstorm. Conclusion. The findings indicate no negative effect of shifting among and updating of mental models in multi remote tower. A possible explanation could be, that a common hybrid mental model for in multi-mode is internally developed so that shifting among mental models is perhaps not necessary.
— We do a field study on controller workload in a conventional tower and a Remote Tower environment (in both single and multiple mode) and give a proof of concept for the validation of indicators on their workload predictability. We analyze the number of ATCO tasks (e.g., arrivals, taxi), the communication times related to different ATCO tasks (and use them as weights for the ATCO tasks), and reaction times to SPAM queries. We show that—while the pure number of ATCO tasks is not a necessary condition for an increase in workload rating—indicators that integrate the communication time related to these ATCO tasks are, that is, each increase in workload rating is accompanied by an increase in these indicators.
Eye-Tracking experiments have proven to be of great assistance in understanding human computer interaction across many fields. Most eye-tracking experiments are non-intrusive and so do not affect t ...
The novel multi remote tower concept involves the control of two airports by one tower controller from one remote workplace at a time. In order to implement a multi remote tower into operations, a safety assessment is crucial to evaluate existing risks. Since there is currently no operational experience available concerning this concept, the hazard identification and risk mitigation remains hypothetical. However, empiric data is needed for evaluating and focusing on the safety-relevant hazards that are multi remote tower specific. To close this gap, we developed the MERASSA concept for gaining evidence on the safety-relevance of hazard using Human-In-The-Loop simulations and stress test scenarios. The method was assessed through a validation study at the multi remote tower case using eight identified hazards that are human-issue originated. In total 32 simulation runs with eight rated and experienced tower controllers were carried out. The results of the study show the ability of the tower controller to compensate risk by slowing down the work speed. No hazard could be verified through a comparison of the multi and single runway baseline scenario. Additionally, the results indicate a clear lack of confidence of the tower controller to control two airports at a time due to the need to share attention across the work environment. The comparison of the empiric and subjective results show equal trends which are a sign for the success of applying the method. However, a major drawback of using simulations and stress test scenarios are the enormous efforts needed to control the conditions of testing. Keywords-Safety Assessment; Socio-Technical Systems; Human Error; Multi Remote Tower; Air Traffic Control
Although where to look, when, and in what order is crucial for situation awareness and task performance in tower control, instructors are lacking support systems that can help them understand opera ...