The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and deployment. Aviation, often regarded as the safest form of transportation, relies on numerous safety-critical systems. For future safety-critical AI-based systems, EASA requires a Safety-by-Design approach, which can be achieved by using Safety Nets that combine neural network compression with lookup tables to ensure 100
Neural activation coverage (NAC) is a recently-proposed technique for out-of-distribution detection and generalization. We build upon this promising foundation and extend the method to work as an uncertainty estimation technique for already-trained artificial neural networks in the domain of regression. Our experiments confirm NAC uncertainty scores to be more meaningful than other techniques, e.g. Monte-Carlo Dropout.
While Artificial Intelligence (AI) offers transformative potential for operational performance, its deployment in safety-critical domains such as aviation requires strict adherence to rigorous certification standards. Current EASA guidelines mandate demonstrating complete coverage of the AI/ML constituent’s Operational Design Domain (ODD)—a requirement that demands proof that no critical gaps exist within defined operational boundaries. However, as systems operate within high-dimensional parameter spaces, existing methods struggle to provide the scalability and formal grounding necessary to satisfy the completeness criterion, and no standardized engineering method exists to bridge the gap between abstract ODD definitions and verifiable evidence. This paper addresses that gap with a structured, multi-step method that integrates parameter grouping, a dependency analysis distinguishing deterministic from stochastic relationships, physically grounded constraint definition, and criticality-based discretization into a single ODD coverage verification process. Applied to a case study on AI-based near mid-air collision avoidance with the VerticalCAS system, the method reduces the verification space by roughly 60
Developers of automated driving systems (ADS) must demonstrate safe operation within a clearly specified Operational Design Domain (ODD). In early development, however, empirical knowledge of system boundaries is limited while an ODD is already needed to guide design and safety activities. At the same time, technical stakeholders depend on a sufficiently detailed ODD to account for relevant operational conditions throughout ADS development. Despite the central role of the ODD, existing approaches to define an ODD either remain highly abstract or presuppose detailed system knowledge that is unavailable in the early ADS development stages. As a result, no practical method currently exists for defining a comprehensive and detailed ODD in this development phase. This paper addresses this methodological gap by introducing a practical and systematic method for deriving an initial yet detailed ODD. Our method unifies regulatory standards, real-world data, and stakeholder decisions via iterative refinement and contextualization. Two real-world use cases illustrate its applicability and show how context-specific design choices are reflected in the resulting ODD. By systematically constraining the operational space, the derived ODD supports targeted design decisions and safety assurance.
AI-based applications already revolutionize our everyday lives; however, it remains unclear how to assess their safety given their black-box nature. Certifying them for safety-critical applications is thus still under current research. Therefore, we developed the Safe AI Pipeline (SAIPi) as an example of a safety-by-design approach to the engineering of AI systems. The pipeline contains several tools for data, model, and monitoring design, fulfilling requirements such as data completeness and model accuracy. We use MNIST as an exemplary showcase and proof of concept for developing deep learning models. We completed an entire iteration of the pipeline, which demonstrates how to engineer an AI system in a safe, iterative way. Future work will improve the tool and pipeline and increase the complexity of use cases.
Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its integration into safety-critical applications requires compliance with the aviation sector's stringent safety standards. For AI and Machine Learning (ML)-based systems, the European Union Aviation Safety Agency (EASA) emphasizes the need to demonstrate the representativeness and completeness of the Operational Design Domain (ODD) and the associated data distributions used during development and verification. Despite this requirement, a structured engineering process for defining target distributions and evaluating representativeness within ODDs remains largely unexplored. This work presents a method for representativeness assessment of AI/ML constituent ODDs in the context of aviation safety assurance. Starting from the methodical identification of suitable target distributions, a process flow is proposed that guides developers from ODD definition and parameter distribution modeling to the quantitative assessment and interpretation of coverage results with respect to EASA's learning assurance objectives. As quantitative measures, the chi-squared goodness-of-fit test is examined and found unsuitable for the large data sets arising in this setting, leading to the adoption of the Kullback–Leibler divergence and Cramér's V for the representativeness assessment. The method is demonstrated using the example of AI-based airborne collision avoidance, employing experimental data from previous Horizontal Collision Avoidance System (HCAS) and Vertical Collision Avoidance System (VCAS) simulations. The results illustrate how statistical distribution comparison methods can support the assessment of representativeness for safety-critical AI applications and contribute toward a systematic Safety-by-Design AI engineering process aligned with emerging EASA guidance.
Artificial Intelligence (AI) has been on the rise in many domains, including numerous safety-critical applications. However, for complex systems in the real world, defining the underlying environmental conditions in which the AI-based system must operate – the Operational Design Domain (ODD) – is extremely challenging. This often results in an incomplete description of the ODD, which contrasts with the requirements of many domains for certifying AI-based systems. Traditionally, the ODD is created in the early stages of the development process, drawing on sophisticated expert knowledge and related standards. This paper presents a novel Safety-by-Design method to a posteriori define the ODD from previously collected data using a multi-dimensional kernel-based representation. This approach is validated through both Monte Carlo methods and a real-world aviation use case for a future collision-avoidance system. Moreover, by defining under what conditions two ODDs are similar, the paper shows that the data-driven ODD can produce a dataset similar to the original, hidden ODD. Deriving the novel, Safety-by-Design, deterministic kernel-based affinity representation of ODDs is fully automated via a bounded, order-independent algorithm. Utilizing the proposed ODD representation enables future certification of data-driven, safety-critical AI-based systems.
The European Union Aviation Safety Agency (EASA) is developing guidelines to certify AI-based systems in aviation with learning assurance as a key framework. Central to the learning assurance are the definitions of a Concept of Operations, an Operational Domain, and an AI/ML constituent Operational Design Domain (ODD). However, because no further guidance on these concepts is provided to developers, this work introduces a framework for defining them. For the concepts of the Operational Domain of the overall system and the AI/ML constituent ODD, a tabular definition language is introduced. Furthermore, processes are introduced to define the different necessary artifacts. During the specification process for the AI/ML constituent ODD, existing steps were identified and consolidated, including the identification of domain-specific concepts for the AI/ML constituent. To validate the framework, it was applied to the pyCASX system, which employs neural-network-based compression. For this use case, the framework produced an AI/ML constituent ODD with finer detail than other ODDs defined for the same airborne collision avoidance use case. Thus, the proposed novel framework is an important step toward a holistic approach aligned with EASA's guidelines.
The Operational Design Domain (ODD) defines the conditions under which an Automated Driving System (ADS) can safely operate. Validating the ODD requires linking its taxonomy concepts to data that represent real operational environments. Current Operational Domains (CODs) provide this link by describing environmental conditions at specific times and locations using ODD-relevant taxonomy concepts. However, systematic methods for defining CODs and their data requirements are still lacking. This paper introduces the novel formal concept of Minimal Taxonomy Sets (MTS), which represent the smallest sets of taxonomy concepts required within a COD to assess whether an ADS operates inside its ODD. We propose a structured approach for identifying these MTS from nested ODDs and present a variant of the Recursive Prime Implicant Enumeration algorithm that enables their extraction regardless of ODD structure or complexity. We illustrate the effectiveness of the approach using ODDs of varying complexity. Our work shows that MTS provide a principled foundation for generating ODD-aligned CODs, supporting scalable, data-driven validation of ADS.
A steep aircraft increase is forecasted in the near future, putting additional strain on en-route air traffic control. To meet the safety and efficiency goals, contemporary research explores the Single Controller Operations (SCOs) concept to replace traditional positioning of two Air Traffic Controllers (ATCOs) per sector. During workshops with ATCOs addressing SCOs, Conflict Resolution (CR) has been identified as one important task that can be supported by automation. Although existing work on CR solvers shows promising results, solvers based on fixed optimization functions are incompatible with the dynamic evolving preferences of ATCOs. This work proposes two additional steps that filter and rank CR solutions based on a set of rules in natural language-forming a flexible policy-to better align with ATCO preferences in automation-supported CR. Inspired from related work on LLM-driven agents, an algorithm using LLMs to filter and sort CR solutions for alignment with natural language policies is presented. The algorithm is tested on a synthetic dataset of policies and solutions for several minimal filtering and sorting scenarios. The experiments show success in solving the task in most cases and a correct understanding of the task by the LLM. Nevertheless, the analysis of failure cases highlights several limitations of LLMs that must be considered in future research and development of similar systems.
Reliable pedestrian detection represents a crucial step towards automated driving systems. However, the current performance benchmarks exhibit weaknesses. The currently applied metrics for various subsets of a validation dataset prohibit a realistic performance evaluation of a DNN for pedestrian detection. As image segmentation supplies fine-grained information about a street scene, it can serve as a starting point to automatically distinguish between different types of errors during the evaluation of a pedestrian detector. In this work, eight different error categories for pedestrian detection are proposed and new metrics are proposed for performance comparison along these error categories. We use the new metrics to compare various backbones for a simplified version of the APD, and show a more fine-grained and robust way to compare models with each other especially in terms of safety-critical performance. We achieve SOTA on CityPersons-reasonable (without extra training data) by using a rather simple architecture.
A continuous increase in artificial intelligence (AI)-based functions can be expected for future aviation systems, posing significant challenges to traditional development processes. Established systems engineering frameworks, such as the V-model, are not adequately addressing the novel challenges associated with AI-based systems. Consequently, the European Union Aviation Safety Agency (EASA) introduced the W-shaped process, an advancement of the V-model, to set a regulatory framework for the novel challenges of AI Engineering. In contrast, the agile Development Operations (DevOps) approach, widely adopted in software development, promotes a never-ending iterative development process. This article proposes a novel concept that integrates aspects of DevOps into the W-shaped process to create an AI Engineering framework suitable for aviation-specific applications. Furthermore, it builds upon proven ideas and methods using AI Engineering efforts from other domains. The proposed extension of the W-shaped process, compatible with ongoing standardizations from the G34/WG-114 Standardization Working Group, a joint effort between EUROCAE and SAE, addresses the need for a rigorous development process for AI-based systems while acknowledging its limitations and potential for future advancements. The proposed framework allows for a re-evaluation of the AI/ML constituent based on operational information, enabling improvements of the system’s capabilities with each iteration.
We benchmark 90 chunker-model configurations across seven arXiv domains (2 520 retrieval runs) and show that a sentence-based splitter with a 512-token window and 200token overlap reaches the highest token-level Intersection-overUnion (IoU similar to 0.099) while remaining compute-efficient. Our study systematically pairs seven open-source embedding models with semantic and fixed-size chunking strategies, measuring their impact on retrieval quality and latency in RetrievalAugmented Generation (RAG) pipelines. Results reveal that (i) sentence splitting consistently outperforms alternative heuristics, (ii) smaller embeddings deliver more stable cross-domain performance than larger ones, and (iii) finance texts benefit most, whereas astrophysics lags. The accompanying code provides practitioners with empirically grounded guidelines for selecting chunking-embedding combinations that balance accuracy and efficiency.
This project focuses on the detection of small rubber inflatables which are regularly used by migrants to cross the central Mediterranean Sea. The physical attributes of such small targets without materials of high dielectricity decrease the probability of detection. We use multi-platform SAR data and a variety of vessel detection algorithms to gain a better understanding of the backscattering properties and to increase detection capabilities of such inflatables. Our special experimental setup benchmarks detectors at different sea states and wave heights. We adapted well known detectors and implemented detector fusion to test and enhance automatic detection.
Advances in Artificial Intelligence (AI) introduce both promising real-world applications in various fields such as aviation but also challenges in assuring safety. Regulators in aviation mandate, as with any other application, compliance of these AI-based systems with high safety standards. This especially holds for Human-AI Teaming. The European Union Aviation Safety Agency (EASA) outlined in their concept paper the definitions of cooperation and collaboration between humans and AI, as well as the differences in terms of authority and task allocation. Therein, the question of safety in situations of collaboration with shared goals, dynamic task allocation, and partial authority within the Human-AI team is paramount. As a safety concept, EASA requires the usage of Operational Design Domains (ODDs) from the automotive field. In this work, the concept of ODDs is transferred to the Air Traffic Control (ATC) domain. An initial ODD is defined for an AI-based digital team partner that supports Air Traffic Controllers in their daily work even for safety-critical tasks such as conflict detection and resolution. Additionally, the required tools for successfully executing this task is demonstrated. Based on the ODD description, conflict scenarios are generated and tested in a simulation environment, showcasing situations inside and outside the ODD. Based on these results, the feasibility of using ODDs in ATC is discussed, outlining a potential step towards the safe application of Human-AI Teaming.
Applications based on artificial intelligence (AI) promise benefits, ranging from improved performance to increased capabilities in many industries. In the aviation domain, one example is the new Airborne Collision Avoidance System (ACAS X). The current investigation aims at combining ACAS X and AI to maintain its performance while decreasing the memory footprint. However, the anticipation of AI being increasingly used confronts regulators with challenges in terms of safety assurance and certification. Consequently, the European Union Aviation Safety Agency (EASA) published a concept paper for machine learning applications in aviation. Both, the Concept of Operation (ConOps) in combination with an Operational Design Domain (ODD), are listed as objectives to be met for the safety analysis. From a developer’s perspective, this raises questions on how to effectively derive the ODD from ConOps and test the given system based on the ODD description. Based on an exemplary use case of a Near Mid-Air Collision avoidance between two aircraft through the advisories of ACAS X, a highly automated framework for generating and testing synthetic data is proposed. Using this framework, 1800 Near Mid-Air Collision scenario files are created and automatically executed in the simulation environment FlightGear. Scenario-based testing is used for the logging of ACAS X advisory data and evaluating it against predefined requirements. By this approach, an efficient way of verifying system requirements and conducting automated testing based on the ODD definition is demonstrated. Throughout this process, Model-Based Systems Engineering (MBSE) is used to reduce and manage complexity. The framework in this paper enables a systematic and highly automated approach for scenario generation based on the ODD.