Maintenance strategies for nuclear power plants (NPPs) must evolve from static, population-based reliability assessments to dynamic, condition-informed frameworks. By leveraging real-time monitoring data, these advanced strategies can ensure more cost-effective facility operations. By relying on component degradation and failure mechanisms derived from both historical data and real-time health assessments, monitoring techniques reduce the perceived stochasticity of failures, providing more deterministic health estimates paired with quantifiable uncertainty. Furthermore, the functional interdependencies among NPP components require aggregating individualized health assessments to evaluate system-level risk. The margin-based framework addresses this need by calculating geometric distances to failure for both individual components and the overall system. Crucially, instead of forcing all data into failure probabilities, this method accommodates the output of any condition-monitoring technique, mapping diverse indicators directly into a distance to failure. This work intends to advance this methodology to make it practically applicable for complex NPP subsystems and systems. To achieve this, we extend the framework to incorporate the propagation of sensor and model uncertainties, along with sensitivity measures. We then introduce a structured maintenance prioritization strategy utilizing a novel metric, the risk reduction measure. Demonstrated on an electrical power distribution system and contrasted with online fault tree analysis, the results show that the extended margin approach effectively optimizes maintenance schedules, explicitly emphasizes system redundancy, and seamlessly integrates heterogeneous health data under uncertainty.
This paper targets the reduction of fixed operations and maintenance (O&M) costs for advanced reactor (AR) designs. The ability to achieve such reduction will be severely constrained if the assumptions that underly O&M approaches and practices are not questioned and reexamined, especially in the present day, when changes can be implemented both effectively and efficiently. Reducing AR O&M costs necessitates a paradigm shift that is unrealizable through incremental, technology-focused approaches alone.The present work addresses this challenge by evaluating the impact of moving to shorter design lives for major structures, systems, and components (SSCs), as well as to shorter, more predictable refurbishment cycles, as modeled by the commercial airline industry. This is termed the build-to-replace approach. The present paper briefly overviews this new perspective on AR O&M, and provides a set of multi-objective optimization-based analytical tools to identify the benefits of this type of approach. As a direct example, we analyze a specific build-to-replace scenario by identifying and evaluating cases involving reduced SSC lifetimes and associated replacement/refurbishment schedules, thus enabling an evaluation of the impacts on O&M costs and other lifecycle elements, such as SSC reliability.
Existing Human Reliability Analysis (HRA) methods, commonly applied for modeling Main Control Room (MCR) operator performance at Nuclear Power Plants (NPPs), are not designed to address the influence of spatiotemporal evolution of environmental conditions on human performance in External Control Room (Ex-CR) scenarios due to the following challenges: (i) limited empirical human performance data in Ex-CR scenarios, (ii) difficulty in obtaining a complete set of Performance Shaping Factors (PSFs) applicable and important for Ex-CR scenarios, (iii) limited ability to adequately address spatiotemporal and bi-directional interactions between human performance, system response, and hazard progression and (iv) difficulty in handling the large uncertainty associated with plant conditions during the time window of Ex-CR scenarios. This paper develops a novel HRA methodology, namely the Integrated Human Reliability Analysis (I-HRA) methodology, to overcome these challenges. Compared to existing HRA methods, I-HRA possesses a unique combination of four key features: (i) it integrates simulation-based human performance modeling and existing non-simulation-based HRA methods under a unified agent-based modeling platform; (ii) it is equipped with a coupling of physics and human performance simulation models to explicitly capture the underlying, bidirectional, and spatiotemporal interactions among these elements; (iii) it enables explicit simulation-based treatment of dependencies; and (iv) it allows for adequate consideration of aleatory and epistemic uncertainties. The I-HRA methodology is demonstrated using a case study that involves deploying Diverse and Flexible Mitigation Strategies (FLEX) equipment in responding to an external flood at an NPP.
Current system reliability methods (typically based on fault trees [FTs] or reliability block diagrams) can effectively propagate reliability data from the asset level to the system level in order to identify system-critical points. However, the asset reliability data employed are an approximated integral representation of past industry-wide operational experience, and thus neglect an asset's present health status (obtainable, for example, from online monitoring data and diagnostic assessments) and forecasted health projection (when available from prognostic models). Asset health should be informed solely by that specific asset's current and historical performance data and should not be an approximated integral representation of past industry-wide operational experience (as currently performed by system reliability models through Bayesian updating processes). Sensor data, diagnostic assessments, and prognostic assessments are in fact not considered in plant reliability models used to inform system engineers as to which assets are the most critical. In addition, propagation of quantitative health data from the asset level to the system level is made challenging by the diverse nature and structure of health data elements (e.g., vibration spectra, temperature readings, and expected failure time). Ideally, in a predictive maintenance context, system reliability models would support decision making by propagating available health information from the asset level to the system level to provide a quantitative snapshot of system health and identify the most critical assets. This paper directly addresses these two goals by proposing a different approach to reliability modeling—one that relies on asset diagnostic/prognostic assessments and monitoring data to measure asset health. Propagation of health data from the asset level to the system level is performed through FT models, not in terms of probability but rather in terms of margin, with margin being the "distance" between the asset's present status and an undesired event (e.g., failure or unacceptable performance). Per a cause-effect lens, while classical reliability models target the effect associated with asset performance, a margin-based approach focuses on the cause of the undesired asset performance (i.e., its health). Hence, thinking of reliability in terms of margins implies decision-making processes based on causal reasoning. We show how FT models can be solved using a margin-based reliability mindset, and how this process can effectively assist system engineers in identifying which assets are the most critical to system performance.
While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.
As the horizon of nuclear energy expands with the advent of small modular reactors, Generation IV reactors, and fusion reactors, there is a growing perspective that the licensing process could benefit from a more comprehensive approach. Moving beyond traditional deterministic and probabilistic risk assessment analyses might pave the way for a novel safety analysis paradigm propelled by the increasing computational power at our disposal. This paper explores different methodologies that can improve the outcomes of nuclear safety analysis. These range from uncertainty quantification techniques, aimed at enhancing the precision of safety margins, to deploying dynamic event trees by driving system code simulations, capturing the potential evolutions of severe accidents. These methodologies offer a better understanding of the management and consequences of nuclear accident scenarios, significantly improving the accuracy and efficiency of safety predictions compared to traditional methods. Specific case studies illustrate the practical application of these advanced techniques, demonstrating substantial improvements in predicting and managing the dynamics of severe accidents. These findings underscore the effectiveness of these methodologies in enhancing risk assessment capabilities and informing decision-making processes for nuclear safety management. The paper also emphasizes the importance of adaptability and continuous evolution, a call for action to address emerging nuclear safety concerns and highlights the utility of advanced tools like RAVEN.
Operating nuclear power plants (NPPs) generate and collect large amounts of equipment reliability (ER) element data that contain information about the status of components, assets, and systems. Some of this information is in textual form where the occurrence of abnormal events or maintenance activities are described. Analyses of NPP textual data via natural language processing (NLP) methods have expanded in the last decade, and only recently the true potential of such analyses has emerged. So far, applications of NLP methods have been mostly limited to classification and prediction in order to identify the nature of the given textual element (e.g., safety or non-safety relevant). In this paper, we target a more complex problem: the automatic generation of knowledge based on a textual element in order to assist system engineers in assessing an asset’s historical health performance. The goal is to assist system engineers in the identification of anomalous behaviors, cause–effect relations between events, and their potential consequences, and to support decision-making such as the planning and scheduling of maintenance activities. “Knowledge extraction” is a very broad concept whose definition may vary depending on the application context. In our particular context, it refers to the process of examining an ER textual element to identify the systems or assets it mentions and the type of event it describes (e.g., component failure or maintenance activity). In addition, we wish to identify details such as measured quantities and temporal or cause–effect relations between events. This paper describes how ER textual data elements are first preprocessed to handle typos, acronyms, and abbreviations, then machine learning (ML) and rule-based algorithms are employed to identify physical entities (e.g., systems, assets, and components) and specific phenomena (e.g., failure or degradation). A few applications relevant from an NPP ER point of view are presented as well.
Dynamic PRA methods couple stochastic tools (i.e., sampling methods) with system simulators (e.g., RELAP5-3D) to determine the risk associated to complex systems such as nuclear power plants. Compared to classical PRA methods they can evaluate with higher resolution the safety impacts of timing and sequencing of events on the accident progression without the need to introduce conservative modeling assumptions and success criteria. This paper provides an overview on how the INL developed code RAVEN can be used to perform DPRA. In addition, it is shown how machine-learning and data mining methods can be successfully employed to reduce the required computational resources and create knowledge out of gigabytes of generated data. Some applications of dynamic PRA methods are also presented.
This presentation was given as part of the American Nuclear Society Risk-informed, Performance-based Policy and Procedures Committee (RP3C) Community of Practice (CoP). For more information on the RP3C CoP, please visit https://www.ans.org/standards/rp3c/cop/ . This presentation took place on February 25, 2022. The presentation video recording can be found online at https://youtu.be/tV9fnyclH3Y .