
Abstract Railway track geometry is monitored by comparing measurements against regulatory thresholds, but this approach does not capture how the distribution of geometry parameters evolves over time or across track segments. Optimal transport (OT) provides a natural framework for such distributional comparisons. This study investigates whether quantum optimal transport (QOT) can provide additional diagnostic information. Track geometry measurements are encoded as quantum states via angle encoding on three qubits, and transport distances are computed through semidefinite programming under four quantum cost functions alongside classical OT, maximum mean discrepancy, energy distance, principal component analysis, and Mahalanobis distance. The framework is applied to three rail datasets. In a temporal analysis of monthly recordings, the quantum swap cost nearly matches classical squared Euclidean (r=0.998), while the Hamming cost ranks transitions differently (Spearman ρ=0.45). Comparison against a monthly track quality index reveals that Hamming QOT is the only metric with a statistically significant correlation after Bonferroni correction (Pearson r=0.701, p<0.001). In a cross-sectional comparison of healthy and degraded track sections, the eigenvalue spectra of density matrices reveal that degraded sections produce higher-rank quantum states (rank 3–4 versus rank 2 for healthy). Cross-validation on a second class 3 track replicates this pattern. Applied as a per-section anomaly score, Hamming QOT achieves an area under the curve of 0.756 for classifying section health. Sensitivity analyses confirm that these findings are largely preserved across parameter choices. These results suggest QOT offers a distinct but complementary diagnostic lens for infrastructure condition assessment.
As a core component in steel production, the steel strip is prone to surface defects during the continuous production process due to its complex conditions. To achieve high-precision real-time monitoring of its surface quality and efficient quality control, the steel strip surface defect detection method is proposed based on uncertainty region fusion and digital twin (DT) model establishment. Specifically, the You Only Look Once version 8 (YOLOv8) and Single Shot MultiBox Detector (SSD) models pretrained on public defect datasets are used for independent surface detection, obtaining initial defect regions with confidence scores and uncertain detection results. To address single-model detection uncertainty and enhance detection accuracy, the decision-level uncertainty region fusion method is designed to fuse initial regions and uncertainty information based on region overlap and confidence, acquiring maximum reliable defect regions for surface defect recognition and localization. Different from existing state-of-the-art (SOTA) steel strip defect detection methods, the proposed method innovatively introduces uncertainty perception and diagonal equal-area decision-level fusion. It effectively reduces unreliable bounding box predictions and single-model detection bias and further realizes high-confidence defect data interaction and virtual-real synchronization with digital twin framework. To evaluate the performance of the proposed method, the public datasets and experimental steel strip defect datasets are utilized with several defects. The results demonstrate that the proposed method yields consistent and significant improvements in detection accuracy, recall, and overlap rate on both the public Northeastern University (NEU) dataset and the self-built dataset, which verifies the value for steel strip quality control and intelligent manufacturing.
Modern industrial systems are characterized as complex socio-technical systems. The behavior of such systems is emergent, making safety management challenging. Traditional risk assessment methodologies have proven insufficient due to their limitations in addressing systemic complexities. The behavior of complex socio-technical systems is more effectively explained through a functional perspective that places performance variabilities at its center. Rather than viewing variabilities as system deficiencies to be eliminated, this approach recognizes them as intrinsic characteristics that require sophisticated management. The present study leverages the functional resonance analysis method (FRAM) with Monte Carlo simulation and a hierarchical fuzzy inference tree (HFIT) to identify critical functions. The study employed a fuzzy similarity aggregate method and multi-objective optimization to systematically select interventions for managing performance variabilities while respecting cost constraints. Moreover, PROMETHEE II (Preference Ranking Organization Method for Enrichment Evaluations) was used to rank interventions based on effectiveness, cost, and functional significance. The methodology offers flexibility for the decision maker in intervention selection, enabling organizations to balance effectiveness with implementation costs. The approach prioritizes interventions that are highly effective, low-cost, and correspond to functions with significant variability potential. The developed methodology was applied to a case study of the anode-changing operation in an aluminum smelter. The study revealed that interventions such as developing a digital pre-operational checklist, a tool inventory management system, and visual aids for pot-tending machine (PTM) operators for lowering the shovel in the pot cavity should be prioritized. The developed framework could serve as a versatile tool for analyzing and managing functional variabilities across various domains.
In high-risk sectors such as oil and gas, time-dependent corrosion can significantly undermine system integrity, leading to unexpected failures and severe accidents. This paper examines why imprecise analysis is used strategically in risk and reliability assessment, the practical barriers to its adoption, and how integrating confidence intervals can overcome these barriers to enhance results. In response, we propose an imprecise approach based on a new unavailability formula to improve conventional fault tree analysis (CFTA) and better account for dependencies among uncertain parameters. In a representative case study, we first apply CFTA with precise inputs, then use interval arithmetic-based fault tree analysis (FTA) with imprecise inputs implemented in matlab through adaptive interval parameterization to quantify uncertainty and reduce over-conservatism caused by neglected variable dependencies. We also develop an affine arithmetic-based FTA in matlab, which preserves first-order dependencies through shared noise symbols. The proposed approach narrows the system unavailability interval by 22.7% compared with interval arithmetic (IA)-based FTA. A subsequent robustness analysis, performed on a different model structure and parameter configuration, demonstrates an even greater reduction of 28.7%, highlighting the method's effectiveness and adaptability. These results confirm that affine arithmetic is an effective method for managing uncertainty in risk and reliability assessment of real engineering systems, as it reduces overestimation, retains computational efficiency, and facilitates decision-maker updates to safe and reliable control actions.
Human reliability analysis (HRA) characterizes the causes and probabilities of human failure events (HFEs) in complex engineering systems. This work seeks to advance the state of the art in HRA by introducing models with: a causally robust theoretical basis, a user-friendly and data-driven quantification scheme, and applicability to a wide range of scenarios. We present a method for developing HRA Bayesian network (BN) model structures applicable to both control room and ex-control room scenarios, and a further method for model parameterization using literature and data. These structures build upon the Information-Decision-Action in Crew Context (IDAC) framework, which considers information-gathering, decision-making, and action execution to be the three main functions of operators in complex systems. We present the development of three causal models of information, decision, and action HFEs: the Causal Human Reliability Operator-Centered Models (CHROMO). The resulting BNs model the HFEs' causal pathways and are capable of quantifying the effects of causal factors on failure probabilities. The BN structures, built with comprehensive data, have been validated through expert discussions, an ATHEANA-based case study on the human failure events at Three Mile Island, empirical data from German nuclear power plant operational experience, and control room scenarios in the U.S. HRA Empirical Study. These models are currently being incorporated into the Phoenix HRA method and are suitable for regulatory use.
Abstract Current fault diagnosis methods for aircraft landing gear retraction mechanisms at the mechanical system level primarily rely on static logical analyses, such as fault tree analysis, which are difficult to effectively integrate with data-driven diagnostic models. This limitation leads to a disconnect between preliminary analytical results and subsequent diagnostic model applications. To address this issue, this paper proposes a fault diagnosis model that integrates Takagi–Sugeno (T–S) fault tree knowledge with an attention mechanism. The proposed model adopts a physically interpretable prior-knowledge integration strategy, thereby establishing a bridge between fault tree analysis and diagnostic modeling. Specifically, a data-and-mechanism-driven attention module is designed to map the importance analysis results of the T–S fault tree into a mechanism-informed weight vector in the neural network. Through adaptive weighting of sensor signal features, the model is guided to focus on high-risk fault features. In addition, a fault dataset covering typical failure modes of the landing gear retraction mechanism is constructed based on multibody dynamics simulations for validation. Experimental results show that the proposed model effectively incorporates prior mechanism knowledge and achieves higher diagnostic accuracy than conventional deep learning models for typical mechanical faults, including joint wear, insufficient joint lubrication, and spring fatigue, thereby demonstrating its feasibility and effectiveness.
Remaining useful life (RUL) prediction is crucial for the predictive maintenance and health management of mechanical transmission systems. Although data-driven RUL prediction methods have advanced significantly, three key challenges remain: (i) constructing a physically interpretable HI that is invariant to operating conditions yet sensitive to incipient degradation; (ii) achieving adaptive time to start prediction (TSP) detection that accommodates non-stationarity and controls false and missed alarms; (iii) accurate and generalizable RUL prediction under cross-domain scenarios. To address these issues, a method combining unsupervised anomaly detection and deep feature transfer for TSP determination and RUL prediction is proposed. A residual health indicator is constructed from an expected trajectory learned by Gaussian process regression (GPR) on the healthy segment and a Mahalanobis distance (MD) deviation measure. For the first time extreme value theory is used to create dynamic thresholds for performance degradation analysis and TSP determination. During the RUL prediction phase, transfer component analysis mitigates feature distribution discrepancies between source and target domains. A dual long short-term memory (dual-LSTM) network with L2 regularization is then employed for RUL prediction across two transfer tasks. Experimental validation on a public benchmark dataset and a laboratory run-to-failure bearing dataset demonstrates the superior predictive accuracy and generalization capability of the proposed approach compared to existing methods.
Today, terms such as sustainable production, industrial cyber-physical systems, cyber-physical production systems (CPPS), software-defined manufacturing, smart manufacturing, Industry 4.0, Industry 5.0, system of systems, internet of things, human-in-the-loop, and digital twins are widely used. These concepts emphasize key characteristics of modern and future production systems, including heterogeneity, structural and behavioral complexity, intelligence, autonomy, reconfigurability, and human centrism. They also highlight the growing importance of reliable and up-to-date risk assessment, safety, and reliability measures, given the significant environmental, economic, and social demands. However, current industrial risk analysis methods lag behind the rising technical sophistication of such systems. It remains unclear whether existing methods can capture complex failure scenarios of dynamic, AI-driven systems with advanced software architectures. This paper presents the main challenges faced by safety engineers in industrial automation and provides a structured classification of risk and reliability analysis methods and metrics, supported by a systematic review of 95 papers up to October 2025. The review addresses questions such as: which CPPS aspects must be considered, which methods are applicable, what are their advantages and limitations, and how can methods be combined? The findings reveal the need to extend classical approaches toward dynamic risk assessment, probabilistic model checking, AI-based techniques, digital twins, and intelligent fault injection. The study provides both a comprehensive overview of current risk and reliability assessment methods for CPPS and a roadmap for advancing their future development.
The rapid development of large language models (LLMs) has introduced a new paradigm for decision-making in mechanical system health management. However, due to the critical lack of interpretable data embeddings and uncertainty-aware, the direct application of LLMs in industrial settings remains ineffective. Therefore, the key to successful LLMs implementation lies in analyzing uncertainty in data embeddings to enhance reliability. This paper proposes a novel model called Probabilistic Hierarchical Perception Networks (PHPN), which first leverages convolutional kernels to extract local features characterizing vibration semantics and then employs a Transformer backbone to capture their global semantic relationships. Furthermore, a probabilistic Bayesian approach is incorporated to enhance the model's perception of prediction uncertainty. In addition, a complete interpretable fault diagnosis framework is proposed to enable the model to automatically quantify and interpret sources of uncertainty when dealing with fault diagnosis tasks. By effectively separating aleatoric and epistemic uncertainty, the framework can identify the sources of unknown factors in the data, thereby improving the accuracy and confidence of the diagnosis. The experimental part is tested on an independently constructed structural fault dataset, which verifies the effectiveness and generalizability of the method proposed in this paper. Comparison results with existing methods show that PHPN can provide more comprehensive diagnostic results and better uncertainty separation performance. Meanwhile, the proposed method can also provide uncertainty analysis and risk assessment for LLMs in industrial health management.
Abstract Given the multisource and uncertainty of tower crane construction safety risks, this study proposes a comprehensive risk analysis method that integrates disaster chain theory (DCT), analytic hierarchy process (AHP), entropy weight method (EWM), and cloud model. Firstly, based on the disaster chain theory, the topology of risk factors is established, and the key risk factors, such as equipment failure, operation error, environmental interference, and management defects in the operation of tower cranes, are systematically identified. An evaluation system consisting of four first-level indicators and 16 s-level indicators was constructed through the topological mapping of the disaster chain, and the analytic hierarchy process-entropy weight method combined weighting method was used to integrate the subjective and objective weights to effectively balance the expert experience and data objectivity. The cloud model theory is introduced to deal with the ambiguity and randomness in risk assessment, and the mapping relationship between the risk level and the cloud characteristic parameters are established to achieve the conversion from qualitative concept to quantitative evaluation. The case application shows that the risk factors with higher weight ranking include: fatigue/deformation of steel structure, and failure of safety devices. This study innovatively combines dynamic weight optimization with uncertainty modeling to provide a visual risk assessment method for tower crane construction companies.
Regarding the problem that the weak fault features in the vibration signals of aircraft engine gearboxes are difficult to stably extract under strong noise interference and nonstationary conditions, a multistage signal processing method oriented toward enhancing weak impact features is proposed. This method builds a full-band multiscale decomposition framework based on the maximum overlap wavelet packet transform and combines the Choi-Williams distribution (CWD) to achieve time-frequency energy reconstruction, in order to improve the focusing ability of fault features in the time-frequency domain. On this basis, the complete ensemble empirical mode decomposition with adaptive noise (CEEMDAN) decomposition, permutation entropy (PE) filtering, and adaptive wavelet threshold denoising are introduced to construct a multilevel denoising strategy. Through the collaborative mechanism of "modal decoupling-information screening-directional denoising," the stable retention and progressive enhancement of weak fault information in the multistage processing process are achieved. Experimental results show that the proposed method can effectively improve the structural distinguishability and feature stability of vibration signals in complex noise backgrounds, enable clearer expression of periodic impact features and fault harmonic structures, and achieve a good balance between noise reduction performance and signal fidelity. The research results indicate that the proposed method enhances the detectability of early weak faults in aircraft engine gearboxes, providing a robust signal processing path and engineering application potential for high-reliability condition monitoring and intelligent diagnosis under complex conditions.
In the context of industry 4.0 and industry 5.0, digital twins (DTs) are key and useful elements for monitoring and adjusting physical entities in manufacturing systems. Although this approach offers clear and well-demonstrated advantages, DTs are affected by inaccuracies and uncertainties coming from multiple sources, such as data sensing, biases introduced by data processing, and model abstraction. Adequately managing uncertainty in DT is crucial for the efficient operations of manufacturing systems. This paper proposes a hybrid framework (i.e., a human-in-the-loop supported by AI techniques) for uncertainty management. The proposed framework uses AI approaches to identify and quantify uncertainties and to propose actions to the human user to reduce the risk of unsafe and/or inaccurate reasoning in DT information processing. The proposed approach incorporates a continuously running module that monitors the DT throughout its life cycle to identify emerging sources of uncertainty. This module is a mechanism for tracking how uncertainties propagate through the model, and for registering how these uncertainties influence and modify the resulting outcomes. Also, the proposed hybrid approach helps identify the DTs risky components and enables human users to manually intervene, facilitating informed and transparent decisions. A simplified case study of predictive maintenance for machine resources in the production line of a lumber factory producing medium-density fiberboard is used to illustrate and validate the proposal. The objective in this case study is to assess whether the proposed framework is able: (i) to identify uncertainty values out of the given confidence threshold defined by the user, (ii) to support the refinement and adaptation of the training phase of the AI approach implemented by the DT to reduce uncertainty, and (iii) to adequately detect malfunctioning sensors requiring replacement.
This paper presents an interpretable framework for constructing health indicators (HIs) to support reliable monitoring and prognostics in complex multicomponent systems. Modern industrial environments generate large volumes of heterogeneous sensor measurements, and extracting meaningful and interpretable indicators from these signals remains a critical step in system health assessment. While data-driven methods have achieved promising predictive performance, they often operate as black-box models, which limits transparency, trust, and practical adoption in safety-critical industrial applications. Explainable artificial intelligence has emerged as a key solution to address this black-box limitation by providing transparent and understandable model behaviors. In this context, this work develops an interpretable-by-design health indicator construction framework that offers clear insight into sensor relevance, component degradation behavior, and system-level health evolution. To this end, the proposed framework introduces three key innovations: (i) a context-aware sensor selection mechanism that identifies the most informative sensing channels under different operating conditions, enhancing robustness to noise and variability; (ii) a flexible module that incorporates expert knowledge to guide the shape and behavior of component-level health indicators when domain insights are available; and (iii) an adaptive aggregation strategy that automatically selects appropriate functions to synthesize component-level health indicators into a system-level representation consistent with the system's structural and operational characteristics. The framework is validated on the Tennessee Eastman Process (tep) and the Commercial Modular Aero-Propulsion System Simulation (c-mapss) turbofan datasets, both featuring high-dimensional, multisensor measurement data with interacting components. Experimental results show that the method produces monotonic, trend-consistent health indicators that capture component and system degradation more accurately than existing approaches, offering improved interpretability, adaptability, and robustness for measurement-informed monitoring in complex industrial settings.
Typically, in engineering projects, aleatory variability of materials, phenomena, and geometry/hardware can only be sampled very sparsely by replicate physical experiments or tests. This precludes accurate identification of the form and parameters of frequency distributions, or more generally, multidimensional random function models, of stochastically varying quantities for ensuing propagation through physics models. These difficult conditions are well handled with a new class of discrete-sample propagation and uncertainty processing approaches described and demonstrated in this paper, instead of trying to infer the random function models from the sparse sample data and then propagate the inferred functions and inference uncertainties. We explain the "Simultaneous Discrete Direct" (SDD) approach for aleatory uncertainty representation and propagation of multiple sources of sparsely sampled aleatory variability and apply SDD to a realistic and challenging test problem involving computationally expensive weld and structural response models. We perform random draws of N = 4 samples of each of five sources of known variability in the synthetic-reality test problem, propagate all samples with N = 4 model runs per the economical SDD approach, and use specialized 1D uncertainty quantification techniques to statistically process the N = 4 samples of each of 40 response quantities of the structural model into reliably conservative bounding estimates on 0.005 tail quantiles and probabilities. We perform 250 random trials to quantify the usefully high reliability/confidence of attaining conservative estimates. SDD is also simpler and less computationally expensive than frequency-distribution inference and propagation methods, so in general appears well suited for sparse-data conditions typical in engineering applications.
For addressing the time-dependent failure possibility (TDFP) under fuzzy uncertainty, the Kriging surrogate models combined with fuzzy simulation (FS) have achieved promising results. However, existing learning functions for Kriging do not comprehensively account for the predictive sign misclassification probability and the contribution of newly selected samples for the failure possibility. To further enhance the computational efficiency, this study proposes a new adaptive Kriging (AK)-based method for solving TDFP. This method first estimates the misclassification probability for critical performance signs via multivariate Gaussian distributions and then selects samples based on both their failure possibility contribution and misclassification probability. This strategy directs new samples toward critical failure regions with maximum joint membership functions (MFs), accelerating Kriging convergence. Three test examples and an engineering application of a simplified turbine blade validate the advantages of the proposed method, which can reduce the number of performance function calls while preserving accurate estimates of TDFP.
In today's modern world, the information processing server is the most demanding and challenging area of research. The theory and methods of information processing servers have developed considerably in recent decades, as demonstrated by several publications. Modern information and communication technology is the result of many years of technological development and is now the driving force behind contemporary operations and solutions. In this study, the structure-function approach is used to calculate the reliability of an information processing server system. The series-parallel complex model has been considered for analyzing the sensitivity of the proposed server system. The researchers attempted to offer the major aspects of reliability and the values of reliability of an information processing server system have been calculated. The B-P index and sensitivity determination are the evaluating components of the suggested system. The future work and some applications are also discussed here.
This study presents the development of a comprehensive thickness prediction framework to enhance risk-based inspection (RBI) for pressure vessels. The research focuses on recommending essential maintenance strategies to maintain structural integrity and prevent failures using a systematic reliability assessment approach. Five predictive models were employed: Bayesian, exponential, linear, power, and quadratic, validated using historical thickness data collected between 2002 and 2008, with projections extended to 2040. The results demonstrated that the power model achieved the highest predictive performance with a coefficient of determination (R2) of 0.9394, while the exponential model showed strong reliability with an R2 of 0.8982 and the lowest mean absolute error (MAE) of 0.1117 +/- 0.0473 mm. The Bayesian model excelled in uncertainty quantification, achieving an R2 of 0.8500 and an MAE of 0.1141 +/- 0.0480 mm. Safety and risk assessments conducted for critical vessel sections (E head 1, nozzle A1, nozzle N4, and shell 1) identified nozzle N4 as requiring immediate attention, with power and quadratic models predicting failure conditions by 2040. Shell 1 demonstrated excellent safety margins, allowing for extended inspection intervals, while nozzle A1 showed lower safety margins, requiring more frequent monitoring. The integration of probabilistic and deterministic predictions into probability of failure (PoF) calculations significantly improved inspection interval optimization and enhanced risk management effectiveness. This framework enables proactive integrity management strategies, potentially achieving cost savings of 20-50% through efficient resource utilization while establishing a robust foundation for improved safety and regulatory compliance in chemical plant operations.
To assess crashworthiness of vehicles, it is important to consider the inherent variability of physical crash tests in the virtual design, i.e., before the first physical tests, through robustness evaluations. For this, forward uncertainty propagation methods have to be established based on computational models of the vehicle for relevant load cases. In the context of occupant safety, the seating position of an anthropometric test device (ATD) and the seatbelt routing in the initial state are influential uncertain parameters for injury criteria. For this purpose, we propose a two-stage probabilistic modeling approach to incorporate these aleatoric uncertainties, based on test data, into the respective computational models for stochastic crash simulations. The ATD seating position is modeled using linear Gaussian Bayesian networks, and the seatbelt routing is represented by multitask Gaussian process models. This method enables the generation of physically consistent geometric variations that are automatically transferred into adapted finite-element models. The approach is applied in a generic frontal crash scenario with the THOR-50M ATD. The results highlight that variations in seatbelt routing affect chest compressions, while uncertainties in ATD seating position significantly influence lower body loading. The diagonal belt path is identified as the dominant factor for chest injury criteria, whereas the horizontal hip point location and correlated leg posture govern pelvis and femur loads. The findings underline the importance of probabilistic modeling geometric uncertainties in virtual robustness assessments of occupant protection systems.
Wind turbines, which are typically complex repairable mechanical systems, fail not only with respect to their own operating time but also depending on the state covariates in practical engineering. We synthesize the lifetime and condition covariates of the equipment and propose a new method for dynamically determining the condition detection interval, which can effectively reduce the cost of wind turbine operation and maintenance. First, the system failure process is described based on the boundary intensity process and combined with the proportional risk model to establish the reliability model of the system. Second, a mathematical expression for the state deterioration threshold is given, and the state detection interval is dynamically adjusted according to the system state changes with the constraint of fault risk. Finally, the proposed method is validated by a comparative modeling analysis, which provides a new approach for the determination of state detection intervals.
The traditional quality traceability methods for electromagnetic trip devices suffer from three major drawbacks. First, there exist information silos among departments, resulting in low transparency. Second, node redundancy leads to a cumbersome process. Thirdly, the methods have inherent bottlenecks, causing delays in traceability. These drawbacks compromise the accuracy and timeliness of quality traceability for electromagnetic trip devices and expand the scope of harm caused by their quality and safety issues. To address these problems, this paper proposes a quality traceability method based on blockchain and Petri nets. First, the tamper-proof and decentralized characteristics of blockchain are utilized to resolve issues related to data credibility and transparency. Then, the Petri net employs an association matrix reorganization algorithm to divide the blockchain into subnets. If the same place has different transitions across multiple subnets, that place will become a target of competition, thus identifying the place and its corresponding transitions as conflict nodes. Next, a Petri-Markov chain for the blockchain subnets is constructed, and redundant nodes are identified by calculating the place occupancy rate and transition utilization rate. Finally, the conflicting and redundant nodes are optimized, and the new method is compared with the traditional quality traceability method through tina software. In the scenario-based simulation, results show that compared with the traditional method, the new method improves traceability efficiency by 35.93% and reduces node delay time by 45.94%.