
This study presents an individualized inverse Gaussian process-based reliability modeling and optimal degradation test design for rubber V-belts, addressing the high cost, long duration, and data scarcity of traditional reliability demonstration testing (RDT). Degradation experiments on five B-type V-belts monitored slip ratio, tension force, and primer crack depth every 12 h, yielding lifetimes between 190 and 209 h (mean: 197 h). A stochastic degradation model incorporating individual heterogeneity was developed, with parameters estimated via Bayesian inference using Markov Chain Monte Carlo. By minimizing the asymptotic variance of the 0.1-quantile lifetime under cost constraints, an optimal RDT scheme was derived. Results show that when the asymptotic variance is below 0.0035, the relative error between estimated and actual lifetimes remained below 1.52%. The optimal plan—four samples, 18 measurements, and a 10-h interval—achieves high evaluation accuracy at a total cost of 7696 CNY. This work advances existing methods by explicitly modeling individual variability, justifying the inverse Gaussian process choice through empirical validation, and offering a practical, cost-efficient RDT framework extensible to other degradation-prone products.
The construction of health indicators (HIs) for failure prognostics is crucial for monitoring and predicting the health state of multi-component systems. Recently, data-driven methods for HI construction and prognostics have gained increasing attention, both relying heavily on high-quality run-to-failure (RTF) data for accurate analysis and reliable prediction. To foster trust in these systems, integrating explainable AI (XAI) into prognostics and health management (PHM), forming XPHM, is highly recommended. However, current literature reveals several research gaps: a lack of RTF data for multicomponent systems, limited consideration of complex component interactions, and insufficient attention to environmental impacts in constructing explainable HIs for prognostics. To address these challenges, this paper introduces new RTF data for the Tennessee Eastman Process (TEP), capturing multicomponent interactions under various environmental conditions through a systematic simulation methodology. Furthermore, the study analyzes this data, elucidating the direct and indirect relationships between sensor measurements and TEP component states across different working conditions. These findings provide valuable insights for researchers and practitioners in developing both component- and system-level HIs and prognostics. Additionally, we illustrate how to use the provided data and analyzed results to construct HIs, further supporting the development of reliable XAI solutions, thereby improving the reliability of engineering systems.
Degradation models are the basis for reliability analysis of products whose performance deteriorates over time. The choice of degradation model is often based on information criteria such as the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC). However, the direct influence of model misspecification on the accuracy of the reliability assessment is rarely quantified. This paper focuses on models based on the Tweedie exponential dispersion (TED) process, chosen for the excellent versatility it demonstrates, and wish to study the problem of model misspecification for the TED process with random effects (TEDR). First, we derive the asymptotic distribution of the quasi-maximum likelihood estimator (QMLE) of the mean time to failure (MTTF). when the model is mis-specified. Then, we quantitatively analyze the influence of the misspecification by calculating the relative bias and variability. Finally, we apply the theoretical elements developed to the case of wear data of a civil aircraft flap and the GaAs laser data, for verification purposes. The results observed through the case study confirm that the misspecification of the TED process model will have a significant impact on the assessment of the average product lifetime. This impact will be more pronounced in the case of small samples. It also shows that the TEDR model is robust. In this paper, we therefore provide theoretical elements to enrich reliability studies of products whose performance deteriorates over time by associating and quantifying the effects of a misspecification of the degradation model.
In this paper, a deep reinforcement learning approach is used to provide a dynamic maintenance model for a load sharing k -out-of- n : G system with identical components, where all remaining components share the load equally. It is assumed that the degradation path of each component in the system follows an Inverse Gaussian process and the system degradation is the only cause of the system failure. The maintenance problem is expressed as a Markov decision process, and the optimal maintenance action is found by solving the problem using a advantage Actor-Critic algorithm, where the algorithm and its modifications on its input are extensively discussed in this paper. Finally, the performance of this modified deep reinforcement learning approach in finding the optimal policy for a load sharing k -out-of- n : G system is demonstrated through a numerical example and compared with the performance of standard Actor-Critic algorithm. Furthermore, this paper uses a decision tree to depict the optimal policy, providing a more interpretable tool than the neural network.
The probability of observing defectives in a production process or lot inspection is usually assumed to be constant. A methodology to determine optimal sampling plans for Lindley distribution is presented by limiting the expected producer and consumer risks. The proposed approach generalizes the traditional single sampling plans when exists prior information on that is modeled under a prior model. In this case, the beta distribution is used to describe the random variation in the proportion defective. The suggested procedure may provide a significant reduction in sample size, as well as a better evaluation of the real producer and consumer risks. Two examples related to the lot acceptability with real lifetime data are shown for illustrative purposes.
Focusing on intelligent heat supply systems (IHSS), this paper proposes a conditional time-varying fused reliability model that extends the scope of conventional heat supply system (HSS) reliability analysis by incorporating software subsystem dynamics alongside the hardware subsystem. The framework integrates the Weibull distribution to characterize hardware lifetime evolution, the non-homogeneous Poisson process (NHPP) to capture software failure dynamics, and a consumer-oriented reliability model based on the minimum path matrix. Validated through Monte Carlo simulations on a real IHSS case, the proposed models enable differentiated reliability improvement across heating paths, offering actionable support for system design optimization and predictive maintenance.
This work presents a statistically validated and data-driven framework for early fault diagnosis of multistage industrial gearboxes, with particular emphasis on pitting defect detection using multidimensional vibration measurements acquired under varying operating conditions (0–30 Nm loads and 1000–1500 RPM speeds). Noise-Assisted Multivariate Empirical Mode Decomposition (NA-MEMD) is employed for adaptive, non-stationary signal decomposition to obtain a set of Intrinsic Mode Functions (IMFs). To improve decomposition robustness and mitigate mode mixing, the influence of multiple white Gaussian noise (WGN) channel configurations (1–3 channels) is systematically investigated. Subsequently, the primary contribution of this study involves the quantitative estimation of an optimal IMF selection criterion. Five distinct thresholding strategies are comparatively evaluated, and the most effective approach is determined through maximization of the Signal-to-Noise Ratio (SNR), enabling retention of fault-sensitive components while suppressing noise-dominant IMFs. The resulting effective IMFs are further analyzed in both time and frequency domains, where Fast Fourier Transform (FFT) facilitates identification of characteristic fault frequencies and spectral signatures. Statistical descriptors, including root mean square (RMS), peak value, and standard deviation, are extracted to construct discriminative feature vectors. Classification performance is comparatively assessed using Decision Tree (DT), Weighted k-Nearest Neighbors (W-kNN), and Tri-layered Neural Network (TNN) models. Finally, Analysis of Variance (ANOVA) is conducted to determine the optimal noise configuration, thresholding strategy, feature set, and classifier. Experimental results demonstrate enhanced SNR and superior diagnostic accuracy, confirming the effectiveness of the proposed methodology for reliable gearbox condition monitoring and predictive maintenance.
The structural reliability of Y25 bogie suspension systems is governed by complex interactions among geometric discontinuities, material variability, manufacturing imperfections, nonlinear stiffness behaviour and assembly-related deviations. While previous studies have largely focussed on component-level stress and fatigue performance, the mechanisms by which local degradation propagates into system-level safety risks remain insufficiently addressed. This study proposes an integrated, risk-informed reliability and system-safety assessment framework that combines Ishikawa root-cause analysis, Quality Function Deployment (QFD), Design Failure Mode and Effects Analysis (DFMEA), and Bayesian Network (BN)-based probabilistic risk propagation. The Ishikawa analysis identifies 66 interacting contributors to structural degradation, highlighting the coupled influence of design, manufacturing and operational factors. QFD results reveal that spring stiffness characteristics, material fatigue resistance and geometric transition quality dominate safety-relevant system behaviour. DFMEA identifies spring seat-housing cracking, bedbox edge yielding, and fatigue crack initiation in cast-steel regions as the most critical failure modes due to their strong association with stiffness degradation and uneven load transfer. These failure mechanisms are subsequently embedded within a BN framework to quantify uncertainty propagation, identify dominant risk drivers and support time-dependent risk prioritisation across service life phases. By explicitly linking component-level degradation mechanisms to system-level instability and maintenance decision support, the proposed framework provides a structured and lifecycle-oriented methodology for reliability governance in safety-critical railway bogie suspension systems.
We present a model for the corrective and preventive maintenance of a system. The latter is based on a bivariate policy that replaces the system either at age T or after the M failure, whichever comes first. A repair follows each of the M − 1 first failures, restoring the system operational state, but with a lower reliability than before failing. We present two scenarios with constant and time-dependent repair costs. The results reveal that systems with low initial reliability can greatly benefit from the bivariate policy. The advantage decreases for poor quality repairs. We also obtain conditions to obtain the optimum number M when T is given. This result is helpful to assess whether a system should be replaced sooner than originally planned.
To address the critical failure of indoor navigation networks during fire or explosion incidents in complex urban public buildings, this study introduces a Building Information Modeling (BIM)-semantics-enhanced dynamic topology road network (DTRN) framework for three-dimensional emergency pathfinding. Departing from conventional static graphs, the proposed method dynamically reconfigures the constrained Delaunay triangulation (CDT) in response to the spatio-temporal evolution of hazards, enabling real-time, safety-oriented and scalable path optimization. Experiments conducted on a three-storey office building under simulated fire-spread scenarios demonstrate that DTRN achieves a balanced trade-off between accuracy and efficiency: the relative path deviation is confined to 5.2%, the average computational time is 0.125 s, and the path length converges to the ground-truth value after four iterations, with the absolute error decreasing from 5.4 to 1.7 m. Moreover, DTRN supports mission-driven pathfinding that enforces traversal of emergency-equipment nodes; although this increases path length by approximately 3.5%, it significantly improves rescue effectiveness. The proposed framework offers a new paradigm for semantically aware and adaptively responsive emergency pathfinding in complex indoor environments.
This paper presents an advanced reliability analysis of a fault-tolerant cellular base station modeled as a hierarchical (m-out-of-r)-within-(k-out-of-n):G system, incorporating nested redundancy to satisfy the stringent availability demands of modern telecommunication networks. The base station comprises three independent sectors (Alpha, Beta and Gamma) configured in a 2-out-of-3:G arrangement, each containing four critical components: a primary power amplifier (PPA), redundant power amplifier (RPA), primary transceiver unit (PTU) and redundant transceiver unit (RTU), structured as a 3-out-of-4:G system. Component failures within each sector exhibit dependency owing to a shared power supply and environmental factors, which are modeled using a copula function. System behavior is captured using a continuous-time Markov process, with the Sumudu transform applied to derive analytical expressions for subsystem reliability, availability, and state probabilities. The numerical analyses considered both repairable and non-repairable scenarios. The overall system reliability was computed using the universal generating function (UGF) technique, which integrates subsystem reliabilities. A sensitivity analysis was performed to assess the effects of the failure rate, repair rate, copula dependence parameter, and hierarchical structural parameters on system performance. An economic evaluation based on the expected profit assessed the cost effectiveness of the design. The results confirmed that hierarchical redundancy significantly enhances the reliability and profitability of telecommunication infrastructure. The proposed copula-UGF framework effectively models the complex dependencies and nested configurations of cellular networks, offering valuable insights for reliability optimization and strategic decision-making.
The U.S. nuclear industry is increasingly modernizing its operations by integrating automation technologies to improve efficiency, safety, and reliability, while minimizing unnecessary costs. However, successfully deploying automation technologies for operations and maintenance (O&M) in Nuclear Power Plants (NPPs) requires addressing several critical challenges, including ensuring that the automation is trustworthy, transparent, and operationally acceptable. This paper focuses on automation transparency, a characteristic that plays a vital role in enabling safe and efficient human-automation interactions and, thus, contributing effectively to operational risk management. Despite its importance, there is currently no consensus on the definition of automation transparency across various domains, and existing methodologies for measuring transparency are predominantly subjective and context specific. This paper presents a literature review to assess the existing definitions and methodologies of automation transparency and identifies key limitations in their applicability to the nuclear domain. In addressing these limitations, this study proposes a novel definition that refers to automation transparency as “ the degree to which underlying information about the inner workings of an automation system is conveyed, relevant to its intended use. ” Aligning with this proposed definition, three evaluation approaches for automation transparency, namely attribute-based, model-replication-based, and entropy-based, are developed. A hypothetical case study that involves an Artificial Intelligence (AI)-based automated anomaly detection system used in NPP monitoring is leveraged to test these three evaluation approaches to better understand their feasibility, applicability, and limitations. Future work will apply the lessons learned from this study for a practical case study to further understand the impacts of automation transparency on human performance.
This paper proposes a dynamic-switching active-learning Kriging (DS-AK) method considering epistemic uncertainty in modeling and active learning processes to enhance the efficiency and accuracy of structural reliability analysis. To address the epistemic uncertainty associated with modeling and active learning, the method dynamically selects the optimal Kriging model and learning function from pools of candidates based on the maximum relative error of failure probability and the Kriging believer criterion. A hybrid stopping criterion that considers the stability of the failure probability estimation is also developed. The proposed method is validated through three numerical examples and a practical engineering application, demonstrating its superior performance over some existing methods.
In manufacturing, researchers have conducted extensive studies to prevent malfunctions of rotating equipment during operation via various Prognostics and Health Management (PHM) techniques. However, most recent studies require sufficient fault data for training, and they usually focus on only a single failure mode or component. Although they may offer novel academic insights, they often have limited practical field applications, mainly due to a lack of ability to access and cope with data in various operating environments. Therefore, it is necessary to develop a method that can detect multiple failure modes early with minimum or at least reasonable data acquisition. To overcome these challenges, this study presents a practical approach for integrated fault diagnosis of rotating machinery by establishing a multi-step diagnostic framework. The framework adopted physically interpretable frequency-domain features extracted from vibration signals, specifically the shaft rotational frequency and bearing characteristic frequencies (BPFO and BPFI), each considered up to the third harmonic, to jointly diagnose shaft and bearing faults. Health indicators were then computed using Mahalanobis distance with respect to a baseline distribution, enabling reliable fault diagnosis using only normal-condition data. To confirm universal field applications of the suggested method, the performance was validated in various test cases with different scenarios. The results proved that the approach was useful for practical purposes by achieving high accuracy, improving from 88% to near-perfect performance. This approach allows for the detection and diagnosis of anomalies and their specific types in both the shaft and rolling element bearings, relying solely on normal operating data.
For the structure with uncertain distribution parameters of random inputs, the prior and posterior augmented failure probability (AFP) can be used to respectively evaluate the prior reliability level and the posterior one under newly observations. However, for the common engineering problem with small magnitude of the prior and posterior AFP and high-dimensional inputs, there is still a lack of efficient algorithm. Therefore, this paper proposes an augmented importance sampling (AIS) method based on the von Mises-Fisher-Nakagami mixture (vMFNM) model for estimating prior and posterior AFP. In the proposed method, the parameterized vMFNM is used to approach the optimal IS density for estimating AFP, and by minimizing the difference of two densities, the parameters of the vMFNM can be determined and then the AFP can be efficiently estimated. To further improve the efficiency of the proposed algorithm, the strategies of layered constructing vMFNM model and adaptively training surrogate model of performance function are incorporated into the proposed method, which further reduces the performance function evaluations for estimating small AFP. The example results show that the computational accuracy and efficiency of the proposed method are significantly better than those of existing methods, especially for the problem with high-dimensional inputs and small magnitude of the AFP.
The degradation of braking torque in elevator block brakes, resulting from the coupled effects of brake shoe wear and spring stress relaxation, poses a critical challenge to operational safety. In this study, a reliability model is developed to characterize the joint degradation behavior of these two mechanisms. The monotonic and irreversible processes of wear and stress relaxation are modeled using two independent gamma processes, and their combined effect on braking force is formulated based on Hooke's law. The relationship between braking force and braking torque is established through experimental identification. Model parameters are estimated and validated using degradation data obtained from field inspections. The results indicate that the reliability evolution exhibits a two-stage characteristic, consisting of an initial stable phase with high reliability (approximately 45 months), followed by a rapid degradation phase. The reliable life, defined as the duration corresponding to a reliability level exceeding 0.9, is approximately 50 months. Sensitivity analysis demonstrates that the equivalent friction coefficient (as represented in the torque-force relationship) and the initial spring stiffness are the dominant factors affecting the service life. The proposed model provides a quantitative framework for reliability assessment and offers theoretical support for condition-based maintenance and life prediction of elevator block brake systems.
Reliability growth modeling plays an important role in the analysis of repairable systems operating under evolving and non-stationary conditions; however, classical formulations such as Crow-AMSAA and Musa-Okumoto generally rely on single-regime assumptions that may limit their ability to represent heterogeneous failure dynamics and imperfect maintenance effects across successive operational phases. This paper proposes a Multi-Rate Reliability Growth Model with Partial Reset (MR-RGM-PR), a multi-regime NHPP-based probabilistic framework integrating regime-dependent dynamics and intensity-level memory propagation mechanisms to represent residual effects induced by imperfect corrective actions. An estimation framework combining maximum likelihood estimation and Bayesian inference is developed to support parameter identification and uncertainty quantification, while the theoretical analysis demonstrates that the proposed formulation preserves the principal properties required for cumulative counting processes, including continuity, monotonicity, and non-negativity, while enabling flexible representation of regime transitions and inter-regime dependency propagation. Validation experiments conducted on both synthetic and real-world datasets suggest that the proposed framework provides improved goodness-of-fit characteristics and reduced prediction discrepancies relative to several classical reliability growth formulations under the considered experimental conditions, as reflected by lower RMSE and MAE values together with improved information criteria. The results additionally highlight the potential relevance of combining adaptive segmentation and memory propagation mechanisms for representing complex reliability dynamics involving heterogeneous operational conditions and partially effective maintenance actions. From a practical perspective, the proposed MR-RGM-PR framework may support the identification of critical operational phases and provide useful insight for maintenance planning, intervention-policy optimization, and reliability-oriented decision support in industrial environments.
As high-speed electric multiple units (EMU) speeds and mileage increase, the electronic traction system has become a prominent part requiring focused attention, as even minor failures can significantly impact the system performance. Consequently, identifying the key components becomes crucial. This study presents a novel Failure Mode, Effects, and Criticality Analysis (FMECA) approach specifically tailored for high-speed rail traction systems. It addresses limitations in conventional FMECA, notably the expert subjectivity involved in Severity (S), Occurrence (O), and Detectability (D) scores, as well as the ambiguity associated with Risk Priority Numbers (RPNs). A two-stage methodology is proposed: (1) An expert qualification process that filters and calibrates initial S, O, D assessments to mitigate individual bias. (2) A game-theoretic weighting scheme optimally integrates weights from Fuzzy Analytic Hierarchy Process (FAHP) and the Method based on the Removal Effects of Criteria (MEREC), objectively determining the relative importance of S, O, D while resolving RPN non-uniqueness and expert uncertainty. A case study conducted on a representative EMU traction system validates the proposed methodology, identifying the Four-Quadrant rectifier, TCU module, traction motor, ventilation system, and carbon skateboard as the key components. Comprehensive sensitivity analysis further demonstrates that the proposed approach exhibits improved robustness over individual methods (FAHP, MEREC, etc.), confirming its enhanced reliability for key component identification in complex engineering systems.
This study examines common-cause failure (CCF) in complex multi-state systems (MSSs), arising from factors such as external environmental influences and the aging of internal units. The random variables associated with system unit parameters are not restricted to exponential distributions and may instead follow various distribution types. The working time of each system unit is modeled using PH distributions, and a matrix-based analytical approach is applied. An improved universal generating function (UGF) is developed on the basis of the PH distribution, forming a reliability analysis method that integrates the enhanced UGF with PH modeling. The reliability of individual units under CCF is evaluated using the weight influence vector method and the factor model, and compared with those obtained under independent unit failures. Conducting CCF analysis is necessary, as it substantially reduces errors in practical applications. The variable speed hydraulic system of a pipe-lifting machine is used to demonstrate the accuracy and applicability of the method, and maintenance strategies are formulated to enhance system reliability. Protective measures are developed based on actual operating conditions to reduce CCF and improve system reliability. The study addresses the research gap in constructing a reliability model and performing corresponding computational analysis for cases in which unit parameter random variables in a multi-state system (MSS) may follow different distribution types under common-cause aging in engineering practice. Accounting for distributional differences in unit parameters provides new approaches for reliability analysis, strengthens the theoretical framework of MSSs, and offers practical guidance for engineering applications.
Bike-sharing systems are becoming increasingly important for sustainable urban mobility, but they need much maintenance, making them less efficient. Although predictive maintenance has been examined using technical and usage data, there has been insufficient focus on user-centric information as an indicator of bicycle reliability. This study examines the incorporation of user ratings into predictive maintenance models for the Lisbon GIRA bike-sharing system. We used linear mixed models and generalized additive models to examine a 1-month dataset with travel records, user feedback, and maintenance logs. Four formulations of Remaining Useful Life were developed, with riding time exhibiting the most reliable prediction capability. The findings show that user ratings are favorably related to Remaining Useful Life, and both models had an R2 of about 0.50. Residual analysis often undervalues severe results, which successfully indicate the need for maintenance sooner. Kernel density estimations and bathtub curve analysis further demonstrated the degradation effects linked to cumulative usage. The results suggest that putting the user at the heart of information enhances predictive maintenance strategies, making operations more dependable and scheduling maintenance easier in large-scale bike-sharing system operations.