Condition-based maintenance (CBM) has attracted significant attention in advanced manufacturing systems. In practice, degradation of production units is often influenced by their production rates. Unlike traditional CBM models that treat maintenance optimization as a stand-alone problem, we develop a joint optimization model of condition-based maintenance and production for multi-stage manufacturing systems. The resulting joint dynamic optimization problem is formulated as a Markov decision process (MDP) with the aim at maximizing the expected total production benefit. To address the curse of dimensionality and improve scalability in decision-making, a hierarchical multi-agent reinforcement learning framework is put forth to solve the MDP, leveraging a two-level architecture that decomposes the problem into manageable subtasks for maintenance and production. A distributed training strategy is introduced to independently train stage-level agents, significantly improving scalability. Experimental results from a semiconductor manufacturing facility demonstrate that the proposed joint dynamic optimization method significantly improves production benefit by maintaining a balance between production efficiency and system reliability.
This paper proposes a reliability-based maintenance optimization method for reliability modeling and maintenance planning of reusable phased mission systems (R-PMSs). Initially, a multi-phased Wiener process-based reliability modeling method, combined with the Binary Decision Diagram-based model, is created to assess the reliability of R-PMSs. The method considers the initial states of components to be perfect/imperfect fixed as a basis to reflect the impacts of multiple usages of R-PMSs. Subsequently, a maintenance optimization model based on the reliability assessed and incorporating multiple maintenance actions is constructed to allocate the maintenance resources into the maintenance activities in advance of subsequent missions of R-PMSs. The model balances the maintenance cost, overall reliability, and task requirements. The superiority and performance of the proposed method are validated by numerical and engineering analysis. The results confirmed that the proposed method contributes to reliability modeling and maintenance scheduling of R-PMSs.
Accurate reliability assessment of advanced manufacturing systems is essential for ensuring production efficiency, reducing downtime, and enabling intelligent maintenance strategies. In practical industrial environments, state observations obtained from sensors or expert evaluations are often imprecise. Effectively utilizing this uncertain information can substantially improve the precision of reliability evaluations. Conventional methodologies often encounter limitations in addressing this challenge, as manufacturing systems are generally characterized by networked production line configurations rather than traditional serial or parallel structures. Moreover, the effective integration of imprecise observational data is essential for the continuous updating of system reliability. This study introduces a novel approach for dynamic reliability evaluation of a multi-state manufacturing system (MSMS), incorporating both rework mechanisms and buffer elements to enhance the accuracy and applicability of system reliability assessments. The MSMS model can effectively depict the gradual degradation processes and diverse performance levels of manufacturing systems, allowing for a more realistic and detailed representation of system behavior over time compared to traditional binary-state models. This study employs the multistate flow network (MFN) model to construct the MSMS reliability assessment framework from a network structure perspective. Dynamic Bayesian networks (DBNs) are developed to update the reliability function of an individual MSMS by incorporating evidential observational data. An illustrative case study on the reliability update of an aluminum alloy wheel production line is presented to demonstrate the proposed methodology. The case study results further confirm the effectiveness of the approach.
The failure of critical components in large-scale, complex equipment such as hydraulic excavators often exhibit dynamic characteristics, including failure priority, sequence dependence, and functional interdependence. These coupled mechanisms complicate fault diagnosis and obscure the identification of failure propagation paths, presenting significant challenges for reliability assessment. To address these challenges, this study develops an advanced reliability evaluation framework based on a discrete-time Bayesian network (DTBN) to model and assess the dynamic failure processes of hydraulic excavators. The framework incorporates multi-state logic into dynamic failure modeling, enabling the integration of temporal and functional dependencies within a unified probabilistic structure. By incorporating dynamic Bayesian reasoning, the framework effectively evaluates system reliability, identifies fatigue-deformation wear synergistic failure as the primary failure mode, and validates the assessment results through a robust consistency analysis. Furthermore, a robustness evaluation method for parameter uncertainty under time-varying failure behavior is proposed, and the uncertainty propagation mechanisms of static and dynamic gates in sequential systems are analyzed. The proposed approach establishes a systematic, data-driven foundation for reliability assessment of hydraulic excavators operating under complex dynamic failure conditions.
Accurately evaluating the reliability of hydraulic excavators has long posed a significant challenge in construction machinery. Existing approaches have difficulty integrating failure information generated throughout the entire product lifecycle, including design, manufacturing, testing, operation, and maintenance. To address this limitation, this paper introduces a multi-phase failure data fusion method for reliability assessment of construction machinery systems, integrating simulation outputs, experimental results, historical failure records, and expert knowledge. An adaptive Markov chain Monte Carlo (MCMC) algorithm is applied to estimate model parameters and support system-level reliability evaluation. Case studies involving hydraulic excavators show that the proposed method enables quantitative reliability assessment for both complete machines and critical subsystems while effectively identifying components with low reliability. The approach provides a theoretical basis and practical framework for enterprises to rapidly and accurately assess the reliability of complex mechanical products, improve design quality, and enhance maintenance strategies, thereby delivering substantial engineering value.
This paper presents a hybrid framework for reliability assessment that connects unsupervised health indicator (HI) construction with stochastic degradation modeling. In the first stage, a temporal convolutional autoencoder is employed to extract multi-scale temporal information from multi-sensor measurements. Attention mechanisms are introduced to emphasize degradation-sensitive features, while a lightweight adaptive denoising module suppresses irrelevant fluctuations and improves the robustness of HI construction. In the second stage, the constructed HI trajectories are modeled using a random-effects nonlinear Wiener process with an acceleration factor. This probabilistic model accounts for unit-to-unit heterogeneity and nonlinear degradation behavior and is used to estimate the time-dependent reliability function and reliable-life measures under a predefined failure threshold. Experiments on C-MAPSS FD001 and N-CMAPSS DS02 demonstrate that the proposed framework produces informative degradation trajectories and achieves favorable reliability-assessment performance compared with the selected baseline models. These results support its potential application in condition-based maintenance and operational decision-making for industrial equipment.
Most existing cross-domain fault diagnosis methods rely on the assumption of class consistency between source and target domains. However, this assumption is frequently violated in practical industrial scenarios, where varying operating conditions often correspond to inconsistent fault type sets (i.e., class mismatch). To address this challenge, this article proposes a prompt-guided disentanglement and fusion (PGDF) framework, which leverages the inherent semantic structure of natural language to guide the disentanglement of machine states under different operating conditions. Specifically, PGDF constructs learnable prompts combined with domain and class tokens to construct structured joint representations. Furthermore, a cross-modal decomposed fusion module is designed to disentangle linguistic features into independent domain and class factors. Crucially, through a disentanglement-recombination strategy, these factors are fused with observational data, enabling the PGDF to learn expressive features that independently capture domain and class relationships. Extensive experiments on four datasets involving cross-domain, cross-unseen domain, and cross-machine scenarios demonstrate that PGDF significantly outperforms state-of-the-art methods in diagnostic performance.
Accurate prediction of unknown labels from feature-label datasets using machine learning is critical for applications spanning drug discovery, disease diagnostics, and climate science. However, challenges persist with limited data, high-dimensional inputs, and multi-fidelity scenarios. We developed multi-fidelity tabular prior-data fitted network (MFTabPFN), a general-purpose multi-fidelity model integrating low- and high-fidelity data through a hierarchical transformer architecture to enhance prediction accuracy and uncertainty quantification (UQ). MFTabPFN captures cross-fidelity correlations while seamlessly adapting to single-fidelity data. An active learning framework further enhances scalability by prioritizing high-value data for model refinement, minimizing resource demands in resource-intensive tasks. Evaluated on various tasks such as forest fire burned area prediction, wine quality assessment, and computational fluid dynamics, MFTabPFN outperforms state-of-the-art methods, achieving varying degrees of prediction accuracy improvement. Its versatility and robust prediction and UQ capabilities across single- and multi-fidelity datasets position MFTabPFN as a promising tool for data-driven discovery in diverse applications.
Cross-domain few-shot fault diagnosis is hindered by feature divergence arising from complex domain shifts, including varying operating conditions, compound fault modes, and individual train discrepancies. Existing meta-learning approaches primarily prioritize inter-class separability but often neglect intra-class compactness, yielding dispersed feature clusters vulnerable to misclassification. To address this limitation, an Intra-inter Class Variation Rectification Network (IICVRN) is proposed. Specifically, a self-attention feature enhancement module is first constructed to extract discriminative local patterns via feature reconstruction, thereby improving initial feature representation. Subsequently, a core bidirectional variation rectification module is designed to synergistically optimize the feature space geometry. It combines a forward reconstruction strategy to maximize interclass distances with a novel backward reconstruction strategy that imposes isotropic geometric constraints to compress intra-class variations. Extensive experiments on three heterogeneous high-speed train datasets demonstrate that IICVRN achieves state-of-the-art performance across cross-operating condition, cross-type, and cross-device scenarios, outperforming advanced baselines by over 5% in challenging cross-device tasks. Both theoretical derivation and experimental visualizations validate the mechanism's efficacy in regulating feature variations.
Reliable fault detection in train transmission systems is essential for safe and efficient railway operation, yet it is challenged by variable operating conditions and scarce fault samples. To address these issues, a novel unified health domain relation learning (UHDRL) framework is proposed. Specifically, a pseudofault sample library is constructed to generate diverse synthetic fault examples, reducing UHDRL's reliance on healthy samples. A unified health domain mechanism is designed to map the different operating conditions into a common feature space, thereby reducing distribution shifts caused by operating variations. Additionally, a health relation learning mechanism is proposed to construct feature pairs between healthy representations and pseudo-faults to uncover intrinsic and discriminative attributes of health states. Experiments on three train transmission systems, conducted under both deterministic and non-deterministic operating condition changes, demonstrate that UHDRL is highly adaptable and robust in zero-fault-sample settings, improving detection accuracy by over 12% compared with existing methods.
Reliability analysis of engineering machinery faces significant challenges due to highdimensional input spaces and scarce failure samples, making accurate reliability assessment difficult. Recognizing the potential of Deep Gaussian Processes in this domain, this paper proposes a surrogate modeling method integrating Deep Gaussian Processes with active learning strategies and develops a dedicated toolbox. The toolbox employs a modular architecture to enable scalable layerwise operations and includes computational efficiency benchmarks. Through two structural reliability case studies, including a numerical example with explicit limit states and a practical engineering application requiring finite element analysis, the proposed method demonstrates superiority over traditional surrogate models in terms of sample efficiency and prediction accuracy. The findings highlight the practical utility of Deep Gaussian Processes in advancing reliability engineering practices.
Wheel profile prediction is critical to the maintenance of high-speed electric multiple units. This study proposes a geometric-constrained residual Bayesian neural network for high-accuracy prediction of worn wheel profiles. A vehicle-track dynamic simulation model is further employed to determine whether predicted worn profiles exceed the limit and to estimate the time-dependent reliability associated with profile degradation. The proposed approach is evaluated via quantitative and qualitative comparisons with a conventional wheelset wear prediction method. Results indicate that the geometric-constrained residual Bayesian neural network improves wheel wear prediction accuracy. In addition, the simulation-based prediction is more conservative, and the wear range observed in service is wider than that obtained from simulation. Wear patterns of wheels within the same wheelset also show strong correlations. Moreover, wheel diameter affects wear behaviour, including the wear degradation rate, equivalent conicity, the vehicle critical speed, and the Sperling index. Hence, Maintenance scheduling should therefore account for wheel diameter effects to ensure robust performance and durability.
Due to the high-speed and heavy-load characteristics of aero engine gear, vibrations and noise are often produced, along with engagement and disengagement impacts. These factors can lead to various failure modes and the tooth surface contact fatigue is usually the primary failure mode. By making minor adjustments to the tooth surface, the relevant impacts during engagement and disengagement can be reduced. For that reason, this paper focuses on the modification of tooth profile at the tip and calculates the optimal amount of modification for the gear through theoretical analysis. Further, a parametric simulation of the tooth tip profile modification is established and the Kriging surrogate model is used to optimize the optimal modification curve and the tip rounding parameters. The relevant results indicate that the maximum contact stress on the tooth surface is reduced by 23.57
Engineering reliability analysis often requires surrogate modeling of multiple correlated structural responses under limited simulation data. Conventional single-output Gaussian processes cannot directly exploit inter-response correlations, while shallow multi-task Gaussian processes may be insufficient for strongly nonlinear engineering systems. To address these issues, this paper develops a Multi-Task Deep Gaussian Process (MTDGP) framework for sparse multi-response surrogate modeling and reliability assessment. The proposed model combines the nonlinear latent representation of deep Gaussian processes with the task-correlation structure of multi-task Gaussian processes. To improve posterior inference, a Gibbs sampling framework with block hidden-layer updates in whitened coordinates is introduced, which enhances sampling stability and alleviates inefficient latent-state updates. The method is evaluated using two synthetic benchmarks and a portal crane reliability case. The numerical examples show that MTDGP can improve predictive accuracy when the target response is sparsely observed but correlated auxiliary responses are available, with particularly clear gains in the five-dimensional benchmark. Sampling diagnostics indicate that the block update strategy reduces non-moving latent updates compared with coordinate-wise sampling. In the portal crane case, MTDGP gives stable response-surface predictions and competitive reliability estimates under incomplete response observations. Nevertheless, its advantage over shallow multi-task Gaussian process models is not universal. Overall, the findings suggest that MTDGP is a promising surrogate framework for sparse multi-response engineering problems, while further improvements in posterior inference and reliability-oriented active learning remain necessary.
Rotating machinery fault detection remains challenging because real fault samples are scarce and operating conditions often fluctuate. Existing methods tend to suffer from false alarms and missed detections under condition-induced distribution shifts, resulting in unstable detection performance. To address these challenges, an innovative adaptive relation network (ARN) is proposed. Specifically, a pseudo-fault sample generator (PFSG) is employed to generate diverse auxiliary samples and alleviate the model’s reliance on healthy-only supervision. An adaptive alignment mechanism is designed to align healthy sample distributions across different operating conditions, thereby mitigating the influence of condition variability. In addition, an attribute feature mining mechanism (AFMM) is developed to extract condition-related features and unique healthy state attributes. This further improves the discriminative capability of ARN. Evaluation on three independent datasets, including aviation plunger pump (PP), train transmission system (TTS), and bearing datasets, shows that ARN achieves an average accuracy improvement of approximately 13% over the second-best state-of-the-art method. These results verify the effectiveness of ARN in complex industrial scenarios.
Complex systems, particularly in industries such as aviation and aerospace, are characterized by high technical challenges, extended development cycles, and substantial capital investments. Consequently, manufacturers in these fields often adopt modification strategies, leveraging the success of existing systems to minimize risks and costs. Traditional reliability-cost optimization models and associated techniques are primarily designed to address reliability allocation challenges in the development of entirely new systems. However, these models face significant limitations when applied to modified systems, as they fail to accurately represent the unique reliability-cost dynamics associated with one-time costs inherent in modifications. Based on the classical binary reliability-cost function, this paper analyzes the reliability-cost decomposition of the modified system, thus establishes a discontinuous nonlinear optimization model applicable to both continuous and discrete reliability cases. Among the cost decomposition, the gradually increasing term in the objective function represents the reliability improvement cost such as reliability design and manufacture, and the one-time term represents the cost such as verification and certification. Then, a differential evolution algorithm is designed to suit both continuous and discrete reliability cases within a unified framework, aiming to find the feasible reliability reallocation solution and to minimize the cost. Finally, two numerical examples utilizing three common cost functions, namely logarithmic, exponential, and power-law types, are conducted to validate the rationality of the proposed model and to demonstrate the effectiveness of the algorithm. Some comparisons with genetic algorithm and particle swarm optimization algorithm are also presented in view of optimality, mean value and deviation in this paper.
Accurate remaining useful life (RUL) prediction is critical for ensuring industrial system safety and effective maintenance; however, it continues to encounter significant challenges, including feature distribution shifts and cross-domain data scarcity. To address these challenges, this study proposes an adversarial enhanced attention domain adaptation network (AEADAN) for robust cross-domain RUL prediction. The framework integrates a Wasserstein generative adversarial network with gradient penalty and a variational autoencoder to perform effective cross-domain data augmentation and extract domain-invariant degradation features. A dedicated multi-component adaptation module integrates domain discrimination with distribution discrepancy minimization to achieve effective feature alignment, while an exponentially decaying weighting mechanism adaptively balances classification and alignment losses throughout training. Extensive experiments conducted on benchmark datasets demonstrate that AEADAN markedly outperforms state-of-the-art methods in terms of prediction accuracy and generalization capability, thereby providing an effective solution for accurate RUL prediction.
Reusable launch vehicles (RLVs), such as the SpaceX Falcon Heavy and Blue Origin New Shepard, have emerged as pivotal advancements in space technology. The operational lifecycle of RLVs comprises discrete sequential mission phases across multiple flights, conceptualized herein as reusable phased mission systems (R-PMSs). This paper proposes a comprehensive reliability and availability modeling framework for R-PMSs. Firstly, a zoned-shock model is formulated to characterize complex shock effects on multiple components, integrating into their multi-phased Wiener process-based degradation process. Secondly, the system reliability is evaluated with the binary decision diagram model for the phased mission system (PMS-BDD) model. The model further quantifies R-PMS availability by incorporating the varying phase duration and maintenance activities executed over successive mission cycles. A joint optimization model is constructed to maximize the system availability, with varying working phase duration and maintenance criteria. Finally, the application of the proposed method is illustrated via an RLV propulsion system case study.
Large language models (LLMs) have demonstrated remarkable in-context learning (ICL) capabilities, enabling flexible utilization of limited historical information to play pivotal roles in reasoning, problem-solving, and complex pattern recognition tasks. Inspired by the successful applications of LLMs in multiple domains, this article proposes a generative design approach by leveraging the ICL capabilities of LLMs with the iterative search mechanisms of metaheuristic algorithms for solving reliability-based design optimization (RBDO) problems. In detail, Kriging surrogate modeling is employed to replace the expensive simulations, and thus Monte Carlo simulation (MCS) can be used to approximate the probability of failure for design alternatives. Then, an RBDO-informed LLM prompt is designed to dynamically provide critical information to the LLMs, which enables the rapid generation of new high-quality design points that satisfy the reliability constraints while improving design efficiency. With the LLMs as a design generator, the RBDO is an iterative process to obtain feasible design solutions with improved performance. With the Deepseek-V3 model, three case studies are used to demonstrate the performance of the proposed approach for solving RBDO problems. The results indicate that the proposed LLM-based generative RBDO approach successfully identifies feasible solutions that meet reliability constraints while achieving a comparable convergence rate compared to traditional genetic algorithms.
Domain generalization (DG) has emerged as a pivotal paradigm for intelligent fault diagnosis of rotating machinery under non-stationary operating conditions. Diverging from prevalent methods that prioritize the elimination of inter-domain discrepancies via macroscopic feature distribution alignment, this article proposes a structurally parsimonious framework predicated on a novel intraclass similarity spectrum (ISS). Distinct from alignment-centric methods, this article shifts the focus toward intrinsic intraclass structural invariance. The theoretical premise is that while the statistical distributions of monitoring signals drift due to varying operating conditions, the relative geometric topology among intraclass samples preserves robust consistency. Accordingly, the proposed method encodes microscopic intraclass similarities into high-order ISS and employs a streamlined six-layer fully connected network to map these topological features to fault classes. Extensive empirical evaluations on three industrial datasets demonstrate that the proposed method achieves an optimal balance between diagnostic precision and computational efficiency. It significantly outperforms state-of-the-art methods and exhibits superior robustness against severe class imbalance.