
Simulation-based calibration aims to infer unknown parameters of complex simulation models by aligning model outputs with real-world observations. When simulation runs are computationally expensive, statistical emulators trained on simulation data are used to efficiently approximate the model. An intelligent, adaptive selection of simulation inputs for building the emulator can substantially improve the efficiency of the calibration process. This task is particularly challenging for stochastic simulations with noisy outputs, since both selecting new input locations (exploration) and allocating repeated runs at existing inputs (replication) are essential for efficiently learning the input-output relationship. In this paper, we introduce an active learning framework that adaptively balances exploration and replication for data-efficient calibration. Our uncertainty-aware acquisition criterion targets learning the posterior density of the unknown simulation parameters, and we derive two corresponding forms of the acquisition function for exploration and replication. Building on these, we propose a strategy that, at each stage of the sequential design, chooses between exploration and replication to most effectively reduce the uncertainty in the estimate of the posterior density of the simulation parameters. Experiments on synthetic benchmarks and a real epidemiological model demonstrate that our approach significantly improves learning of the posterior distribution of the simulation parameters while reducing the number of required simulations, making it well-suited for expensive stochastic simulation settings.
Except a few, the majority of the literature on monitoring ordinal data consider independent and identically distributed processes, where samples of data are collected sequentially in time. However, stationary ordinal processes exhibiting serial dependence are also common in many real-world process monitoring applications. This study proposes three classes of novel control charts for monitoring serially dependent stationary ordinal processes. Instead of sample statistics, individual observations are utilized. Exponentially weighted moving-average smoothing of a sequence of estimates is used for estimating the probability mass (cumulative distribution) function of the process. Defined real-valued functions of the probability mass (cumulative distribution) estimates are then used as the statistic plotted on the control charts. The methods are designed to be sensitive to a shift in the marginal distribution. Average run length performance of the control charts are computed under a comprehensive set of data-generating process models, which are inspired by real-world examples and exhibit quite different serial dependence structures. The performances of the proposed control charts are evaluated and compared to provide recommendations for implementations. The results show that the class of demerit-type charts generally perform better than the others. To illustrate the application and interpretation of the proposed methods, a real-world data example on monitoring of heating, ventilation, and air conditioning systems in passenger rail coaches is discussed.
For the case of sensor health monitoring, keeping the correlation level of two sensors measuring related quantities under surveillance is a promising idea. The corresponding raw data streams behave often unsteady. But having a stable correlation level is a typical in-control pattern, whereas mean and variance may vary from time window to time window. However, there are nearly none control charts for monitoring correlation. Here, we consider EWMA control charts for the linear correlation coefficient & rhov;. Despite it is known for a long time, the usage of the explicit (assuming normally distributed data) distribution of the estimator of & rhov; while setting up a control chart seems to be non-existent. Here, we build an EWMA chart utilizing this estimator, namely the Pearson correlation, and calculate the most popular performance measure, the zero-state average run length (ARL), by means of various numerical methods. Less surprisingly, the two standard methods work poorly for certain chart designs. We solve these problems by utilizing piece-wise collocation. Moreover, we examine further configuration details and provide some guidelines. Two applications illustrate the usefulness of monitoring the & rhov; level.
Order-of-addition experiments are often conducted to study physical systems and production processes in which different component addition orders may yield different final responses. By fitting parametric models to the observed data, active effects can be initially identified to understand how the experimental responses depend on the component addition orders. Based on the identified effects, prediction models can then be built to determine optimal orders. Several efficient experimental designs and analysis methods have been proposed to address such order-of-addition problems in recent years. An R package called OofAExp is introduced in this article to implement these proposals. The package is illustrated by several examples. In addition, a series of simulation studies is conducted to determine the default values of some functions.
Additive manufacturing processes, such as laser powder bed fusion, offer great customization capabilities but often suffer from complexity and variability that can lead to defects and require costly post-process inspections. To enhance real-time, in-situ qualification and reduce the reliance on resource-intensive empirical testing, various machine learning models have been proposed. However, three critical challenges remain: (i) transferring knowledge across different fabrication settings without inducing negative transfer, (ii) achieving robust uncertainty quantification to account for the inherent stochasticity of additive manufacturing processes, and (iii) ensuring computational efficiency and interpretability for real-time applications. In this work, we propose a novel transfer learning methodology based on the partially stochastic machine learning model enhanced with Gaussian Process priors over part-specific metadata. The proposed methodology facilitates flexible knowledge transfer, reduces the risk of negative transfer, and provides well-calibrated uncertainty estimates at reasonable computational costs. Numerical studies on the real-world additive manufacturing dataset and simulation dataset are presented to evaluate the effectiveness and reliability of the proposed method and demonstrate the advantage of the proposed method over existing benchmark approaches. The source code will be publicly available upon paper publication.
Monitoring urban air quality requires statistical tools that can account for strong temporal dynamics, spatial dependence, and within-day variation, while providing timely signals for operational decision-making. In this case study, we consider the problem of monitoring hourly ozone (O3) concentration profiles observed across a network of monitoring stations in Beijing, China. Daily O3 evolution is represented through spatiotemporal functional profiles, whose relationship with meteorological covariates is modeled using a geostatistical functional mixed-effects framework. The monitoring objective is to ensure that this relationship remains stable over time and to detect deviations as they occur within the course of a day. To this end, a residual-based multivariate exponentially weighted moving average (MEWMA) control chart is employed within a Phase I-Phase II statistical process monitoring setting. The procedure is calibrated using a dependence-aware block bootstrap and its finite-sample behavior is assessed through Monte Carlo simulation scenarios tailored to the Beijing application. The proposed monitoring framework is demonstrated on hourly O3 concentrations collected from a ground-level monitoring network, illustrating how functional and spatial modeling can be combined with SPM tools to support real-time environmental monitoring.
The main goal in the analysis of a screening experiment is to correctly identify the set of active factors. Lenth's method (Lenth 1989) has become a standard approach for the analysis of saturated, two-level orthogonal screening designs due to its power and ease of implementation. Although various modifications and alternatives have been proposed in the intervening years, Lenth's method remains a primary option in standard statistical software for the analysis of experiments. Despite its widespread use, limitations to Lenth's method include (1) its inapplicability to nonorthogonal saturated screening designs such as many optimal saturated main effects designs and orthogonal designs with missing observations, and (2) its restrictive reliance on a "pseudo" standard error. This work provides an alternative to Lenth's method that is applicable to both orthogonal and nonorthogonal designs. It's effectiveness relies in part on a novel approach to estimating the error variance (and by extension the standard errors of the regression coefficients). Simulation studies reveal the superiority of our approach to current best alternatives.
In semiconductor manufacturing, accurate and timely anomaly detection is critical. However, the presence of a large number of tightly controlled sensors often results in excessive false alarms, and sensor clustering has emerged as one of the practical approaches to mitigate this issue. This case study develops and validates a synchronization-based similarity framework, referred to as the Sync Ratio, for clustering sensor signals based on the temporal co-occurrence of rare events. Rather than relying on magnitude-based correlations such as Pearson correlation, the proposed framework emphasizes shared anomalous behavior, enabling the identification of physically related sensors with numerically dissimilar responses-without requiring detailed domain knowledge. The contribution lies in reframing inter-sensor similarity through rare-event synchronization and demonstrating its effectiveness at scale in a real-world manufacturing environment. Using real data from a dry etcher, the approach demonstrates improved capability in distinguishing true process anomalies from false alarms and outperforms several benchmark clustering methods in practical fault analysis tasks. Our findings indicate that synchronization-focused sensor clustering can provide an intuitive and effective tool for enhancing process monitoring and quality improvement in semiconductor manufacturing.
Control charts are used in diverse areas of applications, from manufacturing to healthcare to cybersecurity, to name a few, for statistical process monitoring and control. Among these charts, Phase I control charts are designed to retrospectively monitor process data. As has been recognized in the literature, these charts are very important in practice for process analysis and improvement. They serve as a foundation for Phase II control charts that are used for prospective monitoring by providing, for example, reference data, parameter estimates, and ideas about the shape of the underlying distribution. Shewhart control charts are typically recommended in Phase I. This paper introduces the R package PH1XBAR that can be used to construct Phase I Shewhart-type control charts for the mean, which is a common problem in practice. In addition to handling i.i.d. data, this package implements some recently published methodological advances and fills a gap in the currently available software literature, by dealing with individual autocorrelated data as well as within-subgrouped equally correlated (variance components) data. The Phase I control charts are designed (constructed) for a nominal False Alarm Probability, which is the recommended in-control performance metric in Phase I. In addition to a review of the methodology, this paper introduces the contents of the package, illustrates the uses, and presents several applied data examples with relevant code. It may be observed that the proposed methods in the package are useful in many real world applications.
Planning preventive maintenance (PM) actions for a fleet of repairable systems is not simple due to their complex dependent structure and variability among systems. An optimal PM policy can be established by considering a tradeoff between system reliability and maintenance costs over the lifespan of repairable systems. In this paper, we propose age-reduced nonhomogeneous Poisson process (NHPP) models for doubly-censored (left- and right-censored) or bathtub-shaped recurrent failure data from multiple repairable systems. We apply modeling frameworks of mixed-effects and frailty to the proportional age-reduced NHPP model, of which parameters are estimated by the maximum likelihood (ML) method, except for the improvement factor that is assumed to be known. The proposed models explicitly involve between-system variation through random-effects or frailty, along with a common baseline for all the systems through fixed-effects for non-normal data. Given the estimates for the proposed models, we derive the optimal aperiodic PM policy by considering whole population of repairable systems rather than a single system to reflect practical environments where the PM executes imperfect repairs on same-typed systems following the same schedule in a lump. The optimal policy aims at determining irregular PM check-points and useful lifetime with the objective of minimizing the expected total maintenance cost per unit of time. Analytical results of two real-world examples show prominent applications of the proposed models and methods to doubly-censored or bathtub-shaped failure patterns from a fleet of repairable systems for the purpose of reliability prediction and maintenance optimization.
Commercial Radio Frequency Identification (RFID) inventory tracking systems have been successfully deployed in numerous industries, but their viability for complex nuclear material environments is less understood. Success of RFID in this space requires thorough examination of factors which influence its ability to inventory random collections of objects, without much prior knowledge about the objects' properties. In this work, we design and analyze experiments to identify relevant handheld RFID settings which optimize performance for one such challenging scenario, a typical cluttered shelf configuration comprised of numerous metal containers. We specify a full factorial split-plot design, and model the probability of a successful match for each container through a hierarchical generalized linear mixed model. The model is used to infer factors which maximize match probabilities for individual containers and the full configuration, and examine sensitivity of performance to these factors. Crucially, we model match probabilities without explicitly accounting for specific covariates about the containers themselves, in order to mimic practical scenarios in which information about objects being inventoried by RFID is unavailable.
In chemical and engineering sciences, most factors are quantitative, benefiting from investigation at three distinct levels when studying quadratic or non-linear relationships. Consequently, definitive screening designs (DSDs), a class of 3-level foldover designs introduced by Jones and Nachtsheim (2011), have gained prominence over traditional 2-level designs such as fractional factorial and Plackett-Burman designs. This article extends the framework of conference matrix-based DSDs by introducing a new class of circulant weighing matrix-based screening designs, offering greater design flexibility. Like DSDs and the orthogonal minimally aliased response surface (OMARS) designs recently proposed by N & uacute;& ntilde;ez Ares and Goos (2020), these new designs preserve orthogonality among main effects and between main and second-order effects (i.e., quadratic effects and two-factor interactions), while ensuring the absence of full aliasing among second-order effects. We refer to these designs as OMARS designs to highlight their ability to support both factor screening and response surface exploration in a single step. We present a comprehensive catalog of 149 new OMARS designs with desirable projection capability, accommodating up to 50 factors.
Ordering factorial designs (OFDs) play a crucial role when both the level combinations of factors and the addition orders of components affect the responses. All the existing construction methods of OFDs are proposed for optimal designs under some specific models. When the true model is unknown or complex, space-filling designs can be considered. However, the current space-filling criteria cannot be used for OFDs directly. In this article, we define the stratification property of an OFD to measure the low-dimensional space-filling property of a design. Two construction methods are also provided for OFDs with good stratification properties. Some of the resulting designs have fewer run sizes, more flexible numbers of components, or can accommodate more factors with the same run sizes than the existing OFDs. Moreover, simulations on a modified traveling salesman problem demonstrate the superiority of the constructed designs.
Monitoring product lifetimes is essential for ensuring quality and reliability, with control charts based on lifetime testing commonly used to assess whether the production process is in control. To ensure timely feedback, censoring is often applied in lifetime tests. However, the performance of the control chart can heavily depend on the choice of censoring time c. A small c may result in significant information loss, complicating the inference about the production process's state, while a large c prolongs the test duration, defeating the purpose of censoring. In this context, this study proposes a computational framework for evaluating the average time to signal (ATS) to facilitate the determination of optimal c when monitoring Type-I censored lifetimes from the one-dimensional exponential family. Our framework relates the monitored plotting statistics to sufficient statistics of censored lifetimes through two distribution-dependent functions, simplifying the ATS calculation to evaluate the cumulative distribution function (cdf) of the sufficient statistics. To ensure both the accuracy and computational efficiency of our framework, an efficient numerical tool, the recursive discrete convolution algorithm, is developed to approximate complex integrals when evaluating the cdf of the sum of truncated random variables. The flexibility of the proposed framework is demonstrated through its application to censored gamma and lognormal lifetimes. Numerical analysis reveals how ATS varies with c, offering guidance for selecting the optimal c. Two case studies further illustrate the potential of the proposed framework in lifetime test design, providing practical insights for balancing information retention and test duration in lifetime monitoring.
In this article, we extend the continuous Bernoulli to a shape-flexible distribution on the unit interval (0,1). Unlike the well-known beta distribution, it admits a closed-form cumulative distribution function (CDF). The extended model preserves the analytical tractability of the original continuous Bernoulli distribution while allowing greater flexibility. We derive explicit expressions for the moments and, most notably, for the stress-strength reliability function in closed form. Maximum likelihood estimation is implemented via efficient fixed-point and scoring-based routines. We then develop a logistic link regression that treats stress-strength reliability as a covariate-dependent quality index for bounded outcomes, yielding interpretable reliability curves. An illustrative application demonstrates that the model captures complex reliability patterns and offers practical utility for reliability and quality engineering.
Multivariate process monitoring offers numerous advantages for quality practitioners, including the ability to assess multiple variables using a single chart, thereby controlling overall type 1 error rates. However, in practical quality monitoring scenarios, different quality characteristics often possess varying levels of risks, customer value, and costs. Assigning equal weights to all variables may not accurately reflect the reality of quality monitoring practice, potentially reducing the effectiveness of critical measures. To bridge the gap between multivariate process monitoring and real-world quality monitoring, this research proposes a criticality assessment framework to be conducted prior to designing a multivariate monitoring system. By incorporating relative weights, this approach aims to maintain the power to detect changes in critical variables while controlling the overall false alarm rate. To demonstrate the practical implementation of this framework, we apply it to a Hotelling chi 2 chart, resulting in the development of a critical-to-X chart. The proposed criticality assessment framework offers significant advancements in multivariate process monitoring, enabling quality practitioners to effectively prioritize variables based on their levels of risk and importance. This approach not only enhances the accuracy and efficiency of quality monitoring, but also offers a degree of protection of key variables in high-dimensional settings and aligns with the dynamic nature of quality improvement efforts. The findings of this research have far-reaching implications for researchers and industries looking to optimize their quality management systems and drive continuous improvement.
With the explosive growth of data, matrix data has become increasingly common in fields such as genetic engineering, remote sensing satellites, and environmental and atmospheric sciences. Monitoring matrix data has also become increasingly important. Control charts are important tools in the study of statistical process control. Traditional research assumes that the monitored data follows a univariate normal distribution or a vector-valued normal distribution. Matrix sequences can explain the correlations between different row and column variables simultaneously. However, research on the monitoring of such sequences is very limited. In this article, based on the matrix normal distribution, we propose three new matrix control charts. We study the changes in the average run length of three control charts when a matrix normal process shifts. In the application of control charts, the true parameters are rarely known explicitly and are usually obtained through parameter estimation methods. Therefore, estimation errors are inevitable. This article further investigates the effect of parameter estimation errors on control charts. The application of the model and control charts is demonstrated in the example section.
Feedback information concerning automotive quality often suffers from significant delays, making online consumer complaints an invaluable real-time source of information for monitoring and assessing product quality. Given that the frequency of online complaints is influenced by numerous factors, such as automobile quality, sales, Internet development, and public awareness of rights protection, it exhibits significant auto-correlation and dynamics. However, the existing modeling methods have been proven to be unreliable in practical applications, because they often assume that in-control (IC) processes remain static and employ models with fixed parameters. To this end, a dynamic modeling framework that integrates generalized linear regression with an integer-valued auto-regressive (INAR) state space model is proposed to capture the evolving nature of the process. Then, a procedure combining the Extended Kalman Smoothing (EKS) with the Expectation Maximization (EM) algorithm, referred to as EM-EKS, is used to estimate the model parameters. Furthermore, for online monitoring of any upward shifts in the number of complaints, a control chart (denoted as SDC-INAR(1)-G) with one-step-ahead forecasting value as the plotting statistic is constructed. Simulation studies show that the proposed SDC-INAR(1)-G method consistently exhibits much better performance than three benchmark approaches in different scenarios. Finally, the proposed SDC-INAR(1)-G method is applied to monitor online complaints of the Volkswagen Sagitar, focusing on two specific cases: the stationary process of "brake abnormal noise" and the non-stationary process of "transmission abnormal noise." The results further demonstrate that the SDC-INAR(1)-G method outperforms the static approaches in both cases. The state-space framework and adaptive EM-EKS parameter estimation of the proposed SDC-INAR(1)-G method ensure robust sensitivities to different shifts, offering reliable monitoring for both stationary and non-stationary data, at the same time, remaining computationally efficient for real-world applications.
Process monitoring is crucial for ensuring the safe and reliable operation of equipment. Current process monitoring techniques often rely on distribution assumptions, such as the Gaussian distribution, which can be challenging to fulfill or justify in real-world applications. Moreover, many processes are influenced by exogenous covariates. While these covariates are not the primary focus of process monitoring, they offer valuable insights for a comprehensive understanding of process variability. To address these issues, we develop a distribution-free monitoring technique for covariate-regulated processes. Specifically, we propose a multivariate adaptive regression splines (MARS)-based quantile regression model. In contrast to traditional MARS techniques that employ a greedy search to include one pair of basis functions at each step, our approach involves incorporating all basis functions initially. This allows us to transform the problem into a linear programming and effectively incorporate constrains to avoid the quantile-crossing issue. Additionally, we introduce a three-level monitoring framework that progresses from monitoring single observations to quantile values and the entire quantile curves. We conduct numerical and real case studies to demonstrate the effectiveness of this technique.