This paper introduces a blocked general minimum lower-order confounding (B-GMC) criterion and focuses on constructing optimal three-level blocked designs. We obtain some necessary conditions for constructing B-GMC designs. In addition, all 27-run and some 81-run B-GMC designs are tabulated.
Computational capability often falls short when confronted with massive data, posing a common challenge in establishing a statistical model or statistical inference method dealing with big data. While subsampling techniques have been extensively developed to downsize the data volume, there is a notable gap in addressing the unique challenge of handling extensive reliability data, in which a common situation is that a large proportion of data is censored. In this article, we propose an efficient subsampling method for reliability analysis in the presence of censoring data, intending to estimate the parameters of lifetime distribution. Moreover, a novel subsampling method for subsampling from severely censored data is proposed, i.e., only a tiny proportion of data is complete. The subsampling-based estimators are given, and their asymptotic properties are derived. The optimal subsampling probabilities are derived through the L-optimality criterion, which minimizes the trace of the product of the asymptotic covariance matrix and a constant matrix. Efficient algorithms are proposed to implement the proposed subsampling methods to address the challenge that optimal subsampling strategy depends on unknown parameter estimation from full data. Real-world hard drive dataset case and simulative empirical studies are employed to demonstrate the superior performance of the proposed methods.
As a virtual counterpart of physical systems, the digital twin is an advanced technology that has gained significant importance across various industries due to its numerous benefits and wide-ranging applications. However, optimizing digital twins presents a complex challenge, particularly when dealing with multiple responses. This complexity arises from the intricate computational processes required to identify optimal inputs, which can hinder timely decision-making. In this paper, we propose an approach that transforms multi-objective optimization into a single-objective optimization problem and subsequently applies sequential optimization to the single objective. Our method integrates sequential data collection with digital twin learning techniques, encompassing four key steps: data collection, evaluation, optimization, and decision-making. The effectiveness of the proposed approach is demonstrated through three case studies, highlighting its capability to optimize multi-response digital twins. The simplicity of implementation and remarkable adaptability of this approach make it a powerful tool for capturing and representing the intricate nuances of digital twins. This enhanced digital twin facilitates a deeper understanding of the underlying physical system, ultimately accelerating the optimization of its performance.
If Francis Bacon were born today, he might have said "data is power" instead of his original saying, "knowledge is power." In modern society, data is everywhere. In memory of Deming (a guru in quality), this paper attempts to address the fundamental issue of data quality and how Deming would handle it. Specifically, we attempt to explain what data quality really means, and the critical impact that it has on data science. Statisticians, who understand how to collect high quality data, have much more to contribute to both the intellectual vitality and the practical utility of data science. At the same time, data science challenges statisticians to move out of some familiar habits to engage less structured problems, to become more comfortable with ambiguity, and to engage more scientists in a fruitful discussion on what various parties can bring to this new mode of investigation. Some potential avenues for future research in the collection of high-quality data will be proposed.
In an Order‐of‐Addition (OofA) problem, a response depends on the order in which several components are added to a system. This is critical for many modern applications, such as combinatorial drug therapy for cancer patients, chemical engineering, and NP‐hard optimization problems. If the number of components is large, it is impossible to check all possible orders. This motivates the use of Design of Experiments (DOE) to select an optimal and affordable subset of orders for an experiment. In the past few years alone, over 30 statistical publications have emerged that deal with design and modeling for OofA experiments. This large body of literature is difficult for new practitioners to navigate. The main goals of this article are to provide a comprehensive overview of the main models for OofA experiments, to introduce theoretical and algorithmic methods for constructing optimal OofA experiments, and to describe how to analyze the experimental results so that an optimal order‐of‐addition can be obtained. We provide recommendations for using a model‐free approach for solving OofA problems, as well as recommendations for using OofA models that incorporate the effects of other covariates. We also provide some novel discussion for designing experiments in the case where constraints on the pairwise order of components exist. Recent developments in OofA experiments will greatly impact research in pharmaceutical industries, biomedical science, chemical engineering, and computer science, as they allow researchers to efficiently construct designs for larger numbers of components than before. This article reviews the wide variety of approaches for constructing designs and models for OofA experiments, which is critical for continued innovation in this field.
Order-of-addition experiments have emerged as a cornerstone in modern experimental design, yet the critical role of run-order has been entirely neglected in the literature. This oversight is surprising, given that the run order can significantly influence the cost, efficiency, and validity of the experiment. Certain run orders are inherently more economical and effective, while suboptimal orders may introduce unnecessary complexities or compromise results. To ensure experimental integrity, an optimal run order must minimize factor level changes, especially when transitions between levels are costly-and safeguard against the detrimental effects of time trends, which can skew factor effect estimates and undermine conclusions. Despite its importance, the relationship between run order and order-of-addition experiments remains poorly understood. Our research addresses this critical gap by deriving the theoretical lower bound on the total number of factor level changes for ordered order-of-addition designs. We introduce novel construction methods that achieve this minimum, ensuring both cost efficiency and practicality. Additionally, we develop a rigorous framework to eliminate the bias introduced by linear time trends, proposing robust methods to construct time trend-free ordered designs. These methods lead to the establishment of optimal ordered designs that excel under both criteria. The proposed designed orders go beyond theoretical elegance-they are robust to lurking variables in the external environment, making them invaluable for real-world applications. This study not only shines the spotlight on an overlooked dimension of order-of-addition experiments but also delivers solutions that redefine how these experiments should be conducted. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
The order-of-addition (OofA) experiment involves arranging components in a specific order to optimize a certain objective, which is attracting a great deal of attention in many disciplines, especially in the areas of biochemistry, scheduling, and engineering. Recent studies have highlighted its significance, and notable works have aimed to address NP-hard OofA problems from a statistical perspective. However, solving OofA problems presents challenges due to their complex nature and the presence of uncertainty, such as scheduling problems with uncertain processing times. These uncertainties affect processing times, which are not known with certainty in advance. They introduce heteroscedasticity into OofA experiments, where different orders result in varying dispersions. To address these challenges, a unified framework is proposed to analyze scheduling problems without making specific assumptions about the distribution of these certainties. It encompasses model development and optimization, encapsulating existing homoscedastic studies (where different orders produce the same dispersion value) as a specific instance. For heteroscedastic cases, a dual response optimization within an uncertainty set is proposed, aiming to minimize the dispersion of response while keeping the location of response with a predefined target value. However, solving the proposed non-linear minimax optimization is rather challenging. An equivalent optimization formulation with low computational cost is proposed for solving such a challenging problem. Theoretical supports are established to ensure the tractability of the proposed method. Simulation studies are conducted to demonstrate the effectiveness of the proposed approach. With its solid theoretical support, ease of implementation, and ability to find an optimal order, the proposed approach offers a practical and competitive solution to solving general order-of-addition problems.
Digital twins (DTs) provide dynamic, virtual representations of physical systems, facilitating real-time monitoring, prediction, and optimization through the integration of sensor data, computational models, and feedback mechanisms. Despite their potential, deploying DTs in complex systems often entails substantial computational costs. To address this challenge, recent studies have introduced surrogate models-referred to as digital triplets-as computationally efficient approximations of DTs. In particular, a framework that combines Gaussian process (GP) regression with the maximum projection (MaxPro) designs has shown promising predictive performance. However, the impact of alternative surrogate modeling techniques and the sequential experiments on optimization accuracy and convergence remains insufficiently understood. This study systematically evaluates a range of model-design combinations, including GP regression, polynomial regression, neural networks, and automated machine learning, in conjunction with central composite, optimal, MaxPro, and space-filling designs. Performance is assessed in both two- and six-dimensional DT scenarios, with metrics focused on optimization accuracy and convergence behavior. The findings reaffirm the strong performance of the MaxPro-GP combination, particularly in settings with moderate simulation budgets, thereby extending existing insights in the literature.
The space-filling property and orthogonality are perhaps two most desirable design properties for computer experiments.The space-filling property is appropriate for Gaussian process models, while orthogonality allows the estimated effects to be uncorrelated.This paper presents a general approach for constructing a rich class of orthogonal designs with attractive low-dimensional space-filling properties.This is apparently new in the literature.The construction methods are straightforward to implement.Their theoretical supports are established.Moreover, the resulting designs are flexible in the run sizes.
In an order-of-addition (OofA) experiment, the sequence of m different components can significantly impact the experiment's response. In many OofA experiments, the components are subject to constraints, where certain orders are impossible. For example, in survey design and job scheduling, the components are often arranged into groups, and these groups of components must be placed in a fixed order. If two components are in different groups, their pairwise order is determined by the fixed order of their groups. Design and analysis are needed for these pairwise-group constrained OofA experiments. A new model is proposed to accommodate pairwise-group constraints. This paper also introduces a model for mixed-pairwise constrained OofA experiments, which allows one pair of components within each group to have a pre-determined pairwise order. It is proven that the full design, which uses all feasible orders exactly once, is D- and G-optimal under the proposed models. Systematic construction methods are used to find optimal fractional designs for pairwise-group and mixed-pairwise constrained OofA experiments. The proposed methods efficiently assess the impact of question order in a survey dataset, where participants answered generalized intelligence questions in a randomly assigned order under mixed-pairwise constraints.
Screening designs are widely used to identify active effects from a large number of potential factors, ideally with a relatively small number of runs. This process is typically conducted in the early stages of experimentation. Orthogonal designs (ODs) are preferred because they allow the estimated main effects to remain uncorrelated. While two-level and three-level ODs have been extensively studied, ODs for mixed-level experiments have received less attention. This is an important and timely issue, as mixed-level experiments are common in practical applications. This paper proposes a new class of small mixed-level ODs. These designs ensure the uncorrelated estimation of main effects and minimize run size among all ODs for a given number of factors. Additionally, the proposed designs exhibit favorable confounding patterns between main effects and two-factor interactions, with each main effect being orthogonal to most two-factor interactions. The construction methods leverage Paley's conference matrices. This research fills a critical gap in the literature by introducing new theoretical developments for mixed-level experiments.
Orthogonal Latin hypercubes are widely used for computer experiments. They achieve both orthogonality and the maximum one-dimensional stratification property. When two-factor (and higher-order) interactions are active, two- and three-dimensional stratifications are also important. Unfortunately, little is known about orthogonal Latin hypercubes with good two (and higher)–dimensional stratification properties. A method is proposed for constructing a new class of orthogonal Latin hypercubes whose columns can be partitioned into groups, such that the columns from different groups maintain two- and three-dimensional stratification properties. The proposed designs perform well under almost all popular criteria (e.g., the orthogonality, stratification, and maximin distance criterion). They are the most ideal designs for computer experiments. The construction method can be straightforward to implement, and the relevant theoretical supports are well established. The proposed strong orthogonal Latin hypercubes are tabulated for practical needs.
Screening is an important step of experimental design. It aims to identify a few active factors, among a large number of potential factors. In this paper, we propose two classes of mixed-level screening designs with desirable design properties; such as, low correlations between any two design columns, high design efficiencies (e.g., D- or A-efficiencies), and orthogonality between main effects and two-factor interactions. Conference matrices with skew-symmetric structure play an important role in the proposed construction. Two new construction methods for conference matrices with skew-symmetric structure, recursive and algebraic constructions, are provided. It is shown that the proposed mixed-level screening designs have all desirable design properties and do not require any computer search.
Order‐of‐addition (OofA) experiments have gained renewed attention in recent years, especially in regard to their design. For these experiments, the response is determined by the order in which components are added. A particularly useful design introduced in OofA experiments is the component orthogonal array (COA). The COA maintains pairwise balance between any two components while also ensuring each component appears equally often in each position. In this paper, we propose an efficient algorithm for constructing COAs which can be naturally split into blocks of Latin squares. These blocks can be run sequentially in a systematic order, potentially requiring fewer runs to identify optimal orderings, while also preserving good properties should the overall design be needed. We also show how to extend this construction method to create designs for any number of components.
The joint optimization of production, maintenance, and quality control has shown effectiveness in reducing long-term operational costs in production systems. However, existing studies often assume that changes in the mean value of product quality characteristics in a deteriorating system follow a specific distribution while keeping variance constant. To address this limitation, we propose an innovative method based on the continuous ranking probability score (CRPS). This method enables the simultaneous detection of changes in mean and variance in nonconformities, thus removing the assumption of a specific distribution for quality characteristics. Our approach focuses on developing optimal strategies for production, maintenance, and quality control to minimize cost per unit of time. Additionally, we employ a stochastic model to optimize the production time allocated to the inventory buffer, resulting in significant cost reductions. The effectiveness of our proposed joint optimization method is demonstrated through comprehensive numerical experiments, sensitivity analysis, and a comparative study. The results show that our method can achieve cost reductions compared to several other related methods, highlighting its practical applicability for manufacturing companies aiming to reduce costs.
It is difficult to handle the extraordinary data volume generated in many fields with current computational resources and techniques. This is very challenging when applying conventional statistical methods to big data. A common approach is to partition full data into smaller subdata for purposes such as training, testing, and validation. The primary purpose of training data is to represent the full data. To achieve this goal, the selection of training subdata becomes pivotal in retaining essential characteristics of the full data. Recently, several procedures have been proposed to select "optimal design points" as training subdata under pre-specified models, such as linear regression and logistic regression. However, these subdata will not be "optimal" if the assumed model is not appropriate. Furthermore, such subdata cannot be useful to build alternative models because it is not an appropriate representative sample of the full data. In this article, we propose a novel algorithm for better model building and prediction via a process of selecting a "good" training sample. The proposed subdata can retain most characteristics of the original big data. It is also more robust that one can fit various response model and select the optimal model. Supplementary materials for this article are available online.
Sequential Latin hypercube designs (SLHDs) have recently received great attention for computer experiments, with much of the research restricted to invariant spaces. The related systematic construction methods are inflexible, and algorithmic methods are ineffective for large designs. For designs in contracting spaces, systematic construction methods have not been investigated yet. This paper proposes a new method for constructing SLHDs via good lattice point sets in various experimental spaces. These designs are called sequential good lattice point (SGLP) sets. Moreover, we provide efficient approaches for identifying the (nearly) optimal SGLP sets under a given criterion. Combining the linear level permutation technique, we obtain a class of asymptotically optimal SLHDs in invariant spaces, where the L1-distance in each stage is either optimal or asymptotically optimal. Numerical results demonstrate that the SGLP set has a better space-filling property than the existing SLHDs in invariant spaces. It is also shown that SGLP sets have less computational complexity and more adaptability.
and Xi for this thoughtful, revolutionary