A two level orthogonal array design for n observations with k factors and of projectivity P provides an (n, k, P) factor screen for which every projection into P space produces a complete 2 P factorial, possibly with certain points replicated. Box and Tyssedal [1] rigorously investigated the screening properties of such designs derived from fractional factorials and Plackett Burman orthogonal arrays (OA's). For example, they showed that these always provided (12, 11, 3) and (20, 19, 3) screens, but for 16 runs only a (16, 8, 3) screen could be generated. In this paper it is shown that designs derived from a different class of OA's due to Hall [2] can produce sixteen run designs that can screen a larger number of factors and that in particular (16, 12, 3) screens and also a (16, 14, 3) screen can be obtained.
Industrial success requires efficient experimentation both for the improvement of existing products and processes and for development of new ones. Because results are usually known quickly, the natural way to experiment is to use information from each group of runs to plan the next. Such investigation employs a scientific paradigm in which data drives an alternation of induction and deduction. This process can suggest at each stage how questions that are still at issue can be resolved. Response surface methods are a group of statistical techniques specifically designed to catalyze scientific learning of this kind. In this paper, the scientific paradigm for discovery and sequential learning is contrasted with the mathematical paradigm for the proof of theorems. It is argued that, because statistical training unduly emphasizes mathematics at the expense of science, confusion between the two paradigms occurs. This has resulted in emphasis on the development and use of "one-shot" statistical procedures which mimic the mathematical paradigm-examples are hypothesis testing and the use of alphabetically optimal designs. Such one-shot procedures, where the model is assumed known a priori and fixed, are appropriate for some practical problems and are attractive because they allow rigorous development of theories of statistics based on mathematics alone. By contrast, discovery of new knowledge requires the use of the scientific paradigm in which the model is continually changing. Scientific method is thus mathematically incoherent. The importance of robustness is discussed both for analysis and design, and the relationship between these two kinds of robustness is clarified. Implications for teaching are discussed.
Innovation in the design and manufacture of processes and products usually comes about as a result of careful investigation—a directed process of sequential learning. Many practitioners, although familiar with “one-shot” statistical procedures, have little knowledge of the power of statistical techniques designed to catalyze investigation itself. A simple means of demonstrating and experiencing this learning process is illustrated using response surface methods to find an improved design for a paper helicopter.
Criteria are derived for objective choice among rival models of multiresponse processes. Formulas are given for the relative posterior probabilities of candidate models and for their goodness of fit as tested on a common data set with Normally distributed errors. The formulas are demonstrated with examples from chemical kinetics and catalysis.
It is argued that the domination of Statistics by Mathematics rather than by Science has greatly reduced the value and the status of the subject. The mathematical "theorem-proof paradigm" has supplanted the "iterative learning paradigm" of scientific method. This misunderstanding has affected university teaching, research, the granting of tenure to faculty and the distributions of grants by funding agencies. Possible ways in which some of these problems might be overcome and the role that computers can play in this reformation are discussed.
Click to increase image sizeClick to decrease image size Additional informationNotes on contributorsGeorge E. P. BoxDr. Box is Professor Emeritus and Director of Research. He is a Fellow of ASQC.David E. ColemanMr. Coleman is a Senior Technical Specialist in Applied Math & Computer Technology. He is a Member of ASQC.Robert V. BaxleyMr. Baxley is a Fellow in the Fibers Business Unit. He is a Senior Member of ASQC.
The role of statistics in quality and productivity improvement depends on certain philosophical issues that the author believes have been inadequately addressed. Three such issues are as follows: (1) what is the role of statistics in the process of investigation and discovery; (2) how can we extrapolate results from the particular to the general; and (3) how can we evaluate possible management changes so that they truly benefit an organization Therefore, statistical methods appropriate to investigation and discovery are discussed as distinct from those appropriate to the testing of an already discovered solution. It is shown how the manner in which the tentative solution has been arrived at determines the assurance with which experimental conclusions can be extrapolated to the application in mind. Whether or not statistical methods and training can have any impact depends on the system of management. A vector representation which can help predict the consequences of changes in management strategy is discussed. This can help to realign policies so that members of an organization can better work together for the benefit of the organization.
The inverse probability theorem of Bayes is used, along with sampling theory, to obtain objective criteria for choosing among rival models. Formulas are given for the relative posterior probabilities of candidate models and for their goodness of fit, when the models are fitted to a common data set with Normally distributed errors. Cases of full, partial and minimal variance information are treated. The formulas are demonstrated with three examples, including a kinetic study of a catalytic reaction.
Highly fractionated factorial designs and other orthogonal arrays are powerful tools for identifying important, or active, factors and improving quality. We show, however, that interactions, and important factors involved in those interactions, may go unidentified when conventional methods of analysis are used with these designs. This is particularly true of Plackett-Burman designs where the number of runs is not a power of two. A Bayesian method that allows for the possibility of interactions is developed to compute the marginal posterior probability that a factor is active. The method can be applied to both orthogonal and nonorthogonal designs, as well as other troublesome situations, such as when data are missing, extra data are available, or factor settings for certain runs have deviated from those originally planned. The value of the new technique is demonstrated with three examples in which potential interactions and factors involved in those interactions are uncovered.
SUMMARY It is often necessary to adjust some variable X, such as the concentration of consecutive batches of a product, to keep X close to a specified target value. A second more complicated problem occurs when the independent variables X in a response function η(X) are to be adjusted so that the derivatives ∂η/∂X are kept close to a target value zero, thus maximizing or minimizing and the paper is devoted mainly to the estimation from past data of the “best” adjustments to be applied in the first problem.
In this chapter we analyse two sets of data using some of the methodology described in this book. Practical experience is important in the development of new statistical tools. Through such experience, strengths and weaknesses of methodology are exposed, and important directions for further research and development are discovered. Both for our sake and the reader’s, the examples in this chapter and other chapters in the book were not carefully selected to show off the tools described in this book. Rather, we required examples that show honestly the strengths of additive modelling as well the inherent difficulties.
This article studies how to identify hidden factors in multivariate time series process. This problem is important because, when the series are driven by a set of common factors, (a) a large number of parameters may be needed to obtain an adequate representation of the system and (b) the estimated parameters will be highly correlated. Therefore, a complex and badly defined relationship can appear when, in fact, a simpler and parsimonious model in terms of a few common factors can be operating. This article develops a methodology to identify the number of factors and to build a simplifying transformation to represent the series. It is proved that the number of factors is equal to the rank of the covariance matrices and the parameter matrices of the infinite moving average representation of the process. The eigenvectors of these matrices will provide the canonical transformation. The method is illustrated with one example, using series of the price of wheat in five provinces of Spain in the 19th century. The standard approach to build a vector autore-gressive integrated moving average model showed a complex relationship with all kinds of feedback operating. When the methodology developed in the article was applied, however, two factors were identified and a clearer and simpler representation of the system was achieved.
We consider the problem of trend estimation from an ARIMA-model-based perspective. Given that the time series of interest is well represented by a multiplicative seasonal ARIMA model (Box and Jenkins 1970), a method is presented for estimating the trend of this series as a component of the model's forecast function. This method is applied to the “airline model,” a commonly occurring model form for economic and social time series. For those numerous series that are rendered stationary by logging and first-differencing, the resulting quantity supplies an estimate of the rate of change, or growth rate, of the series. As this trend estimate is given in terms of the underlying series' forecast function, we analyze forecast functions of ARIMA models in general and of the airline model in particular, extending the development in Box and Jenkins (1970, chaps. 5 and 9). Representing the forecast function as the solution of a difference equation, the trend estimate is the component of the forecast function that is a linear combination of those roots of the observable series' autoregressive operator that are associated with trend (an allocation that is usually standard and unambigous in practice). Thus it is important to determine the coefficients in this linear combination, which are adaptive in the time origin of the forecast, and we describe three general methods for doing this: in terms of initial forecasts, directly from the present and recent values of the observed series, and recursively. For the airline model the trend estimate is a straight line, with both intercept and slope changing in each time period, adapting to the new information becoming available. We also connect this trend estimate with trend forecasts in unobserved-components ARIMA models, which underlie much seasonal adjustment research of the last decade. Trend components in such models are not unique, even given the association of certain of the autoregressive roots with the trend. We show, however, that all admissible trend-seasonal-irregular decompositions have identical trend-component forecasts and that this common forecast is supplied by our trend estimate. Thus the “true trend” (that which our trend estimate estimates) is given only with reference to a particular decomposition. One such decomposition, commonly used in model-based seasonal adjustment, is the one (whose existence and uniqueness are known under fairly general conditions) that maximizes the irregular-component variance or, equivalently, minimizes the trend- and seasonal-component innovation variances. We show that, among all admissible decompositions, this “canonical decomposition” minimizes the mean squared error of our trend estimate.
A distinguishing feature of Japanese quality improvement techniques is an emphasis on the designing of quality into the product and into the process that makes the product. In particular, experimental design is used to discover conditions that minimize variance and appropriately control the mean level. The direct estimation of variance by replication at each of the design points, however, can be excessively expensive in experimental runs. In this article we show how it is sometimes possible to use unreplicated fractional designs to identify factors that affect variance in addition to those that affect the mean.
Loss of markets to Japan has recently caused attention to return to the enormous potential that experimental design possesses for the improvement of product design, for the improvement of the manufacturing process, and hence for improvement of overall product quality. In the screening stage of industrial experimentation it is frequently true that the “Pareto Principle” applies; that is, a large proportion of process variation is associated with a small proportion of the process variables. In such circumstances of “factor sparsity,” unreplicated fractional designs and other orthogonal arrays have frequently been effective when used as a screen for isolating preponderant factors. A useful graphical analysis due to Daniel (1959) employs normal probability plotting. A more formal analysis is presented here, which may be used to supplement such plots and hence to facilitate the use of these unreplicated experimental arrangements.
Fractional factorial designs have been used successfullyin industry and elsewhere todetect and estimate sparse factor effects , The effectsusually evisioned measure changes in location associated with the experimental factors . Here we consider the possibility of detecting and estimating sparse dispersion effects measuring changes in variance associated with the factors . ( 2 ) In industrial experimentation it is frequently true thata large proportion of process v ariation is associated with a smallproportion of the process variables . In such circum stancs of“effect sparsity”unreplicated fractional designs have frequently been effectivein islolating preponderant factors.A very useful graphical analysis for such experiments due to Cuthbert Daniel(1959)employs normal probability plotting.A more formal analysis is presented here which might be used to supplement such plots.