When COVID-19 vaccines were introduced in late 2020 and widely distributed in early 2021, states were responsible for collecting and managing the data. In the best situation, states kept accurate records of each person who received the vaccine, including the age and the county of residence. States reported the cumulative number of those vaccinated in each county, although there were substantial numbers of vaccine recipients (within a given state) whose county of residence was unknown. Some states have very low numbers of vaccine recipients with unknown county, while other states reported upwards of 50% "unknown county of residence." At the extreme, Texas did not report the county of residence until October 2021, although they did report the state-wide total. There were a number of states that reported a nearly simultaneous jump in the cumulative number of those vaccinated whose county of residence was known and a drop in the number of "unknowns," likely caused by a retrospective analysis and reallocation of those whose county of residence was unknown. A further problem occurs when the cumulative number of vaccines drops. We describe how we created a database for county-level vaccine data that addresses these data quality issues.
ABSTRACT The moving average control chart has been used to detect changes in a process mean more quickly than the Shewhart chart. The MA chart takes an unweighted average of the most recent values and uses this to make an inference regarding the stability of the process. The double moving average chart, which applies the moving average to the already‐computed moving average, has recently been proposed. Triple and even quadruple moving average charts have also been proposed. We determine the optimal MA charts for these compound control charts and find that the optimal double or triple moving average is a single moving average chart. These higher‐order moving average charts provide no benefit over the single moving average chart.
Process monitoring involves applying control charts to infer whether the process is stable, or in control, at a given time. There are many control charting procedures available for practitioners. These procedures are often compared by looking at the average time to raise a signal. It is then possible to select an optimal chart for any given shift. Recently, however, focus has shifted from considering the average run length (ARL) for a given shift to the average or integrated ARL across an interval of shifts. This involves approximating an integral. For most monitoring techniques, the ARL cannot be determined analytically; rather, Monte Carlo simulation is needed to obtain an approximation. We consider three numerical techniques for approximating this integral: (1) Gaussian quadrature with simulation required to approximate the ARL for a shift, (2) a purely Monte Carlo estimate obtained by recognizing that the integral is an expectation, and (3) an adaptive method that allocates the simulation effort across the nodes in Gaussian quadrature. The adaptive quadrature method seems to produce the greatest accuracy. The purely Monte Carlo method, which is often used in the literature, is the least accurate.
The generally weighted moving average (GWMA) control chart has been proposed as a generalization and alternative to the exponentially weighted moving average (EWMA) chart. Proponents of the GWMA claim that it is more efficient in detecting shifts in the process mean than the EWMA. Most research on the GWMA chart compares its out-of-control properties against an EWMA chart that is not appropriate for the given situation. Detractors of the GWMA chart point out that (1) the GWMA is usually compared against an EWMA chart that is not ideal for the given situation, (2) the GWMA has no recursive formula so all previous data values must be stored and used to compute the next GWMA chart, and (3) the GWMA can have weights that do not decrease as the age of the data increases. We compare this optimal GWMA chart against the EWMA chart that is optimal for the same shift. For a given shift, a GWMA chart can be constructed that has a shorter cyclic steady-state average run length than the optimal EWMA chart. The optimal GWMA chart will usually have poor out-of-control performance for shifts other than those for which the GWMA was designed to be optimal. The GWMA chart is the preferred chart only in very specific circumstances.
We investigate Metropolis–Hastings (MH) algorithms to approximate the distribution of independent binomial random variables conditioned on the sum. Let Xi∼BIN(ni,pi). We want the distribution of [X1,…,Xk] conditioned on X1+⋯+Xk=n. We propose both a random walk MH algorithm and an independence sampling MH algorithm for simulating from this conditional distribution. The acceptance probability in the MH algorithm always involves the probability mass function of the proposal distribution. For the random walk MH algorithm, we take this distribution to be uniform across all possible proposals. There is an inherent asymmetry; the number of moves from one state to another is not in general equal to the number of moves from the other state to the one. This requires a careful counting of the number of possible moves out of each possible state. The independence sampler proposes a move based on the Poisson approximation to the binomial. While in general, random walk MH algorithms tend to outperform independence samplers, we find that in this case the independence sampler is more efficient.
The exponentially weighted moving (EWMA) control chart has been used successfully to monitor a process mean. Recently, extensions of the EWMA have been proposed that use an EWMA chart on the usual EWMA statistic; this is called the double EWMA, or DEWMA, chart. This effectively “double smooths” the original data. A multivariate version of the DEWMA chart, called the MDEWMA, is a straightforward extension of the DEWMA chart. The process output vectors are doubly smoothed, and a T^2 statistic is computed and plotted. Sufficiently large values of the T^2 statistic indicate a process shift. We compare the MEWMA chart with the newly proposed MDEWMA chart using the criterion of first-to-signal or “firstness.” For a given stream of output data, the chart that is most likely to signal first is considered the better chart. We estimate the firstness for the MEWMA and MDEWMA charts, along with the probability of a simultaneous signal. These firstness curves are estimated using simulation. We find that in most cases, the MEWMA chart signals first, or there is a simultaneous signal. Only rarely does the MDEWMA signal before the MEWMA.
Control chart performance is often measured using average run length or median run length, which gives the expected or median number of samples to signal. It is often argued that on average, one chart will signal a process change quicker than another, and is therefore a better choice. Average and median run length do not, however, answer the question of which method will be more likely to signal first. We introduce the idea of "first to signal" and compare charts based on this criterion.
With more and more data related to driving, traffic, and road conditions becoming available, there has been renewed interest in predictive modeling of traffic incident risk and corresponding risk factors. New machine learning approaches in particular have recently been proposed, with the goal of forecasting the occurrence of either actual incidents or their surrogates, or estimating driving risk over specific time intervals, road segments, or both. At the same time, as evidenced by our review, prescriptive modeling literature (e.g., routing or truck scheduling) has yet to capitalize on these advancements. Indeed, research into risk-aware modeling for driving is almost entirely focused on hazardous materials transportation (with a very distinct risk profile) and frequently assumes a fixed incident risk per mile driven. We propose a framework for developing data-driven prescriptive optimization models with risk criteria for traditional trucking applications. This approach is combined with a recently developed machine learning model to predict driving risk over a medium-term time horizon (the next 20 min to an hour of driving), resulting in a biobjective shortest path problem. We further propose a solution approach based on the k-shortest path algorithm and illustrate how this can be employed.
Abstract The double exponentially weighted moving average (DEWMA) control chart has been proposed as an alternative to the usual EWMA chart. The DEWMA is a compound chart, in the sense that the output of one charting procedure (the EWMA) is used as input to another charting procedure (again, the EWMA). Triple, and even quadruple, EWMA charts have been proposed. The DEWMA can be written as a weighted average of all previous data and the initial chart statistic. These weights can have the counterintuitive property that they are not monotonically decreasing as the age of the data value increases. In addition, they require additional computations compared to the EWMA. It has been found that the optimal EWMA chart is as good, or nearly as good, as the optimal DEWMA chart. When the DEWMA chart does outperform the EWMA, in the sense of shorter average run lengths for a given process shift, the DEWMA usually performs poorly for shifts other than the one for which it was optimized.
All animals are equal, but some animals are more equal than others. —George Orwell
Applied Stochastic Models in Business and IndustryEarly View COMMENTARY Discussion of “Specifying prior distributions in reliability applications,” by Qinglong Tian, Colin Lewis-Beck, Jarad B. Niemi, and William Meeker Necip Doganaksoy, Corresponding Author Necip Doganaksoy [email protected] orcid.org/0000-0002-8641-0814 Siena College, Loudonville, New York, USA Correspondence Necip Doganaksoy, Siena College, Loudonville, NY, USA. Email: [email protected]Search for more papers by this authorSteven E. Rigdon, Steven E. Rigdon Saint Louis University, Saint Louis, Missouri, USASearch for more papers by this author Necip Doganaksoy, Corresponding Author Necip Doganaksoy [email protected] orcid.org/0000-0002-8641-0814 Siena College, Loudonville, New York, USA Correspondence Necip Doganaksoy, Siena College, Loudonville, NY, USA. Email: [email protected]Search for more papers by this authorSteven E. Rigdon, Steven E. Rigdon Saint Louis University, Saint Louis, Missouri, USASearch for more papers by this author First published: 30 June 2023 https://doi.org/10.1002/asmb.2796Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinkedInRedditWechat Open Research DATA AVAILABILITY STATEMENT Data sharing is not applicable to this article as no new data were created or analyzed in this study. REFERENCES 1Nelson WB. Applied Life Data Analysis. Paperback ed. John Wiley & Sons; 2003. 2Nelson WB. Accelerated Testing: Statistical Models, Test Plans, and Data Analysis. Paperback ed. John Wiley & Sons; 2004. 3Tian Q, Lewis-Beck C, Niemi JB, Meeker WQ. Specifying prior distributions in reliability applications. Appl Stochastic Models Bus Ind. 2023. doi:10.1002/asmb.2752 4 Reliasoft Weibull ++. https://help.reliasoft.com/weibull20/bayesian_weibull_analysis.htm 5 Rogers Commission. Presidential Commission on the Space Shuttle Challenger Accident Report. Vol 1 & 2. Rogers Commission; 1986. 6Rigdon SE, Pan R, Montgomery DC, Freeman L. Design of Experiments for Reliability Achievement. Vol 1. John Wiley & Sons; 2022. Early ViewOnline Version of Record before inclusion in an issue ReferencesRelatedInformation
The double EWMA (DEWMA) has been proposed as a more efficient control charting procedure for monitoring the mean of a process. Comparisons of the DEWMA and the EWMA charts, which often indicate the superiority of the DEWMA, are often flawed because the same smoothing constant is used in both charts. We take the approach of first selecting a shift that we would like to detect, and then compare the optimal DEWMA chart and the optimal EWMA chart for that particular shift. We consider the DEWMA chart whose smoothing constants are restricted to be the same and the general DEWMA. We find that there are situations where the optimal DEWMA outperforms, in the sense of a shorter out-of-control average run length (ARL) for a fixed in-control ARL, but the improvement is slight. The optimal EWMA chart usually performs much better than the optimal DEWMA chart when the actual shift differs from the shift used to optimize the chart. The poor performance of the DEWMA chart away from the shift for which it was optimized, the nonmonotonicity of the DEWMA weights, and the additional computations required of the DEWMA chart indicate that the EWMA is a better overall choice than the DEWMA chart.
Introduction to Probability and Statistics for Data Science provides a solid course in the fundamental concepts, methods and theory of statistics for students in statistics, data science, biostatistics, engineering, and physical science programs. It teaches students to understand, use, and build on modern statistical techniques for complex problems. The authors develop the methods from both an intuitive and mathematical angle, illustrating with simple examples how and why the methods work. More complicated examples, many of which incorporate data and code in R, show how the method is used in practice. Through this guidance, students get the big picture about how statistics works and can be applied. This text covers more modern topics such as regression trees, large scale hypothesis testing, bootstrapping, MCMC, time series, and fewer theoretical topics like the Cramer-Rao lower bound and the Rao-Blackwell theorem. It features more than 250 high-quality figures, 180 of which involve actual data. Data and R are code available on our website so that students can reproduce the examples and do hands-on exercises.
Abstract Experiments for reliability usually involve lifetimes. Because lifetimes are usually not normally distributed, and because most life tests involve censoring, designing life tests for experiments is somewhat different from the usual normal‐theory designs. We present a review of regression models for life testing experiments, provide some examples, and discuss the optimal design of life tests.
The assumption of normality is usually tied to the design and analysis of an experimental study. However, when dealing with lifetime testing and censoring at fixed time intervals, we can no longer assume that the outcomes will be normally distributed. This generally requires the use of optimal design techniques to construct the test plan for specific distribution of interest. Optimal designs in this situation depend on the parameters of the distribution, which are generally unknown a priori. A Bayesian approach can be used by placing a prior distribution on the parameters, thereby leading to an appropriate selection of experimental design. This, along with the model and number of predictors, can be used to derive the D-optimal design for an allowed number of experimental runs. This paper explores using this Bayesian approach on various lifetime regression models to select appropriate D-optimal designs in regular and irregular design regions.
Background Socially vulnerable communities are at increased risk for adverse health outcomes during a pandemic. Although this association has been established for H1N1, Middle East respiratory syndrome (MERS), and COVID-19 outbreaks, understanding the factors influencing the outbreak pattern for different communities remains limited. Objective Our 3 objectives are to determine how many distinct clusters of time series there are for COVID-19 deaths in 3108 contiguous counties in the United States, how the clusters are geographically distributed, and what factors influence the probability of cluster membership. Methods We proposed a 2-stage data analytic framework that can account for different levels of temporal aggregation for the pandemic outcomes and community-level predictors. Specifically, we used time-series clustering to identify clusters with similar outcome patterns for the 3108 contiguous US counties. Multinomial logistic regression was used to explain the relationship between community-level predictors and cluster assignment. We analyzed county-level confirmed COVID-19 deaths from Sunday, March 1, 2020, to Saturday, February 27, 2021. Results Four distinct patterns of deaths were observed across the contiguous US counties. The multinomial regression model correctly classified 1904 (61.25%) of the counties’ outbreak patterns/clusters. Conclusions Our results provide evidence that county-level patterns of COVID-19 deaths are different and can be explained in part by social and political predictors.
Principal components have been used in conjunction with Hotelling's T-2 chart to monitor multivariate processes. It is known that prohibitively large sample sizes are needed to estimate the process parameters with enough precision to deploy the chart. We investigate whether principal components can be used to reduce the dimensionality of a process so that multivariate process control can be performed using estimated parameters. The chart based on the first k principal components, which we will refer to as the TPC,k2$T<^>2_{{\rm PC},k}$ chart, is investigated. Specifically, we explore three research questions in this paper: (1) can the TPC,k2$T<^>2_{{\rm PC},k}$ chart with estimated parameters be applied with moderate preliminary sample sizes?, (2) are there situations where the TPC,k2$T<^>2_{{\rm PC},k}$ charts are able to detect a shift in the mean vector quicker than the T-2 charts?, and (3) is it possible to exploit assumptions about the covariance matrix, such as equal covariances, to improve the performance of the TPC,k2$T<^>2_{{\rm PC},k}$ chart. Using simulation, we find that for high dimensions, the TPC,k2$T<^>2_{{\rm PC},k}$ chart with estimated parameters can be used to detect shifts in the direction of the first principal component. Otherwise, the chart requires very large preliminary sample sizes. When it is reasonable to assume equal covariances, the number of parameters to be estimated is substantially reduced and the performance of the TPC,k2$T<^>2_{{\rm PC},k}$ chart is improved.