Finding key drivers in regression modeling via Bayesian Sensitivity-Specificity and Receiver Operating Characteristic is suggested, and clearly interpretable results are obtained. Numerical comparisons with other techniques show that this methodology can be useful in practical statistical modeling and analysis helping to researchers and managers in making meaningful decisions.
Discrete choice modeling is one of the main tools of estimation utilities and preference probabilities among multiple alternatives in economics, psychology, social sciences, and marketing research. One of popular DCM tools is the Best-Worst Scaling, also known as Maximum Difference. Data for such modeling is given by respondents presented with several items, and each respondent chooses the best alternative. Estimation of utilities is usually performed in a multinomial-logit modeling which produces utilities and choice probabilities. This article describes how to obtain probability estimation adjusted to possible absence of items in actual purchasing. We apply Markov chain modeling in the form of Chapman-Kolmogorov equations and its steady-state solution for stochastic matrix can be obtained analytically. An adjustment to choice probability with network effects is also considered. Numerical example by marketing research data is used.
A description of Likert scales can be given using the multipoles technique known in quantum physics and applied to behavioral sciences data. This paper considers decomposition of Likert scales by the multipoles for the application of decreasing the respondents' heterogeneity. Due to cultural and language differences, different respondents habitually use the lower end, the mid-scale, or the upper end of the Likert scales which can lead to distortion and inconsistency in data across respondents. A big impact of different kinds of respondent is well known, for instance, in international studies, and it is called the problem of high and low raters. Application of a multipoles technique to the row-data smoothing via prediction of individual rates by the histogram of the Likert scale tiers produces better results than standard row-centering in data. A numerical example by marketing research data shows that the results are encouraging: while a standard row-centering produces a poor outcome, the dipole-adjustment noticeably improves the obtained segmentation results.
Maximum difference (MaxDiff) is a discrete choice modeling approach widely used in marketing research for finding utilities and preference probabilities among multiple alternatives. It can be seen as an extension of the paired comparison in Thurstone and Bradley–Terry techniques for the simultaneous presenting of three, four or more items to respondents. A respondent identifies the best and the worst ones, so the remaining are deemed intermediate by preference alternatives. Estimation of individual utilities is usually performed in a hierarchical Bayesian (HB)-multinomial-logit (MNL) modeling. MNL model can be reduced to a logit model by the data composed of two specially constructed design matrices of the prevalence from the best and the worst sides. The composed data can be of a large size which makes logistic modeling less precise and very consuming in computer time and memory. This paper describes how the results for utilities and choice probabilities can be obtained from the raw data, and instead of HB methods the empirical Bayes techniques can be applied. This approach enriches MaxDiff and is useful for estimations on large data sets. The results of analytical approach are compared with HB-MNL and several other techniques.
Predictor importance in applied regression modeling gives the main operational tools for managers and decision-makers. The paper considers estimation of predictors' importance in regression using measures introduced in works by Gibson and R. Johnson (GJ), then modified by Green, Carroll, and DeSarbo, and developed further by J. Johnson (JJ). These indices of importance are based on the orthonormal decomposition of the data matrix, and the work shows how to improve this approximation. Using predictor importance, the regression coefficients can also be adjusted to reach the best data fit and to be meaningful and interpretable. The results are compared with the robust to multicollinearity, but computationally difficult, Shapley value regression (SVR). They show that the JJ index is good for importance estimation, but the GJ index outperforms it if both predictor importance and coefficients of regression are needed; hence, this index (GJ) can be used in place of the more computationally intensive estimation by SVR. The results can be easily estimated by the considered approach that is very useful in practical regression modeling and analysis, especially for big data.
Best-Worst Scaling (BWS) modeling is widely used for finding probabilities of choice among multiple items. The paper considers how to apply BWS data to another problem – of finding the items׳ cannibalization and synergy. For a product of primary interest, we estimate its probability to be chosen as the best one out of all the data, and also conditionally to each product׳s presence or absence. For a given product, each other one behaves as a catalyzer or inhibitor of the choice. Constructing the entire matrix of such relations for all the products, we compare its symmetrical elements for each pair of products. It shows which pairs of products are mutually synergic, or complementary, so their chances to be chosen as the best ones are higher in the presence of each other. In other cases, the products can be of negative impact on one another, so one is a cannibalizer of another; or both products suppress each other. Estimations on real marketing data are considered, including the Shapley value for key driver analysis.
An AHP matrix of the quotients of the pair comparison priorities is transformed to a matrix of shares of the preferences which can be used in Markov stochastic modeling via the Chapman-Kolmogorov system of equations for the discrete states. It yields a general solution and the steady-state probabilities. The AHP priority vector can be interpreted as these probabilities belonging to the discrete states corresponding to the compared items. The results of stochastic modeling correspond to robust estimations of priority vectors not prone to influence of possible errors among the elements of a pairwise comparison matrix.
Best-Worst Scaling (BWS), sometimes also called Maximum Difference (MaxDiff), is a discrete choice modeling method widely used for finding utilities and choice probabilities among multiple alternatives. It can be seen as an extension of the paired comparison techniques for the simultaneous presentation of several items together to respondents. A respondent identifies the best and the worst ones and estimation of utilities is performed using a multinomial-logit (MNL) model in numerical nonlinear estimations. The main contribution of this paper consists in finding an analytical closed-form solution producing an approximation of the results for utilities and choice probabilities that are obtained using MNL models. The analytical formulae permit the inference of the characteristics of the model׳s quality, including standard errors of the utilities and choice probabilities, the residual deviance and pseudo-R2. This approach enriches the BWS methods and is useful for theoretical descriptions and practical applications.
Shapley value regression consists of assessing relative importance and accordingly adjusting regression coefficients. It is argued that adjustment of coefficients is unnecessary and even misleading for practically relevant situations. Examples are given, and an alternative procedure is proposed for situations for which the coefficients are requested to have a certain sign. Copyright © 2009 John Wiley & Sons, Ltd.
Decision makers in organizations frequently need to determine relative importance of multiple predictors of an important Dependent Variable. Using Multiple Regression for this purpose is often challenging when predictors are intercorrelated. In this
Applied Stochastic Models in Business and IndustryVolume 26, Issue 2 p. 203-204 Letter to the Editor Reply to the paper ‘Do not adjust coefficients in Shapley value regression’ by U. Gromping, S. Landau, Applied Stochastic Models in Business and Industry, 2009; DOI: 10.1002/asmb.773 Stan Lipovetsky, Stan Lipovetsky GfK Custom Research North America, 8401 Golden Valley Rd, Minneapolis, MN 55427, U.S.A.Search for more papers by this authorW. Michael Conklin, W. Michael Conklin MarketTools, Inc., 6465 Wayzata Blvd, Suite 170, St. Louis Park, MN 55426, U.S.A.Search for more papers by this author Stan Lipovetsky, Stan Lipovetsky GfK Custom Research North America, 8401 Golden Valley Rd, Minneapolis, MN 55427, U.S.A.Search for more papers by this authorW. Michael Conklin, W. Michael Conklin MarketTools, Inc., 6465 Wayzata Blvd, Suite 170, St. Louis Park, MN 55426, U.S.A.Search for more papers by this author First published: 12 April 2010 https://doi.org/10.1002/asmb.829Citations: 4AboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat No abstract is available for this article.Citing Literature Volume26, Issue2March/April 2010Pages 203-204 RelatedInformation
Coefficients of ordinary least squares (OLS) multiple regression present the best linear combination of the independent variables for fitting a dependent variable. However, the OLS model had never been designed to produce meaningful individual predic
Singular value decomposition (SVD) is widely used in data processing, reduction, and visualization. Applied to a positive matrix, the regular additive SVD by the first several dual vectors can yield irrelevant negative elements of the approximated matrix. We consider a multiplicative SVD modification that corresponds to minimizing the relative errors and produces always positive matrices at any approximation step. Another logistic SVD modification can be used for decomposition of the matrices of proportions, when a regular SVD can yield the elements beyond the zero-one range, while the modified SVD decomposition produces all the elements within the correct range at any step of approximation. Several additional modifications of matrix approximation are also considered.
We consider latent class regressions for the simultaneous construction of several regression models by the data clusters. Maximum likelihood objective of observations belonging to at least one data segment is developed. Solution is reduced to the iteratively reweighted least squares (IRLS) procedure that defines coefficients of all models and the characteristics of fitting. Together with the regression models, this approach yields probabilities of each observation belonging to each of the classes. This technique can also be used for finding parameters of mixed distributions. The suggested approach enriches results of the regression modeling and clustering in practical applications.
We consider simultaneous minimization of the model errors, deviations from orthogonality between regressors and errors, and deviations from other desired properties of the solution. This approach corresponds to a regularized objective that produces a consistent solution not prone to multicollinearity. We obtain a generalization of the ridge regression to two-parameter model that always outperforms a regular one-parameter ridge by better approximation, and has good properties of orthogonality between residuals and predicted values of the dependent variable. The results are very convenient for the analysis and interpretation of the regression. Numerical runs prove that this technique works very well. The examples are considered for marketing research problems. Copyright © 2005 John Wiley & Sons, Ltd.
A regular problem in regression analysis is estimating the comparative importance of the predictors in the model. This work considers the ‘net effects’, or shares of the predictors in the coefficient of the multiple determination, which is a widely used characteristic of the quality of a regression model. Estimation of ;the net effects can be a difficult task because multicollinearity among the regressors can produce negative inputs to multiple determination. This paper suggests estimating the incremental net effects as subsequent marginal inputs to the coefficient of multiple determination, and it is shown that the results coincide with estimation by cooperative game theory. This approach guarantees positive and interpretable net effects, which offers a better interpretation of the regression results.
We consider a problem of marketing decisions for the choice of a product with maximum customer appeal. A widely used technique for this purpose is TURF, or Total Unduplicated Reach and Frequency, which evaluates a union of the events defined by the sample proportion of many products, or flavors of one product. However, when using TURF, it is often impossible to distinguish between subsets of different flavor combinations with practically the same level of coverage. An appropriate tool can be borrowed from cooperative game theory, namely, the Shapley Value, that permits the ordering of flavors by their strength in achieving maximum consumers' reach and provides more stable results than TURF. We describe marketing strategy reasons for using these techniques in the identification of the preferred combinations in media or product mix.
contains a few good case studies directly related to quality engineering, many of the examples have nothing to do with industry (e.g., measurements on the shapes and sizes of painted turtle carapaces, weight of middle-aged men in fitness clubs, diagnosis of liver disease). These examples are used to illustrate the methodologies as they are being developed and make up the bulk of the results given in the book. On the bright side, the case studies are engaging. The data analysis objectives of the case studies are clearly described, and the presentations of results show how multivariate analysis can address the objectives. One could argue that the goal of a book describing application of multivariate methods to specific types of problems should aim to provide reads with a “feel” for how these methods work for their problems. For example, I have spent years encouraging microbiologists to use “industrial statistics,” such as response surface methods or control charts. I have found that once they see indepth case studies within their area of science that address the nuances of their data, they are almost completely on board. In this sense, the authors moderately succeed using their case studies. However, I would suggest that they make the data available to the public, because many people learn by reproducing examples that they study. To enable quality organizations to better use multivariate methods, this text should be supplemented with others. For example, the chapter on discrimination hits the major points for classical multivariate methods; however, a vast array of tools for discrimination currently exist, and the classical methods may not fit the problem well and may be far from optimal. For instance, many on-line measurement systems produce large amounts of data on each sample, such as optical coordinate measurement machines, so there may be more variables than observations. Classical methods cannot handle this situation well (e.g., LDA with more variables than observations). I suggest the text by Hastie, Tibshirani, and Friedman (2001) to supplement this book if more expertise is required for discrimination, clustering, and principal components analysis, and the text by Mason and Young (2001) for a more in-depth resource for multivariate control charting. All in all, the authors meet their intended goals somewhat. This text provides the motivation to use multivariate analysis, but could do a better job providing the means.
We consider simultaneous minimization of the model errors, deviations from orthogonality between regressors and errors, and deviations from other desired properties of the solution. This approach corresponds to a regularized objective that produces a consistent solution not prone to multicollinearity. We obtain a generalization of the ridge regression to two-parameter model that always outperforms a regular one-parameter ridge by better approximation, and has good properties of orthogonality between residuals and predicted values of the dependent variable. The results are very convenient for the analysis and interpretation of the regression. Numerical runs prove that this technique works very well. The examples are considered for marketing research problems. Copyright (c) 2005 John Wiley & Sons, Ltd.
It is known that two-group linear discriminant function can be constructed via binary regression. In this article, it is shown that the opposite relation is also relevant - it is possible to present multiple regression as a linear combination of a main part, based on the pooled variance, and Fisher discriminators by data segments. Presenting regression as an aggregate of the discriminators allows one to decompose coefficients of the model into sum of several vectors related to segments. Using this technique provides an understanding of how the total regression model is composed of the regressions by the segments with possible opposite directions of the dependency on the predictors.