Given the widely available online customer ratings on products, the individual-level rating prediction and clustering of customers and products are increasingly important for sellers to create targeting strategies for expanding the customer base and improving product ratings. However, the massive missing data problem is a significant challenge for modeling online product ratings. To address this issue, we propose a new co-clustering methodology based on a bipartite network modeling of large-scale ordinal product ratings. Our method extends existing co-clustering methods by incorporating covariates and ordinal ratings in the model-based co-clustering of a weighted bipartite network. We devise an efficient variational EM algorithm for model estimation. A simulation study demonstrates that our methodology is scalable for modeling large datasets and provides accurate estimation and clustering results. We further show that our model can successfully identify different groups of customers and products with meaningful interpretations and achieve promising predictive performance in a real application for customer targeting.
Purpose This paper aims to examine the evolution of a competitive market structure over time through the lens of competitive group membership dynamics. Design/methodology/approach A new hidden Markov modeling approach is devised that accounts for the three sources of competitive heterogeneity involving managerial strategy, corporate performance and the impact of strategy on performance. In addition, some observed "entry" and "exit" states are considered to model firms' entry into and exit from the market. The proposed model is illustrated with an investigation of the US banking industry based on a data set created from the COMPUSTAT database. This paper estimated the model within the Bayesian framework and devised a reversible jump Markov chain Monte Carlo estimation procedure to determine the number of latent competitive groups and uncover the characteristics of each group. Findings This paper shows that the US banking industry, contrary to the prior findings of having a relatively stable structure, has, in fact, gone through dramatic changes in the past number of decades. Originality/value Contrary to prior work that has primarily focused on managerial strategy to study market evolutions, the competitive groups perspective accounts for all three sources of intra-industry competitive heterogeneity. In addition, unlike prior research, the analysis is not limited to firms remaining in the panel of study for the entire observation period. Such limitation results in missing the various changes that occur in the competitive market structure because of the new entrants or the struggling firms that do not survive in the market.
We propose a new spatial modeling approach to calibrate the potential impact of spatial dependency and heterogeneity on the underlying drivers of customer service and/or satisfaction measurement. The newly proposed procedure derives regionally varying coefficients, provides more flexible fitting, improves calibration fit and predictive validation, and can potentially result in augmented managerial implications compared to existing procedures by utilizing a hierarchical Bayes framework with geographical boundary effects. Using synthetic datasets, we illustrate how the proposed model outperforms four relevant benchmark models including ordinary linear regression, a Spatially Dependent Segmentation model (Govind, Rabikar, and Mittal 2018), classic Geographically Weighted Regression, and Bayesian Geographically Weighted Regression. The improved performance is most prominent when there exist significant differences between geographic boundaries and/or irregular patterns of observation locations. In our automobile customer satisfaction application study, the proposed approach also demonstrates favorable performance compared to these benchmark models. We find a dramatically heterogeneous pattern regarding two covariates in the Mountain U.S. geographic division: dealership service is more important in urban areas (e.g., Phoenix, Salt Lake City and Denver) than in rural areas, but vice-versa concerning vehicle quality.
Consumer dispersion analysis divides aggregate markets into smaller geographic units that marketers can target with their promotional mix. However, dispersion patterns are not always contiguous. Using survey data from National Football League (NFL) fans, we introduce a new hierarchical expectation-maximization (EM) bi-level clustering model that iteratively classifies both teams and fans (nested within teams) based on the spatial heterogeneity of fans in terms of both distance and direction. The proposed multi-level latent class model with a variable number of classes at the lower level outperforms benchmark models in a Monte Carlo simulation study and points to three non-contiguous team segments with a varying number of fan group vectors in the NFL application. We present these results in two-dimensional consumer dispersion maps and report corresponding differences in consumer behavior.
Graphical models have received an increasing amount of attention in network psychometrics as a promising probabilistic approach to study the conditional relations among variables using graph theory. Despite recent advances, existing methods on graphical models usually assume a homogeneous population and focus on binary or continuous variables. However, ordinal variables are very popular in many areas of psychological science, and the population often consists of several different groups based on the heterogeneity in ordinal data. Driven by these needs, we introduce the finite mixture of ordinal graphical models to effectively study the heterogeneous conditional dependence relationships of ordinal data. We develop a penalized likelihood approach for model estimation, and design a generalized expectation-maximization (EM) algorithm to solve the significant computational challenges. We examine the performance of the proposed method and algorithm in simulation studies. Moreover, we demonstrate the potential usefulness of the proposed method in psychological science through a real application concerning the interests and attitudes related to fan avidity for students in a large public university in the United States.
Model-based market segmentation analyses often involve an ordinal dependent variable as ordinal responses are frequently collected in marketing research. In the Bayesian segmentation literature, there are models for an interval- or ratio-scaled dependent variable but there is not any general model for an ordinal dependent variable. In this manuscript, the authors propose a new Bayesian procedure to simultaneously perform segmentation and ordinal regression with variable selection within each derived segment. The procedure is robust to outliers and it also provides an option to include concomitant variables that allows the simultaneous profiling of the derived segments. The authors demonstrate that the practice of treating ordinal responses as interval- or ratio-scales to apply existing Bayesian segmentation procedures can lead to very misleading results and conclusions. Through simulation studies, the authors show that the proposed procedure outperforms several benchmark Bayesian segmentation models in parameter recovery, segment retention, and segment membership prediction for such data. Finally, they provide a commercial business customer satisfaction empirical application to illustrate the usefulness of the proposed model.
The field of marketing has made significant strides over the past 50 years in understanding how methodological choices affect the validity of conclusions drawn from our research. This paper highlights some of these and is organized as follows: We first summarize essential concepts about measurement and the role of cumulating knowledge, then highlight data and analysis methods in terms of their past, present, and future. Lastly, we provide specific examples of the evolution of work on segmentation and brand equity. With relatively well-established methods for measuring constructs, analysis methods have evolved substantially. There have been significant changes in what is seen as the best way to analyze individual studies as well as accumulate knowledge across them via meta-analysis. Collaborations between academia and business can move marketing research forward. These will require the tradeoffs between model prediction and interpretation, and a balance between large-scale use of data and privacy concerns.
Assessing market structure by deriving a brand positioning map and segmenting customers is essential for supporting brand-related marketing decisions. We propose adaptive multidimensional scaling (ADMDS) for simultaneously deriving a brand positioning map and market segments using customer data on cognitive decision sets and brand dissimilarities. In ADMDS, the judgment task is adapted to the individual customer where dissimilarity judgments are collected only for those brands within a customers’ awareness set. Thus, respondent fatigue and unfamiliarity with the brands are circumvented thereby improving the validity of the dissimilarity data obtained, as well as the multidimensional spatial structure derived from them. Estimation of the ADMDS model results in a spatial map in which the brands and derived segments of customers are jointly represented as points. The closer a brand is positioned to a segment’s ideal brand, the higher the probability that the brand is considered and chosen. An assumption underlying this model representation is that brands within a customers’ consideration set are relatively similar. In an experiment with 200 respondents and 4 product categories, this assumption is validated. We illustrate adaptive multidimensional scaling model on commercial data for 20 midsize car brands evaluated by 212 members of an on-line consumer panel. Potential applications of the method and future research opportunities are discussed.
While the smooth transition (ST) model has become popular in business and economics, the treatment of unobserved heterogeneity within these models has received limited attention. We propose a ST finite mixture (STFM) model which simultaneously estimates the presence of time-varying effects and unobserved heterogeneity in a panel data context. Our objective is to accurately recover the heterogeneous effects of our independent variables of interest while simultaneously allowing these effects to vary over time. Accomplishing this objective may provide valuable insights for managers and policy makers. The STFM model nests several well-known ST and threshold models. We develop the specification, estimation, and model selection criteria for the STFM model using Bayesian methods. We also provide a theoretical assessment of the flexibility of the STFM model when the number of regimes grows with the sample size. In an extensive simulation study, we show that ignoring unobserved heterogeneity can lead to distorted parameter estimates, and that the STFM model is fairly robust when underlying model assumptions are violated. Empirically, we estimate the effects of in-game promotions on game attendance in Major League Baseball. Empirical results show that the STFM model outperforms all its nested versions.for this article are available online.
Consumers’ preferences for various product attributes change over time. Modeling such temporal changes through a single process assumes that all the attributes’ preferences change together with the same dynamics; however, this assumption is not appropriate when there are several processes with distinct characteristics. We propose a new non-homogeneous factorial hidden Markov model (FHMM) for choice models to dynamically segment consumers into distinct states while each preference parameter may follow a distinct Markov process. The transition probabilities are modeled as time-varying at the individual level, affected by covariates of a feedback term of the consumer’s previous purchase decision, specific to each Markov process. We motivate the proposed approach by an application to a scanner panel choice dataset and find two processes with entirely different characteristics governing the shifts in two preference attributes. Model fit and prediction power based on Brier scores show the superiority of the proposed non-homogeneous FHMM in capturing temporal changes in preferences compared to a traditional hidden Markov model as well as a benchmark comparison model.
There is a vast behavioral decision theory literature that suggests different individuals may utilize and/or weigh different attributes of an object to form the basis of their opinions, attitudes, choices, and/or evaluations of such stimuli. This heterogeneity of information utilization and importance can be due to several different factors such as differing goals, level of expertise, contextual factors, knowledge accessibility, time pressure, involvement, mood states, task complexity, communication or influence of relevant others, etc. This phenomenon is particularly pertinent to the evaluation of stimuli involving large numbers of underlying attributes or features. We propose a new hierarchical Bayesian multivariate probit mixture model with variable selection accommodating such forms of choice heterogeneity. Based on a Monte Carlo simulation study, we demonstrate that the proposed model can successfully recover true parameters in a robust manner. Next, we provide a consumer psychology application involving consideration to buy choices for intended consumers of large Sports Utility Vehicles. The application illustrates that the proposed model outperforms several comparison benchmark choice models with respect to face validity and choice predictive validation performance.
The hidden Markov model (HMM) provides a framework to model the time-varying effects of marketing mix variables. When employed in a panel data context, it is important to properly account for unobserved heterogeneity across individuals. We propose a new random coefficients mixture HMM (RCMHMM) that allows for flexible patterns of unobserved heterogeneity in both the state-dependent and transition parameters. The RCMHMM nests all HMMs found in the marketing literature. Results of two simulation studies demonstrate that 1) averaging across a large number of different data generating processes, the RCMHMM outperforms all its nested versions using both in-sample and out-of-sample performance and 2) the RCMHMM is more robust than its nested versions when underlying model assumptions are violated. In addition, we apply the RCMHMM to an empirical application where we examine the effectiveness of in-game promotions in increasing the short-term demand for Major League Baseball (MLB) attendance. We find that the effectiveness of four promotional categories varies over the course of the season and across teams and that the RCMHMM performs best.
Purpose - Joint space multidimensional scaling (MDS) maps are often utilized for positioning analyses and are estimated with survey data of consumer preferences, choices, considerations, intentions, etc. so as to provide a parsimonious spatial depiction of the competitive landscape. However, little attention has been given to the possibility that consumers may display heterogeneity in their information usage (Bettman et al., 1998) and the possible impact this may have on the corresponding estimated joint space maps. This paper aims to address this important issue and proposes a new Bayesian multidimensional unfolding model for the analysis of two or three-way dominance (e.g. preference) data. The authors' new MDS model explicitly accommodates dimension selection and preference heterogeneity simultaneously in a unified framework. Design/methodology/approach - This manuscript introduces a new Bayesian hierarchical spatial MDS model with accompanying Markov chain Monte Carlo algorithm for estimation that explicitly places constraints on a set of scale parameters in such a way as to model a consumer using or not using each latent dimension in forming his/her preferences while at the same time permitting consumers to differentially weigh each utilized latent dimension. In this manner, both preference heterogeneity and dimensionality selection heterogeneity are modeled simultaneously. Findings - The superiority of this model over existing spatial models is demonstrated in both the case of simulated data, where the structure of the data is known in advance, as well as in an empirical application/ illustration relating to the positioning of digital cameras. In the empirical application/illustration, the policy implications of accounting for the presence of dimensionality selection heterogeneity is shown to be derived from the Bayesian spatial analyses conducted. The results demonstrate that a model that incorporates dimensionality selection heterogeneity outperforms models that cannot recognize that consumers may be selective in the product information that they choose to process. Such results also show that a marketing manager may encounter biased parameter estimates and distorted market structures if he/she ignores such dimensionality selection heterogeneity. Research limitations/implications - The proposed Bayesian spatial model provides information regarding how individual consumers utilize each dimension and how the relationship with behavioral variables can help marketers understand the underlying reasons for selective dimensional usage. Further, the proposed approach helps a marketing manager to identify major dimension(s) that could maximize the effect of a change of brand positioning, and thus identify potential opportunities/threats that existing MDS methods cannot provides. Originality/value - To date, no existent spatial model utilized for brand positioning can accommodate the various forms of heterogeneity exhibited by real consumers mentioned above. The end result can be very inaccurate and biased portrayals of competitive market structure whose strategy implications may be wrong and non-optimal. Given the role of such spatial models in the classical segmentation-targeting-positioning paradigm which forms the basis of all marketing strategy, the value of such research can be dramatic in many marketing applications, as illustrated in the manuscript via analyses of both synthetic and actual data.
In a sorting task, consumers receive a set of representational items (e.g., products, brands) and sort them into piles such that the items in each pile "go together." The sorting task is flexible in accommodating different instructions and has been used for decades in exploratory marketing research in brand positioning and categorization. However, no general analytic procedures yet exist for analyzing sorting task data without performing arbitrary transformations to the data that influence the results and insights obtained. This manuscript introduces a flexible framework for analyzing sorting task data, as well as a new optimization approach to identify summary piles, which provide an easy way to explore associations consumers make among a set of items. Using two Monte Carlo simulations and an empirical application of single-serving snacks from a local retailer, the authors demonstrate that the resulting procedure is scalable, can provide additional insights beyond those offered by existing procedures, and requires mere minutes of computational time.
The Segmentation-Targeting-Positioning (STP) process is the foundation of all marketing strategy. This chapter presents a new constrained clusterwise multidimensional unfolding procedure for performing STP that simultaneously identifies consumer segments, derives a joint space of brand coordinates and segment-level ideal points, and creates a link between specified product attributes and brand locations in the derived joint space. This latter feature permits a variety of policy simulations by brand(s), as well as subsequent positioning optimization and targeting. We first begin with a brief review of the STP framework and optimal product positioning literature. The technical details of the proposed procedure are then presented, as well as a description of the various types of simulations and subsequent optimization that can be performed. An application is provided concerning consumers' intentions to buy various competitive brands of portable telephones. The results of the proposed methodology are then compared to a naive sequential application of multidimensional unfolding, clustering, and correlation/regression analyses with this same communication devices data. Finally, directions for future research are given.
While the sport industry is a multibillion dollar industry, there is a paucity of academic marketing research regarding the various aspects of the industry, especially concerning fan avidity—the level of interest, involvement, passion, enthusiasm, and loyalty a fan exhibits to a sport entity. This is somewhat surprising given that avid fans are the lifeblood of any sport organization, spending significantly more money, time, and effort on sport-related products than other consumers. Thus, given its importance to the sport industry, we examine the relationship between fan avidity and its various behavioral manifestations. Recognizing the existence of consumer heterogeneity among fans, we present a new parametric constrained segmentation methodology and corresponding estimation algorithm that incorporates managerial constraints pertinent to the sport industry (or any other industry) while simultaneously segmenting the market and profiling each segment. We conducted a Monte Carlo simulation, which demonstrates the successful performance of the estimation algorithm across various models, data, and error structures. Then, we applied our proposed methodology to college football data for a major US university and found evidence for two distinct market segments. Finally, we performed a series of model comparisons and showed that our parametric constrained segmentation methodology outperforms existing alternatives.
Positioning is among a marketer’s preeminent strategic responsibilities. Positioning helps to clarify brand strengths among competitors and identify potential challenges of similar brands and possible substitutability. Assessments of positioning, from initial marketplace efforts to resources directed at modifications and re-positioning, are frequently assisted by the graphical representations of brands in multidimensional space. Such perceptual maps are constructed to reflect the closeness of brands and therefore the extent to which they are seen as interchangeable, versus distances between brands representing their relative positioning distinctiveness. To create perceptual maps, data are frequently obtained that comprise a sample of respondents rating a series of brands with respect to their perceived similarities and differences, as well as the status of each brand along multiple attributes. This research uses the variability inherent in such three-dimensional data to construct confidence regions around point estimates in perceptual maps. Current maps tend to be simply descriptive, with positions reflected by point estimates, but multivariate models including multidimensional scaling and multi-mode factor analysis can be modified to extract the subject heterogeneity and derive inferential perceptual maps. Confidence regions that overlap will indicate more clearly an inference of brand similarity, whereas non-overlapping regions imply statistically differentiated brand perceptions.
The interrelationships between two sets of measurements made on the same subjects can be studied by canonical correlation. Originally developed by Hotelling, the canonical correlation is the maximum correlation between linear functions or canonical factors of two sets of variables. An alternative pair of statistics to investigate the interrelationships between two sets of variables are the redundancy indices, developed by Stewart and Love. A redundancy coefficient is an index of the average proportion of variance in the variables in one set that is reproducible from the variables in the other set. Unlike canonical correlation, redundancy indices are non-symmetric in that a measure can be calculated for each set of variables (predictor and criterion) and need not be equal to each other. Van Den Wollenberg has developed a method of extracting factors that maximize redundancy, as opposed to canonical correlation. DeSarbo, Johansson, and Israels have developed extensions of this methodology. Takane and Hwang developed extended redundancy analysis to generalize redundancy analysis to investigate asymmetric or directional associations among more than two sets of variables, analogous to the work of Carroll and Kettenring regarding generalized canonical correlation analysis. A sports marketing application is provided examining the relationship between the different ways consumers/fans follow their college football team and their various attitudes, opinions, and lifestyles (i.e., psychographics) regarding sports.