We propose an approach for uncovering characteristic response paths of a population from an individual-level multivariate time series data set. The approach is based on a model that accommodates a set of arbitrary distributions for endogenous variables and interstep intervals and variables. The model enables reliable estimation of individual-level parameters by uncovering and statistically pooling clusters of similar individuals. We show that using such a model one can distribute the response of an outcome variable to an impulse over all possible preceding activity sequences. When a few such sequences explain most of the response, they describe the population’s characteristic response paths from the impulse to the outcome. We apply the proposed approach to a customer touchpoint data set from a large multichannel specialty retailer. This application uncovers six customer segments, each with unique characteristic paths to purchase. These paths provide insights into the behavior of customers and the optimal over-time communication strategy for different customer segments. Summary of Contribution: Uncovering users’ paths through physical and virtual spaces has been of considerable interest in the computing and operations research domain. The existing research suggests a demand for visualizing the primary paths of agents through geographic, online, and activity spaces. Thus far, most of the research have developed approaches that are unique to specific domains providing insight into the domain in the process. There is a need for a general statistically robust approach that can be applied to a broad range of domains to uncover variable sequences that lead to outcomes of interest. We propose a computational approach to uncover characteristic response paths of a population from an individual-level multivariate time series data set. The approach is based on a statistical model that accommodates arbitrary and mixed set of distributions for the endogenous variables, accommodates intersession intervals and variables, and reliably estimates individuals’ parameters through statistical pooling by uncovering clusters of similar members. These features make the proposed model suitable for a large variety of real-world datasets. We show that using such a model one can extract characteristic paths over possible activity sequences starting from an impulse leading up to a target variable of interest.
This article studies the question and answer (Q&A) technology of electronic commerce platforms, an increasingly common form of user-generated content that allows consumers to publicly ask product-specific questions and receive responses, either from the platform or from other customers. Using data from a major online retailer, the authors show that Q&As complement consumer reviews: unlike reviews, questions are primarily asked prepurchase and focus on clarification of product attributes rather than discussion of quality; answers convey fit-specific information in a predominantly sentiment-free way. Drawing on these observations, the authors hypothesize that Q&As mitigate product fit uncertainty, leading to better matches between products and consumers and, therefore, improved product ratings. Indeed, when products suffering from fit mismatch start receiving Q&As, their subsequent ratings improve by approximately .1 to .5 stars, and the fraction of negative reviews that discuss fit-related issues declines. The extent of the rating increase due to Q&As is proportional to the probability that purchasers will experience fit mismatch without Q&A. These findings suggest that, by resolving product fit uncertainty in an e-commerce setting, the addition of Q&As can be a viable way for retailers to improve ratings of products that have incurred low ratings due to customer–product fit mismatch.
The term Market Vandalism Attack refers to large scale review manipulation, not for the benefit or detriment of individual sellers, but in order to damage the market itself. In this paper we present a theoretical framework for such attacks that allows us to provide reasonable estimates on the outcomes of vandalism attacks in the presence of countermeasures from the market operator, even when complete review-level data are not available. Based on our theoretical foundation, we estimate the cost of such attacks in four different markets (airline, beer, movies, and cocaine) assuming that in each case the best known generic review manipulation countermeasures are employed. We find that in all cases, an attacker who can afford to post between 10\% and 40\% as many reviews by injected (fake) profiles as there are by legitimate market participants can equalize all review scores, making it very difficult for the market to operate. We discuss an application of market vandalism in the context of Dark Net Markets (DNMs) and consider the potential for vandalism attacks to be used in conjunction with other methods in the policing of these markets by Law Enforcement Agencies (LEAs).
Although many researchers in Information Systems and Marketing have studied the effect of product reviews on sales, few have looked at their effect on product returns. We hypothesize that, by affecting the quality of purchase decisions, product reviews influence the probability of the eventual return of the purchased products. We elaborate this hypothesis by developing an analytical model that shows how changes in the precision of product quality and fit information affect the return probabilities of risk- averse, but rational, consumers. We empirically validate the predictions of our theory using a transaction level dataset from a multi-channel, multi-brand specialty retailer operating in North America. Harnessing data on multiple purchases and returns of the same products, but with varying sets of product reviews over a period of two years, we find that the availability of higher volumes of reviews, as well as review bodies that contain a higher percentage of reviews that are designated as ‘helpful’ by consumers, lead to lower incidence of product returns, after controlling for customer, product and other context-related factors. These results are consistent with the predictions of our theoretical model, and suggest that online reviews indeed help consumers make better purchase decisions leading to lower product returns.
Social media have great potential to support diverse information sharing, but there is widespread concern that platforms like Twitter do not result in communication between those who hold contradictory viewpoints. Because users can choose whom to follow, prior research suggests that social media users exist in "echo chambers" or become polarized. We seek evidence of this in a complete cross section of hyperlinks posted on Twitter, using previously validated measures of the political slant of news sources to study information diversity. Contrary to prediction, we find that the average account posts links to more politically moderate news sources than the ones they receive in their own feed. However, members of a tiny network core do exhibit cross-sectional evidence of polarization and are responsible for the majority of tweets received overall due to their popularity and activity, which could explain the widespread perception of polarization on social media.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
In this paper, we study the question and answer (Q&A) feature of electronic commerce platforms, an increasingly common form of user-generated content (UGC) that allows consumers to publicly ask product-specific questions and receive responses, either from the platform or from other customers. Using data from a major online retailer, we show that Q&As complement reviews and ratings: unlike reviews, Q&As primarily happen pre-purchase, focus on clarification of product attributes (rather than discussion of quality), and convey fit-specific information in a sentiment-free way. Our main hypothesis is that Q&As mitigate product fit uncertainty, leading to better matches between products and consumers, and therefore improved product ratings. We show that when low-rated products start receiving Q&As, their subsequent ratings improve by approximately 0.5 stars. We further show that the extent of the rating increase due to Q&As is moderated by the degree of ex-ante fit uncertainty. Overall, our findings suggest that, by resolving product fit uncertainty in an e-commerce setting, the addition of Q&As can be a viable way for retailers to improve ratings and sales of low-rated products, particularly those products that have incurred low ratings due to customer-product fit mismatch.
News aggregators have emerged as an important component of digital content ecosystems, attracting traffic by hosting curated collections of links to third-party content, but also inciting conflict with content producers. Aggregators provide titles and short summaries (snippets) of articles they link to. Content producers claim that their presence deprives them of traffic that would otherwise flow to their sites. In light of this controversy, we conduct a series of field experiments whose objective is to provide insight with respect to how readers allocate their attention between a news aggregator and the original articles it links to. Our experiments are based on manipulating elements of the user interface of a Swiss mobile news aggregator. We examine how key design parameters, such as the length of the text snippet that an aggregator displays about articles, the presence of associated images, and the number of related articles on the same story, affect a reader’s propensity to visit the content producer’s site and read the full article. Our findings suggest the presence of a substitution relationship between the amount of information that aggregators offer about articles and the probability that readers will opt to read the full articles at the content producer sites. Interestingly, however, when several related article outlines compete for user attention, a longer snippet and the inclusion of an image increase the probability that an article will be chosen over its competitors. This paper was accepted by Lorin Hitt, information systems.
A digital consumer’s purchase journey, referred to as the path to purchase, is non-linear and heterogeneous. Despite a strong interest in this concept, there are few published approaches to empirically extract consumers’ path to purchase (in terms of a sequence of different types of activities leading to purchase), especially in settings where consumers engage in multiple simultaneous activities in each period. We address this gap by proposing a methodology that identifies consumers’ paths to purchase from commonly available CRM touch point data. We propose a generalized multivariate autoregressive (GMAR) model to capture the interactions among distinct but potentially simultaneous activities of a consumer over time. Using the proposed model we show how to attribute parts of the purchase volume to consumer activity sequences, or paths, starting from an initial marketing stimulus leading to the maximal purchase response. We embed the GMAR model in a clustering framework that endogenously identifies segments of consumers who exhibit similar paths to purchase. We apply the methodology to a dataset from a multi-channel North American Specialty Retailer to uncover the distinct paths of five consumer segments: loyal and engaged shoppers, digitally-driven offline shoppers, holiday shoppers, infrequent offline shoppers, and frequently targeted occasional shoppers. Using out-of-sample forecasts, we demonstrate improved predictions of future purchases compared to extant methods. Finally, we perform policy simulations to show that managers can use the uncovered path information to dynamically optimize marketing campaigns for each segment.
To what extent do social media like Twitter support diversity of ideas and communication between those who hold contradictory viewpoints? Despite the evident potential for social media in this regard, two prominent theories predict that such diverse discourse may be very limited. We review the ideas of
Vector Autoregression Models (VAR) are widely used by researchers to capture the linear interdependencies among multiple time series. We propose a novel method called Clustered VAR (CVAR) to identify components of the data generated by a mixture of K VAR processes. By applying a CVAR model to a consumer-level time series dataset on shopping behavior at a retailer, we segment consumers based on their path-to-purchase. We estimate the CVAR model using the EM (expectation–maximization) algorithm that assigns each consumer into a segment that maximizes the likelihood and optimizes the VAR parameters for each segment given the membership assignments. We verify the effectiveness of the Clustered VAR model on a simulated dataset. Following successful evaluation, we apply the Clustered VAR model to a retail dataset from a major multi-channel, multi-brand North American Retailer. Our study could segment 2,000 randomly selected consumers into 4 clusters and offers insights on two issues: 1. Potential interdependencies among online marketing, offline marketing and their effects for each group, 2. Differences in the above effects across consumer segments. As a result, the consumer clusters in our study will guide managers in tailoring the marketing mix for different customer segments to help them move forward on the path-to-purchase.
We propose a novel method to identify predominant paths-to-purchase of retail consumers from activity level dataset collected in CRM systems. We verify the effectiveness of the proposed model on a simulated dataset. Following successful verification, we apply the model on a retail dataset from a major multi-channel, multi-brand North American Retailer. We uncover three different types of consumers based on how they respond to external stimuli over time: catalog driven shoppers, email driven shoppers, and holiday driven online shoppers. We also find significant activity across channels by these consumers. Finally, we use the path information in the segments to identify the groups that are most sensitive to a certain type of marketing contact. By analyzing the response of customers in different groups in a test dataset, we show that managers can optimize marketing budget allocation using our proposed segmentation approach.
Several past studies have commented on the uneven distribution of contributions in online social production communities while at the same time highlighting the successful end products of many such communities. These two seemingly paradoxical situations are made possible through smaller groups of highly devoted volunteers who act as catalysts in organizing and maintaining community outputs. These volunteers have been referred to as “knowledge janitors.” There is currently limited understanding of how the group composition and interaction patterns of knowledge janitors affect social production quality outcomes. This study provides answers to these questions in the context of Wikipedia. By analyzing 11,359 changes in Wikipedia article quality, we found that cohesiveness, diversity, and equal distribution of communication turn-taking of an article’s janitors increase the likelihood of that article’s quality improvement. These main findings are further refined by considering how the main effects differ at different development stages of an article. The study’s contributions to research and implications to practice are discussed.
Motivated by the growing practice of using social network data in credit scoring, we analyze the impact of using network -based measures on customer score accuracy and on tie formation among customers. We develop a series of models to compare the accuracy of customer scores obtained with and without network data. We also investigate how the accuracy of social network -based scores changes when consumers can strategically construct their social networks to attain higher scores. We find that those who are motivated to improve their scores may form fewer ties and focus more on similar partners. The impact of such endogenous tie formation on the accuracy of consumer scores is ambiguous. Scores can become more accurate as a result of modifications in social networks, but this accuracy improvement may come with greater network fragmentation. The threat of social exclusion in such endogenously formed networks provides incentives to low-type members to exert effort that improves everyone's creditworthiness. We discuss implications for managers and public policy.
This paper delineates the main characteristics of the evolution of the organization as a social business in response to the socially networked marketplace. We advance the notion that the modern day firm is increasingly organized as a community according to the principle of collaboration. The main message is that the prominence of organizational structure is not redundant but needs to be complemented by collaborative community in response to market demands. In order to fulfill this complementary role, the concept of organization is profoundly changing. Based on recent theorizing, we review the role of collaborative community as a key characteristic of social business, provide an overview of its principles, show how social media can effectively facilitate and support collaborative community, and introduce the concept of expressive individuality. We provide illustrative examples that feature Dell. We conclude by identifying an agenda for further academic inquiry, and by specifying a large number of issues that researchers may address.
In Section 3 of the paper, we use a reduced form model to account for how consumers select their anchor node. In this section, we provide a formal micro-model of a process that provides justi cation for our assumptions. We also consider the possibility that some consumers may switch their anchor nodes when encountering a link to a higher quality site and show that including this feature in our model does not qualitatively change our results. In common with other analyses of web-browsing behavior, we employ a Markovian model to abstract the anchor node selection process. Model states i = 0, 1, .., N designate which site a consumer uses as her anchor node. State 0 corresponds to the situation where the consumer anchors herself at the outside option. We de ne transition probabilities from one state to another, representing the likelihood that consumers change their anchor node. We follow the behavior of one randomly selected consumer and measure the probabilities that the consumer anchors herself at a particular node. Let pi de ne the probability that the consumer is anchored at site i = 0, 1, ..N ; p0 corresponds to the outside option. Let wij measure the transition probability from node i to j, that is, the probability that the consumer switches her anchor from i to j, given that her current anchor is j. We assume
Considering new business models for massive open online courses.