Capturing valuable product/service improvement ideas is helpful for the development of new features. However, the existing methods for capturing such improvement ideas have the disadvantages of high cost, long time lag, information overload, and difficulty in getting a response. We propose an innovative framework based on lead user theory for capturing product/service improvement ideas from user-generated content on social media (henceforth called "chatter"). To identify the chatter containing improvement ideas, we design a machine-learning-based imbalanced classification model. Additionally, we use text summarization technology to get a rough sense of improvement ideas from the selected chatter. We validate the proposed framework by a case study in the automotive industry. The results demonstrate that the ideas extracted by our framework are breakthrough innovative, useful, feasible, and adoptable.
Social media provides customers with platforms on which to express usage experiences, opinions, preferences and expectations of product quality. Thus, a large number of reviews available on social media become an effective information source for quality management. In this study, we attempt to identify reviews that are helpful from the perspective of total quality management, called Helpful Quality-related Reviews (HQRs). First, we propose a definition and taxonomy of HQRs based on attractive quality theory. Then, we construct a Helpful Quality-related Review Identification (HQRI) model to mitigate the information overload represented by social media reviews. The HQRI model incorporates an imbalanced data classification method and a multi-label classification method based on the characteristics of HQRs. Experimental results demonstrate the effectiveness of the HQRI model in terms of six performance metrics in comparison with three state-of-the-art methods. Finally, we employ the Latent Dirichlet Allocation (LDA) topic generation model and word clouds to analyse and present the topics, specific manifestations, customer behaviours and related product components mentioned in HQRs.
User-generated content (UGC) is becoming increasingly available on social media for a wide range of products and services. Such UGC contains rich information about customer attitudes, opinions, and experiences. We propose a novel method for product competitive advantage analysis, which provides an essential basis for quality management and marketing strategy development, by mining UGC. Compared to traditional product performance analysis methods based on manufacturers' internal data and expert reviews, our method better reflects the perspective of customers. While a few recent methods based on UGC analysis assess the performance of one product in isolation, our method reveals the competitive advantages (and disadvantages) of a target product relative to its competitors. Our method uses supervised learning to identify competitors from UGC and domain-specific sentiment analysis to quantify customer attitudes. A case study in the automotive industry demonstrates the utility of the method.
[目的]从社交媒体用户生成内容中发现未知情感词,构造领域情感词典,应用于汽车评论的情感分析.[方法]选取HowNet情感词典作为种子,以实际汽车评论作为语料,分别利用PMI和Word2Vec算法识别新词情感极性,根据集成规则对二者识别结果综合判定,通过情感分类实验对比显示本文算法的有效性.[结果]按照该方法构造的情感词典准确率比HowNet情感词典提高21.6%,较分别使用PMI和Word2Vec算法构建的词典分别提升3.7%和2.1%,同时正面、负面情感词数量均有大幅增加.[局限]语料来源单一,应用于其他领域具有一定局限性.[结论]该方法构造的情感词典可有效应用于社交媒体文本情感分析.
[目的]从产品论坛中识别潜在客户,对产品论坛中的用户生成内容特征进行分析,识别有购买意愿的产品潜在客户.[方法]将不均衡数据集转换为n个均衡数据集,结合Stacking分类算法识别潜在客户,分别使用基分类器算法和本文提出的针对不均衡数据集的Stacking分类算法对样本数据进行测试,并通过对比F值验证本文算法的有效性.[结果]本文提出的算法的F值较贝叶斯网络、逻辑回归、C4.5决策树、SMO和朴素贝叶斯5种基分类器算法分别提高17.4%、26.5%、24.1%、29.3%、40.9%,较Stacking、Bagging和Boosting三种集成学习算法分别提高10.1%、5.9%、13.1%.[局限]研究语料来源于汽车行业,具有一定的领域局限性.[结论]该方法能有效识别潜在客户.
As social media are continually gaining more popularity, they have become an important source for manufacturers to collect information related to defects on their products from consumers. Researchers have started to develop automated models to identify mentions of product defects from social media, such as online discussion forums. In this paper, we propose a novel method for product defect identification from online forums, addressing two inadequacies in previous studies, namely, the inadequate use of information contained in replies and the straightforward use of standard single classifier methods. Our method incorporates contextual features derived from replies and uses a multi-view ensemble learning method specifically tailored to the problem on hand. A case study in the automotive industry demonstrates the utilities of both novelties in our method.
Traditional behavioral scoring models applying classification methods that yield a static probability of default may ignore the borrowers' dynamic characteristics because borrower repayment behavior evolves dynamically. In this study, we propose a novel behavioral scoring model based on a mixture survival analysis framework to predict the dynamic probability of default over time in peer-to-peer (P2P) lending. A random forest is utilized to identify whether a borrower will default, and a random survival forest is introduced to model the time to default. The results of an empirical analysis on a Chinese P2P loan dataset show that the proposed ensemble mixture random forest (EMRF) has a better performance in terms of predicting the monthly dynamic probability of default, while compared with standard mixture cure model, Cox proportional hazards model and logistic regression. It is also concluded that the proposed EMRF model provides a meaningful output for timely post-loan risk management. (C) 2017 Elsevier B.V. All rights reserved.
Reviews posted to social media are an effective source of information for helping quality managers to improve product quality. However, because helpful quality-related reviews may involve various aspects of product quality, previous studies confusing these aspects cannot provide targeted information regarding different aspects of product quality and production system improvement. In this paper, we propose a method of multi-class classification for helpful quality-related reviews corresponding to different aspects of product quality and production systems. Furthermore, the efficient and accurate identification of helpful quality-related reviews remains a critical challenge because of the sparseness of such reviews, which significantly influences classifier performance. To address these problems, we develop a model for the identification of helpful reviews called Helpful Quality-related Review Mining (HQRM) that incorporates a multi-class classification architecture and imbalanced data classification methods. The experimental results show that HQRM enables the multi-class classification of helpful quality-related reviews with significantly improved precision, recall and F-measure values.
通过社会媒体信息预测股票行为已经成为近年来金融和知识管理等领域的研究热点.考虑到社会媒体参与人员和讨论话题的多样性,传统的基于整体层面分析社会媒体信息来预测股票行为的方法过于粗糙.本文根据社会媒体信息在写作风格和内容特征上的不同,利用文本特征提取技术、主成分分析法、EM聚类技术等分析参与社会媒体的干系人和他们关注的话题.进一步,我们针对每类干系人和话题,从信息活动强度和情感倾向两个方面提取四个社会媒体变量构建股票行为的回归预测模型,用以分析各干系人和话题在社会媒体上的活动状况对公司股票行为的影响.最后,本文以雅虎金融论坛的Bank of America板块为实验平台进行实验研究,验证了所提出方法的有效性和实用性.