The accurate prediction of industry trends has become increasingly challenging because of unforeseen events. To address this challenge, this study proposes a deep learning approach to generate an economic sentiment index by integrating Natural Language Processing (NLP) models and image-clustering techniques. We first employ sampling techniques to create standardized online news datasets. Feature engineering techniques from the Korean Bidirectional Encoder Representations from Transformers (KoBERT) model are then used to generate relevance and sentiment scores for the textual data. Further, to enhance visualization and clustering, we transform the textual data into joint plot images, which are grouped into distinct clusters based on news categories. Finally, using Multi-criteria Decision Analysis, the various scores and cluster information are synthesized to generate the final economic sentiment index. This approach improves visualization and enhances the interpretability of the generated index. The proposed algorithm is applied to construct a new economic sentiment index for the Information and Communications Technology (ICT) industry in South Korea.
Curvilinear growth trajectories of products/services are common in the tourism and hospitality industry. To fit nonlinear growth patterns with multilevel structures, this study proposed Bayesian hierarchical growth curve models (BHGCMs) in line with the increasing adoption of Bayesian analysis in tourism and hospitality academia. We provided the basic form of BHGCM along with its unique advantages. For an empirical test, this study applied several growth curves to approximate online reviews of U.S. hotels from August 2020 to January 2021. After selecting a Gompertz curve as a mean function of the GCM, a Bayesian hierarchical approach was employed to estimate growth parameters—namely, base and maximum volume of hotel reviews, inflection week, and relative growth rate—and identify their determinants. Our findings demonstrate the superiority of the proposed BHGCM in fitting the growth patterns of hotel reviews while revealing the effect of price and accumulated reviews on the parameters.
Previous studies have extensively investigated the effects of online word-of-mouth (eWOM) factors such as volume and valence on product sales. However, studies of the effect of eWOM factors on product prices are lacking. It is necessary to examine how various eWOM factors can either explain or affect product prices. The objective of this study is to suggest explanatory and predictive analytics using a regression analysis and ensemble-based machine learning methods for eWOM factors and hotels booking prices. This study utilizes publicly available data from a hotel booking site to build a sample of eWOM factors. The final study sample was comprised of 927 hotels. The important eWOM factors found to affect hotel prices are the review depth and the review rating, which are moderated by a number of reviews to affect prices. The effect of the number of positive words is moderated by the review helpfulness to affect the price. The review depth and rating, along with the number of reviews, should be considered in the design of hotel services, as these provide the rationale for adjusting the prices of various aspects of hotel services. Furthermore, the comparison results when applying various ensemble-based machine learning methods to predict prices using eWOM factors based on a 46-fold cross-validation partition method indicated that ensemble methods (bagging and boosting) based on decision trees outperformed ensemble methods based on k-nearest neighbor methods and neural networks. This shows that bagging and boosting methods are effective ways to improve the prediction performance outcomes when using decision trees. The explanatory and predictive analytics using eWOM factors for hotel booking prices offers a better understanding in terms of how the accommodation prices of hotel services can be explained and predicted by eWOM factors.
IntroductionWith rapid advancements in natural language processing (NLP), predicting personality using this technology has become a significant research interest. In personality prediction, exploring appropriate questions that elicit natural language is particularly important because questions determine the context of responses. This study aimed to predict levels of neuroticism—a core psychological trait known to predict various psychological outcomes—using responses to a series of open-ended questions developed based on the five-factor model of personality. This study examined the model’s accuracy and explored the influence of item content in predicting neuroticism.MethodsA total of 425 Korean adults were recruited and responded to 18 open-ended questions about their personalities, along with the measurement of the Five-Factor Model traits. In total, 30,576 Korean sentences were collected. To develop the prediction models, the pre-trained language model KoBERT was used. Accuracy, F1 Score, Precision, and Recall were calculated as evaluation metrics.ResultsThe results showed that items inquiring about social comparison, unintended harm, and negative feelings performed better in predicting neuroticism than other items. For predicting depressivity, items related to negative feelings, social comparison, and emotions showed superior performance. For dependency, items related to unintended harm, social dominance, and negative feelings were the most predictive. DiscussionWe identified items that performed better at neuroticism prediction than others. Prediction models developed based on open-ended questions that theoretically aligned with neuroticism exhibited superior predictive performance.
In recent years, virtual online communities have experienced rapid growth. These communities enable individuals to share and manage images or websites by employing tags. A collaborative tagging system (CTS) facilitates the process by which internet users collectively organize resources. CTS offers a plethora of useful information, including tags and timestamps, which can be utilized for recommendations. A tag represents an implicit evaluation of the user’s preference for a particular resource, while timestamps indicate changes in the user’s interests over time. As the amount of information increases, it is feasible to integrate more detailed data, such as tags and timestamps, to improve the quality of personalized recommendations. The current study employs collaborative filtering (CF), which incorporates both tag and time information to enhance recommendation precision. A computational recommender system is established to generate weights and calculate similarities by incorporating tag data and time. The effectiveness of our recommendation model was evaluated by linearly merging tag and time data. In addition, the proposed CF method was validated by applying it to big data sets in the real world. To assess its performance, the size of the neighborhood was adjusted in accordance with the standard CF procedure. The experimental results indicate that our proposed method significantly improves the quality of recommendations compared to the basic CF approach.
It has become increasingly important to consider the efficiency of movies in creating box revenue while using fewer movie resources. Further, there is a lack of eWOM (online-word-of-mouth) studies regarding using the production efficiency of movies as a dependent outcome measure replacing box revenue. This study shows that production efficiency can be suggested by comparing movie resources powers, i.e., powers of actors, directors, distributors, and production companies, which are input for movie production, and the box office. For testing the validity of the measure of production efficiency, this study examines the effect of eWOM attributes, i.e., review depth, volume, rating, review sentiment, and helpfulness on production efficiency. Data envelopment analysis is adopted to produce the efficiency of movies. This study provides insights into a current movie study on eWOM by showing the effect of interaction between eWOM (review rating) and helpfulness on production efficiency. Further, this study purports to test the prediction power in predicting production efficiency using decision trees, neural networks, and logistic regression. These results show that k nearest neighbor and automated neural networks outperform the other machine learning methods in classifying efficient movies.
Purpose This study aims to provide a way to derive inter-brand similarities from user-generated content on online brand forums, which enables the authors to analyze the market structures based on consumers' actual information searching and sharing behavior online. This study further presents a method for deriving inter-brand similarities from data on how the sales of competing brands covary over time. The results obtained by the above two methods are compared to each other. Design/methodology/approach In drawing similarities between brands, the authors utilized a newly proposed measure that modified the lift measure. The derived similarity information was applied to multidimensional scaling (MDS) to analyze the perceived market structure. The authors applied the proposed methodology to the imported car market in South Korea. Findings In light of some clear information such as the country of origin, the market structure derived from the presented methodology was seen to accurately reflect the consumer's perception of the market. A significant relevance has been found between the results derived from user-generated online content and sales data. Originality/value The presented method allows marketers to track changes in competitive market structures and identify their major competitors quickly and cost-effectively. This study can contribute to improving the utilization of the overflowing information in the big data era by proposing methods of linking new types of online data with existing market research methods.
BackgroundSelf-report multiple choice questionnaires have been widely utilized to quantitatively measure one’s personality and psychological constructs. Despite several strengths (e.g., brevity and utility), self-report multiple choice questionnaires have considerable limitations in nature. With the rise of machine learning (ML) and Natural language processing (NLP), researchers in the field of psychology are widely adopting NLP to assess psychological construct to predict human behaviors. However, there is a lack of connections between the work being performed in computer science and that of psychology due to small data sets and unvalidated modeling practices.AimsThe current article introduces the study method and procedure of phase II which includes the interview questions for the five-factor model (FFM) of personality developed in phase I. This study aims to develop the interview (semi-structured) and open-ended questions for the FFM-based personality assessments, specifically designed with experts in the field of clinical and personality psychology (phase 1), and to collect the personality-related text data using the interview questions and self-report measures on personality and psychological distress (phase 2). The purpose of the study includes examining the relationship between natural language data obtained from the interview questions, measuring the FFM personality constructs, and psychological distress to demonstrate the validity of the natural language-based personality prediction.MethodsPhase I (pilot) study was conducted to fifty-nine native Korean adults to acquire the personality-related text data from the interview (semi-structured) and open-ended questions based on the FFM of personality. The interview questions were revised and finalized with the feedback from the external expert committee, consisting of personality and clinical psychologists. Based on the established interview questions, a total of 300 Korean adults will be recruited using a convenience sampling method via online survey. The text data collected from interviews will be analyzed using the natural language processing. The results of the online survey including demographic data, depression, anxiety, and personality inventories will be analyzed together in the model to predict individuals’ FFM of personality and the level of psychological distress (phase 2).
The automotive industry evaluates various success factors to achieve competitive advantage in selling products. Existing studies have predicted the success of newly launched automobiles based on an economic perspective. However, factors such as dynamic changes in consumer preferences and the emergence of numerous automobile brands pose difficulty in understanding product quality. This study proposes a method of understanding the automotive market using text mining techniques and online user opinions for newly launched cars. By analyzing customer experiences and expectations through their opinions, we can anticipate automobile demand in the market more easily. The proposed method is based on online reviews from an online portal for automobiles. Based on a literature review, this study presents a framework for analyzing input versus output word-of-mouth (WOM). It also integrates the success factors from existing automobile studies and derives functional categories and relevant keywords. The analysis identifies differences in consumer-interest factors that lead to short-term success or normal results in automobile sales. In addition, it confirms that the elements of WOM produces varying results depending on the timing these are employed in relation to the product launch (i.e., before or after a product’s launch). It revealed which dimensions of automobile characteristics are important factors in identifying sales volume and market share for specific types and brands of automobile models. The results of this study provide theoretical advantage in predicting market success in the automobile industry. In addition, the study derives practical insights into characteristics of classification information for market forecasts in the automotive industry. The paper provides empirical insights about how input WOM and output WOM which are analyzed differently can have predictive power in forecasting market share and sales volume for automobiles.
As social platforms become essential in promoting songs, many artists create an official channel on YouTube and encourage their fans' engagements. However, little is known about the effectiveness of official videos in generating fans' media engagements. We conduct two empirical studies to investigate the relationship between official video attributes, media engagements and channel subscribers. First, we propose a model to explain how music video attributes facilitate the relative social media engagements of a video. Second, we test whether three types of social engagements are associated with an increase in official channel subscribers. To do so, we collect social media engagement data for the 2896 music videos uploaded between May 2016 and April 2019 by 105 artists who own their YouTube official artist channels. The empirical evidence shows that the official videos incorporating visual, performance and storytelling components can generate more positive engagement from the viewers, while audio-only videos exhibit lower overall engagement intensity. Furthermore, we find that active media engagement, such as leaving 'comments', contributes to the increase in channel subscribers beyond the effect of the number of registered videos. Our results suggest that highly involved music fans may show the active types of social engagements, such as leaving a comment on live and follow-up videos. In practice, our findings imply that YouTube creators need to incorporate visual-focused platform characteristics to stimulate in-depth social engagement.
Speech signals are being used as a primary input source in human–computer interaction (HCI) to develop several applications, such as automatic speech recognition (ASR), speech emotion recognition (SER), gender, and age recognition. Classifying speakers according to their age and gender is a challenging task in speech processing owing to the disability of the current methods of extracting salient high-level speech features and classification models. To address these problems, we introduce a novel end-to-end age and gender recognition convolutional neural network (CNN) with a specially designed multi-attention module (MAM) from speech signals. Our proposed model uses MAM to extract spatial and temporal salient features from the input data effectively. The MAM mechanism uses a rectangular shape filter as a kernel in convolution layers and comprises two separate time and frequency attention mechanisms. The time attention branch learns to detect temporal cues, whereas the frequency attention module extracts the most relevant features to the target by focusing on the spatial frequency features. The combination of the two extracted spatial and temporal features complements one another and provide high performance in terms of age and gender classification. The proposed age and gender classification system was tested using the Common Voice and locally developed Korean speech recognition datasets. Our suggested model achieved 96%, 73%, and 76% accuracy scores for gender, age, and age-gender classification, respectively, using the Common Voice dataset. The Korean speech recognition dataset results were 97%, 97%, and 90% for gender, age, and age-gender recognition, respectively. The prediction performance of our proposed model, which was obtained in the experiments, demonstrated the superiority and robustness of the tasks regarding age, gender, and age-gender recognition from speech signals.
The COVID-19 pandemic has significantly changed individuals' daily life due to increased risk aversion, which has affected their consumption patterns and preferences. To understand the effect of the pandemic on consumer behavior through risk aversion, this study investigated the relationships among the pandemic, social distancing, online information search, and firm performance in the hospitality and tourism industries. For data analysis, we developed two joint models and estimated the models using the fixed-effects method. The results of the first model showed that social distancing triggered by COVID-19 news stories affected firm value. The second regional-level analysis revealed that the number of confirmed cases and COVID-19 news stories influenced individuals’ social distancing and online information search for tourist attractions and the changed social distancing and online search, in turn, affected the volume of online hotel reviews.
BACKGROUND:The coronavirus disease 2019 (COVID-19) pandemic has prompted a global-scale public health response. Social distancing, along with intensive testing and contact tracing, has been considered an effective vehicle to reduce new infections. In this study, we aimed to estimate the effect of South Korean public health measures on behavioral changes with respect to social distancing without a nationwide lockdown. The results of this study may provide insights to countries who are planning to relax their aggressive restrictions though still having an unflattened curve of infections.METHODS:To estimate how the closure of educational and social welfare facilities and the disclosure of confirmed patients' contact history affected social distancing behaviors, we analyzed public transportation data in Seoul, Korea. For the modeling analysis, we used linear mixed-effects estimation.RESULTS:Our estimation showed that the average daily number of bus passengers decreased by 21.8% in February 2020 as compared to the previous year with an additional decrease observed in the areas around educational and social welfare facilities. The highest drop in the daily number of passengers was observed in areas with religious facilities. We also found that individuals avoided areas that were proximate to or within the locations that constituted the confirmed patients' contact history.CONCLUSION:Our results demonstrate that public health measures can lead people to practice social distancing. Among them, the measures that strongly encourage voluntary social distancing behaviors would play a critical role in suppressing the infections as it becomes increasingly difficult to continue imposing aggressive restrictions due to practical and economic reasons.
While many business intelligence methods have been applied to predict movie box office revenue, the studies using an ensemble approach to predict box office revenue are almost nonexistent. In this study, we propose decision trees, k-nearest-neighbors (k-NN), and linear regression using ensemble methods and the prediction performance of decision trees based on random forests, bagging and boosting are compared with that of k-NN and linear regression based on bagging and boosting using the sample of 1439 movies. The results indicate that ensemble methods based on decision trees (random forests, bagging, boosting) outperform ensemble methods based on k-NN (bagging, boosting) in predicting box office at week 1, 2, 3 after release. Decision trees using ensemble methods provide better prediction performance than ensemble methods based on linear regression analysis in the box office at week 1 after release. This is explained by the results that after comparing the prediction performance between ensemble methods and non-ensemble methods. For decision tree methods, unlike the other methods, the prediction performance of ensemble methods is greater than that of non-ensemble methods. This shows that decision trees using ensemble methods provide better application effectiveness of ensemble methods than k-NN and linear regression analysis.
The enormous volume and largely varying quality of available reviews provide a great obstacle to seek out the most helpful reviews. While Naive Bayesian Network (NBN) is one of the matured artificial intelligence approaches for business decision support, the usage of NBN to predict the helpfulness of online reviews is lacking. This study intends to suggest HPNBN (a helpfulness prediction model using NBN), which adopts NBN for helpfulness prediction. This study crawled sample data from Amazon website and 8699 reviews comprise the final sample. Twenty-one predictors represent reviewer and textual traits as well as product traits of the reviews. We investigate how the expanded list of predictors including product, reviewer, and textual characteristics of eWOM (online word-of-mouth) has an effect on helpfulness by suggesting conditional probabilities of the binned determinants. The prediction accuracy of NBN outperformed that of the k-nearest neighbor (kNN) method and the neural network (NN) model. The results of this study can support determining helpfulness and support website design to induce review helpfulness. This study will help decision-makers predict the helpfulness of the review comments posted to their websites and manage more effective customer satisfaction strategies. When prospect customers feel such review helpfulness, they will have a stronger intention to pay a regular visit to the target website.
The social engagement of eWOM (electronic word-of-mouth) can reduce the threat of adverse selection in e-commerce. As studies that examine the social influence of eWOM are rare, the present work suggests the moderating effect of review or reviewer helpfulness and product type (experience or search goods) on the relationship between eWOM and product sales. The volume of eWOM, which is defined as the multiplication of the average length by the number of reviews, is shown to be moderated by review and reviewer helpfulness and search goods to affect product sales. Review ratings are moderated by reviewer helpfulness, and review extremity is positively (negatively) moderated by search (experience) goods and review helpfulness to affect product sales. As previous studies of differentiated sampling strategies that consider review helpfulness for predicting product sales using eWOM are lacking, this study compares the prediction power of business intelligence methods for different subsamples of products created according to high or low review and reviewer helpfulness levels. The subsample with high review or reviewer helpfulness demonstrates greater prediction performance than the subsample with low review or reviewer helpfulness when eWOM variables are used as predictors of product sales. Hence, preliminary filtering data preprocessing should consider review or reviewer helpfulness as a crucial criterion of the data quality. This will contribute to the sampling or preprocessing strategy used to predict product sales using eWOM.
Purpose This paper aims to intend to study the effect of movie production efficiency on eWOM and the moderating effect of efficiency on the relationship between eWOM and review helpfulness for movies. Design/methodology/approach Production efficiency is suggested by comparing the power of movie resources (e.g. the power of actors, directors, distributors, production companies) against box-office revenue through a data envelopment analysis (DEA). Findings The study results present that the number of reviews, the number of reviews by reviewers and review extremity are greater in an efficient subsample than in an inefficient subsample. For efficient movies, the review depth and the strength of the sentiments in the reviews are more positively related to review helpfulness. The prediction results for review helpfulness using the k-nearest neighbor method and automatic neural networks show that the efficient subsample provides a significantly lower prediction error rate than the inefficient subsample. The study results can support the effective facilitation of helpful online movie reviews. Originality/value As the numbers of online reviews are increasingly used to provide purchase decision support, it becomes crucial to understand which attributes represent average helpful reviews for movies. While previous studies have examined eWOM (online word-of-mouth) variables as predictors of helpfulness on movie websites, the role of the production efficiency of movies has not been examined considering the relationship between eWOM and review helpfulness for movies.
Process mining in the context of information systems, which consists of information flows, has been one of the major research areas in the past decade. One of the most common objectives of process mining is the automated business process discovery. There are many challenges in the business process discovery, such as spaghetti models, same-name activities, and discovering loop structures. The researchers have presented a variety of methods that focus on one or more challenges. Due to the importance of commercial systems and the diversity of flows in them, in this research, the process mining problem in the context of social commerce systems is studied. Moreover, the research objective is to present a new method for commercial process discovery that has not been considered before. The proposed method is based on network analysis methods, multi-layered networks (networks with heterogeneous relations), and attributed networks. The results obtained from the proposed method are more precise and more comfortable to understand than the previous ones.
The studies are almost nonexistent regarding production efficiency of movies which is determined based on the relationship between movie resources powers (powers of actors, directors, distributors, and production companies) and box office. Our study attempts to examine how efficiency moderates the relationship between eWOM (online word-of-mouth) and revenue, and to show the difference in prediction performance between efficient and inefficient movies. Using data envelopment analysis to suggest efficiency of movies, movie efficiency negatively moderates the effects of review depth and volume on subsequent box office revenue compensating negative effects of smaller box office in previous period while efficiency exert a positive moderating effect on the influences of review rating and the number of positive reviews on revenue. This shows that review depth and volume are affected by the slack of movie resources powers for inefficient movies, and high rating and positive response for efficient movies to affect revenue. The results of decision trees, k-nearest-neighbors, and linear regression analysis based on ensemble methods using eWOM or movie variables indicate that the movies with the inefficient movie resources powers are providing greater prediction performance than movies with efficient movie resources powers. This show that diverse variation in the efficiency of movie resources powers contributes to prediction performance.