Competitive intelligence is essential for operations management decision-making. Beyond traditional offline information channels, firms increasingly gather online data and resources to generate comprehensive competitive intelligence. This study derives competitive intelligence in large markets by developing an interpretable machine learning framework that integrates multifaceted user behavior data, including user favorites, user-commented products, and user textual comments. Considering the complementary nature of these data sources, we first combine latent features derived from user favorites and user-commented products to improve submarket inference. Using these inferred submarkets as supervised signals, we connect user-commented products and associated textual comments to uncover consumer perceptions. We estimate the model using multifaceted data on online user behavior in the automotive domain. The results demonstrate that our model effectively improves submarket identification, captures consumer perceptions, and predicts competitive positions for new entrants. The derived competitive intelligence helps managers make more informed decisions in product operations and marketing strategies.
Dataset recommendation is pivotal for streamlining data selection and accelerating scientific discovery. In this study, we propose the Sparse-Link Dataset Recommendation Model (SLDRM), an explainable framework that maps textual content, authors, and datasets into a unified topic space. Specifically, SLDRM captures the correlations among words, research communities, and dataset usage patterns by linking their respective latent topics. To handle the inherent sparsity of the research landscape, we incorporate a Spike-and-Slab prior. We validate our model using a real-world dataset collected from the PapersWithCode website. Experimental results show that our model not only improves recommendation accuracy but also enhances interpretability. The proposed model provides researchers with an efficient tool for dataset discovery and deepens the understanding of the knowledge production process in scientific networks.
Web attacks are one of the major and most persistent forms of cyber threats, which bring huge costs and losses to web application-based businesses. Various detection methods, such as signature-based, machine learning-based, and deep learning-based, have been proposed to identify web attacks. However, these methods either (1) heavily rely on accurate and complete rule design and feature engineering, which may not adapt to fast-evolving attacks, or (2) fail to estimate model uncertainty, which is essential to the trustworthiness of the prediction made by the model. In this study, we proposed an Uncertainty-aware Ensemble Deep Kernel Learning (UEDKL) model to detect web attacks from HTTP request payload data with the model uncertainty captured from the perspective of both data distribution and model parameters. The proposed UEDKL utilizes a deep kernel learning model to distinguish normal HTTP requests from different types of web attacks with model uncertainty estimated from data distribution perspective. Multiple deep kernel learning models were trained as base learners to capture the model uncertainty from model parameters perspective. An attention-based ensemble learning approach was designed to effectively integrate base learners' predictions and model uncertainty. We also proposed a new metric named High Uncertainty Ratio-F Score Curve to evaluate model uncertainty estimation. Experiments on BDCI and SRBH datasets demonstrated that the proposed UEDKL framework yields significant improvement in both web attack detection performance and uncertainty estimation quality compared to benchmark models.
This study constructs an explainable recommendation via combining product images, textual descriptions, user reviews, and user-item interactions. We propose a theory-based multimodal deep learning architecture. Specifically, we first extract the element-level features from product display information, including region-level visual and word-level textual features. To measure the impacts of these element-level features on user preferences, we introduce attention mechanisms. In addition, we inject the user reviews to disentangle the effects of visual and textual features on user preferences. Finally, we introduce a stick-breaking method to measure the asymmetrical influence of images and textual descriptions at the holistic level. To evaluate our model’s utility, we conduct experiments from four perspectives: recommendation performance, explanatory analysis, mechanism, and robustness. Experimental results show that our model can improve recommendation performance and give explanations. Our findings provide valuable insights for recommendation system design and offer guidance to marketers for optimizing product display pages
While consumer-generated reviews deliver substantial business intelligence value for advancing recommender systems, they also create an attack surface in review-based recommender systems (R-RSs). Subtle textual perturbations through review tampering, forgery, and other adversarial attacks can manipulate recommendation outcomes. Nevertheless, prior scholars primarily focus on numerical ratings or visual-input adversarial manipulations, offering limited guidance for R-RSs that rely on discrete and semantically interdependent textual reviews. Following adversarial robustness theory, we develop a computational design science framework that jointly assesses and enhances adversarial robustness for R-RSs. For assessment, we design a novel Shapley value-guided Adversarial Review Generation (SARG) method that mimics rigorous adversarial environments by generating adversarial reviews aligned with the manipulation objective. For enhancement, we design an Attack-Lifecycle Robustness Enhancement (ALRE) method aligned with the attack reconnaissance and execution lifecycle, integrating stochastic recommendation process to reduce reconnaissance-phase information leakage, sensitivity-aware input dropout with certified robustness bounds to reduce overreliance on high-sensitivity tokens and adversarial contrastive retraining to strengthen representation robustness during attack execution-phase. Extensive experiments on Amazon and Yelp datasets demonstrate the severity of adversarial vulnerabilities in existing R-RSs and the effectiveness of our framework across diverse attack settings. This study advances the IS literature by offering a unified approach to assessing and enhancing adversarial robustness in R-RSs, with broader implications for online platforms reliant on user-generated textual content.
Fake news on social media platforms poses a significant threat to societal systems, highlighting the urgent need for advanced detection methods. The existing detection methods can be divided into machine-intelligence-based, crowd-intelligence-based, and hybrid-intelligence-based methods. Among these, hybrid-intelligence-based methods achieve the best performance but fail to consider the uncertainty issue in detection. In light of this, we propose a novel uncertainty-aware hybrid-intelligence (UAHI) method for fake news detection. Our method comprises three integral modules. The first module employs a Bayesian deep learning model to capture the inherent uncertainty within machine intelligence. The second module uses an item response theory-based user response aggregation to account for the uncertainty in crowd intelligence. The third module introduces a new distribution fusion mechanism, which takes the distributions derived from both machine and crowd intelligence as input and outputs a fused distribution that provides predictions along with the associated uncertainty. Experiments on the two datasets demonstrate the advantages of our method. This study has practical implications for three key stakeholders: internet users, online platform managers, and the government.
This study improves tag recommendation for marketer-generated contents (MGCs) by integrating multimodal information from both textual descriptions and visual information. To address this issue, we develop a behavior-driven multimodal neural network that jointly captures content semantics and marketer-specific tagging preferences. Specifically, we first introduce a text-guided attention mechanism to model the interaction between the textual contents and images in MGCs. Second, we incorporate internal attention modules to identify the impacts of specific words, image regions, and individual images on tag prediction. Third, we design a time-decaying attention mechanism to account for marketers’ historical tagging behaviors, thereby capturing the temporal dynamics and heterogeneity in their preferences. We evaluate our model using a large-scale real-world dataset collected from Taobao. Quantitative results demonstrate that our approach significantly outperforms state-of-the-art baselines in tag recommendation. Qualitative analyses further reveal interpretable insights into how marketers’ historical behaviors and multimodal cues contribute to their tagging decisions.
This study focuses on multimodal topic modeling and attempts to separate public topics (shared across modalities) from private topics (unique to each modality) hidden in text and image data. To address this issue, we propose a novel Disentangled Multimodal Neural Topic Model (DMNTM). Specifically, we design the modality-specific encoder with an independence constraint to capture private topics, and the public encoder with a product-of-experts module to extract cross-modal shared topics. We conduct extensive experiments on six public datasets, including multimodal online reviews from Amazon, posts from Flickr, tweets from Twitter, and webpages from Wikipedia. Compared with state-of-the-art methods, we find that DMNTM significantly improves topic modeling performance in terms of perplexity, coherence, diversity, and topic quality over the best baseline. In two downstream tasks, including recommendation and sentiment classification, DMNTM further improves the performance. These results show that disentangling public and private topics effectively enhances both the quality and utility of multimodal representations.
Bundle recommendation is a widely used marketing strategy designed to suggest related items to users, providing users with one-stop convenience while increasing platform efficiency. Yet, aligning bundle recommendations with user preferences remains challenging due to the complex tripartite relationships among users, bundles, and items. This challenge is further exacerbated by sparse user–bundle interactions, where the large number of bundles and limited user engagements hinder accurate preference estimation and often result in suboptimal outcomes. To address these issues, we propose an advanced multi-graph contrastive framework powered by large language models (LLMs). The framework integrates multi-view expansion and LLM-based data augmentation to enrich interaction signals, while employing multi-graph contrastive learning to improve the accuracy and robustness of representation learning. Empirical evaluations demonstrate that our approach consistently outperforms state-of-the-art baselines across overall performance and varying sparsity conditions, underscoring its effectiveness across three benchmark real-world datasets. Moreover, the observed alignment between sparse and dense users with similar preferences suggests that the LLM-enhanced augmentation enables the model to learn meaningful representations even under sparse settings.
Large Language Models (LLMs) have brought unprecedented innovation opportunities to the marketing field. However, the practical applications of LLMs within the marketing landscape currently exhibit a fragmented and scattered nature. In this study, we aim to aggregate these scattered literature to create a holistic view of LLMs capabilities for marketing research. Specifically, we present an overview of LLMs using the evolution of LMs. Subsequently, we explore their application in the marketing domain across five distinct dimensions: data annotation, idea inspiration and content generation, substitution of human participants, user behavior learning and prediction, and evaluation of LLM feedback. Finally, we discuss the new trends and challenges for LLMs in marketing. This study enriches the theoretical foundations of integrating generative AI with marketing practices.
Music recommender systems play a critical role in music streaming platforms by providing users with music that they are likely to enjoy. Recent studies have shown that user emotions can influence users' preferences for music moods. However, existing emotion-aware music recommender systems (EMRSs) explicitly or implicitly assume that users' actual emotional states expressed through identical emotional words are homogeneous. They also assume that users' music mood preferences are homogeneous under the same emotional state. In this article, we propose four types of heterogeneity that an EMRS should account for: emotion heterogeneity across users, emotion heterogeneity within a user, music mood preference heterogeneity across users, and music mood preference heterogeneity within a user. We further propose a Heterogeneity-aware Deep Bayesian Network (HDBN) to model these assumptions. The HDBN mimics a user's decision process of choosing music with four components: personalized prior user emotion distribution modeling, posterior user emotion distribution modeling, user grouping, and Bayesian neural network-based music mood preference prediction. We constructed two datasets, called EmoMusicLJ and EmoMusicLJ-small, to validate our method. Extensive experiments demonstrate that our method significantly outperforms baseline approaches on metrics of HR, Precision, NDCG, and MRR. Ablation studies and case studies further validate the effectiveness of our HDBN. The source code and datasets are available at https://github.com/jingrk/HDBN.
The global proliferation of social media has provided a unique platform for cross-cultural exchange, greatly enhancing interactions between users from different cultural backgrounds through friend recommendation systems. However, the highly complex and intrinsically coupled nature of factors driving friendship formation makes it difficult for traditional methods to effectively predict and recommend genuinely deep social connections. Therefore, this study proposes leveraging emerging information technologies, specifically deep learning, to optimize and improve friend recommendation systems on social media platforms. This paper introduces a novel personality trait disentanglement method. By using large language models to extract personality factors from user text, we constructed a multi-subgraph convolutional method driven by personality traits. This enables the model to clearly distinguish the mechanisms of different personality factors. Additionally, we designed a shared attention layer to adaptively learn the importance weights of different personality traits, and implicit representations to capture non-personality-driven factors. Our research combines deep learning with personality trait analysis to foster deeper interpersonal understanding and cultural exchange, thereby enhancing the quality and breadth of interactions on social networks globally.
Federated recommender systems (FedRSs) effectively tackle the trade-off between recommendation accuracy and privacy preservation. However, recent studies have revealed severe vulnerabilities in FedRSs, particularly against untargeted attacks seeking to undermine their overall performance. Defense methods employed in traditional recommender systems are not applicable to FedRSs, and existing robust aggregation schemes for other federated learning-based applications have proven ineffective in FedRSs. Building on the observation that malicious clients contribute negatively to the training process, we design a novel contribution-aware robust aggregation scheme to defend FedRSs against untargeted attacks, named contribution-aware Bayesian knowledge distillation aggregation (ConDA), comprising two key components for the defense. In the first contribution estimation component, we decentralize the estimation from the server side to the client side and propose an ensemble-based Shapley value to enable the efficient calculation of contributions, addressing the limitations of lacking auxiliary validation data and high computational complexity. In the second contribution-aware aggregation component, we merge the decentralized contributions via a majority voting mechanism and integrate the merged contributions into a Bayesian knowledge distillation aggregation scheme for robust aggregation, mitigating the impact of unreliable contributions induced by attacks. We evaluate the effectiveness and efficiency of ConDA on two real-world datasets from movie and music service providers. Through extensive experiments, we demonstrate the superiority of ConDA over the baseline robust aggregation schemes.
Composed image retrieval (CoIR) involves a multi-modal query of the reference image and modification text describing the desired changes, allowing users to express image retrieval intents flexibly and effectively. The key of CoIR lies in how to properly reason the search intent from the multi-modal query. Existing work either aligns the composite embedding of the multi-modal query and the target image embedding in the visual domain through late-fusion or converts all images into text descriptions and leverage large language models (LLM) for text semantic reasoning. However, this single-modality reasoning approach fails to comprehensively and interpretably capture the users’ ambiguous and uncertain intents in the multi-modal queries, incurring the inconsistency between retrieved results and ground truth. Besides, the expensive manually annotated datasets limit the further performance improvement of CoIR. To this end, this article proposes an LLM-enhanced Intent Uncertainty-Aware Linguistic-Visual Dual Channel Matching Model (IUDC), which combines the strengths of multi-modal late-fusion and LLMs for CoIR. We first construct an LLM-based triplet augmentation strategy to generate more synthetic training triplets. Based on this, the core of IUDC consists of two matching channels: the semantic matching channel is responsible for intent reasoning on the aspect-level attributes extracted by an LLM, and the visual matching channel accounts for the fine-grained visual matching between multi-modal fusion embedding and target images. Considering the intent uncertainty presented in the multi-modal queries, we introduce Probability Distribution Encoder (PDE) to project the intents as probabilistic distributions in the two matching channels. Consequently, a mutually enhanced module is designed to share knowledge between the visual and semantic representations for better representation learning. Finally, the matching scores of two channels are added to retrieve the target image. Extensive experiments conducted on two real datasets demonstrate the effectiveness and superiority of our model. Notably, with the help of the proposed LLM-based triplet augmentation strategy, our model achieves a new record of state-of-the-art performance among all datasets.
Inferring consumers' preferences provides a better understanding of their purchase behavior, which is very important for business success, e.g., recommendation systems and targeted advertising. In this paper, we propose an explainable machine learning approach, namely Multi-view Latent Dirichlet Allocation (MVLDA), to infer and interpret consumer preferences. In the proposed model, we assume that there exists a downstream relationship between consumers' motivations and their purchase behaviors. We model this downstream relationship by linking two types of topics (i.e., textual topics for motivations and product-related topics for purchase behaviors), to quantify and explain consumers' choices. We validate our modeling framework using a real-world dataset collected from the online retailer Amazon. The experimental results show that the proposed model identifies a set of high-quality textual topics but also interprets its effect on consumer choices based on product-related topics. In addition, we demonstrate that the proposed model quantifies the consumer's preferences. The proposed model yields interesting insights about user preferences, and provides several important managerial implications, e.g., e-commerce platforms, brand managers, and marketers.
The advent of the internet has facilitated the wide spread of online disinformation and thus poses severe threats to the trustworthiness of cyberspace. Two types of methods are proposed to detect online disinformation: traditional machine learning-based and deep learning-based, where the former is limited due to the shallow representation and the latter is hindered by its lack of interpretability. In this study, we develop a novel model named interpretable wide and deep model for text (IWDMT) for disinformation detection which incorporates the interpretability benefits of traditional machine learning and the representation advantages of deep learning. Furthermore, we advance the interpretability of existing models by utilizing neural topic models to capture topical semantic representations and the attention mechanism to extract sequential syntactic representations. The proposed IWDMT is a mixture of a generative model and a discriminative model, and we devise a novel learning algorithm for it. Experiments on deceptive reviews and fraudulent emails demonstrated the proposed IWDMT not only outperformed baselines but was also able to provide a rich set of angles of interpretation for management insights. The higher accuracy and improved interpretability in detecting online disinformation will benefit four stakeholder groups: internet users, managers, researchers, and the government.
Grasping complementary relationships between fashion product pairings is gaining increasing attention in the e-commerce field. Current methods primarily utilize visual cues to assess compatibility, which, despite their efficacy, often lack sufficient explainability. Meanwhile, the rich semantic details embedded in product attributes remain largely unexplored. To tackle this, we propose a novel framework called Explainable Attribute-augmented Neural framework (EAN), which integrates comprehensive attribute and visual data, enabling explainability in fashion product compatibility modeling. We conduct quantitative and qualitative experiments to demonstrate the effectiveness and explainability of our proposed framework. The practical significance of our research is twofold. Firstly, it helps consumers understand the underlying reasons for fashion item pairings, thereby assisting them in refining their dressing combinations. Secondly, it provides novel perspectives for product design and assists e-commerce platforms in creating more effective product marketing combinations.
Professional generated content (PGC) serves as a vital and reliable online source that provides large-scale information about various aspects of brands and products. This study focuses on acquiring product-level competitive intelligence from large-scale PGCs. Specifically, we aim to simultaneously identify competitive relationships among products, extract representative topics shared by competing products, and estimate content preferences. To this end, we present a topic model that jointly leverages textual content and their associated product tags in PGCs. Owing to large-scale and lengthy PGCs, we propose a collapsed variational Bayesian inference algorithm to improve the model learning. We analyze over 100,000 PGCs and 3,000 associated products for empirical application in automobiles. Experimental results show that the proposed approach can accurately analyze market competition. Our findings have significant implications for product managers, enabling them to identify competitors, assess experts’ opinions on their products and competitors, and select high-quality content creators to improve promotions.
Images play a vital role in social media platforms, which can more vividly reflect people’s inner emotions and preferences, so visual sentiment analysis has become an important research topic. In this paper, we propose a Supervised Contrastive Learning-based model for image emotion classification, which consists of two modules of low-level feature extraction and deep emotional feature extraction, and feature fusion is used to enhance the overall perception of image emotions. In the low-level feature extraction module, the LBP-U (Local Binary Patterns with Uniform Patterns) algorithm is employed to extract texture features from the images, which can effectively capture the texture information of the images, aiding in the differentiation of images belonging to different emotion categories. In the deep emotional feature extraction module, we introduce a Supervised Contrastive Learning approach to improve the extraction of deep emotional features by narrowing the intra-class distance among images of the same emotion category while expanding the inter-class distance between images of different emotion categories. Through fusing the low-level and deep emotional features, our model comprehensively utilizes features at different levels, thereby enhancing the overall emotion classification performance. To assess the classification performance and generalization capability of the proposed model, we conduct experiments on the publicly FI (Flickr and Instagram) Emotion dataset. Comparative analysis of the experimental results demonstrates that our proposed model has good performance for image emotion classification. Additionally, we conduct ablation experiments to analyze the impact of different levels of features and various loss functions on the model’s performance, thereby validating the superiority of our proposed approach.