Brand competitive analysis is crucial for businesses to understand their market position and develop strategies to outperform competitors. As consumer behavior and market trends evolve, companies increasingly rely on data-driven insights to stay ahead. In this context, online reviews on E-commerce platforms offer a wealth of information, providing valuable opportunities for analyzing consumer preferences. With the rise of multimodal reviews (combining both text and images), it has become essential to harness these diverse data sources effectively. To address this need, we first propose a multimodal and multi-corpora latent Dirichlet allocation (MM-LDA) model for static data. It uncovers both general topics and brand-specific topics, enabling fine-grained insights into consumer concerns and brand positioning. We further extend MM-LDA into a dynamic framework and propose the multimodal and multi-corpora dynamic topic model (MM-DTM), which incorporates temporal dynamics to capture topic evolution over time. The effectiveness of our approaches are demonstrated using two datasets from ZOL and Baidu Tieba discussions.
Survival analysis problems are crucial in business area. Most of the existing research conducts business survival problem with structured data, yet overlooks the potential rich information provided in text data. Therefore, we creatively extract useful information from online reviews and study the influence of textual features on the occurrence of a business event. We propose a novel dynamic survival topic model (DSTM) that first extract topic proportions from text corpus as topic features in each time slice, then estimate the coefficients each feature contributes to the hazard ratio under time-dependent circumstances. Experiments on two real-word business datasets show that our proposed model can not only identify the true event time in the corpus and how the features influence the event, but also outperform two baseline models in six evaluation metrics. Our findings provide significant implications for decision makers to understand users’ personal attitude towards corresponding event behind their corpus and personal features.
Rating prediction is an important task in review mining. Aspect category sentiment analysis (ACSA) and overall rating prediction (ORP) are closely related tasks, yet existing joint models are typically limited to text-only inputs, while existing multimodal review models usually address only a single task. To bridge this gap, we propose MMRP, an end-to-end multimodal multitask framework for jointly modeling ACSA and ORP from image-text reviews. The model integrates textual and visual information through a cross-modal Transformer and further introduces modality-aware learnable positional encoding and a gated multi-head cross-modal attention mechanism to better handle the structure of multimodal reviews, where text forms token sequences and images appear as variable-sized sets. We evaluate MMRP on the ZOL mobile review dataset and an additional cross-domain multimodal benchmark, ViMACSA. Extensive experiments, including comparative evaluation, ablation analysis, sensitivity analysis, and case studies, show that MMRP consistently achieves strong performance and exhibits good robustness, cross-domain applicability, and practical deployment potential. These results demonstrate the effectiveness of jointly leveraging multimodal information for aspect-level and overall rating prediction in user reviews.
A smoothed maximum rank correlation (MRC) estimator for ordinal choice models is introduced, combining a linear function with a nonlinear component modeled by deep neural networks to achieve both identifiability and interpretability. A two-step estimation algorithm is designed that maintains the order relations among outputs without relying on the parallelism assumption, making it appealing in practical applicability. The statistical properties of the smoothed MRC estimator are established under regular conditions, including identification, convergence rate, and minimax optimality, while allowing the number of categories to increase with sample size. Our theoretical results extend beyond ordinal choice models and apply to a broad range of generalized regression models. Extensive simulations demonstrate the superiority of the proposed method in classification accuracy and interpretability. Its effectiveness is further validated through applications to twelve benchmark datasets and an online education dataset.
Deep trajectory modeling has garnered significant attention across various applications, particularly in trajectory user linking (TUL), which aims to associate trajectories with specific users by analyzing complex mobility patterns. Despite its advancements, the lack of explainability remains a critical challenge. In this article, we propose a general Information Bottleneck framework, TUL-IB, designed to enhance the explainability of TUL models for both sequence and graph data, with tractable optimization bounds to solve the TUL-IB objective. We further demonstrate that TUL-IB can be effectively applied to two distinct types of trajectory data: (1) waypoint trajectories, for which we extend TUL-IB into a dual-view approach, TUL-DV-IB, integrating both driving behavior sequences and trajectory road graphs. To ensure temporal continuity in subsequence selection, we employ dynamic programming during post-processing; (2) staypoint trajectories, for which we adapt TUL-IB to the graph node level and apply it to global trajectory graph model, resulting in TUL-GTG-IB. This adaptation identifies key neighboring trajectories that significantly contribute to explaining the user-linking results. Experimental results on three real-world datasets demonstrate that our method outperforms existing explainable approaches, providing deeper insights into trajectory user-linking models.
Over the past twenty years, topic modeling has gradually become popular as a powerful tool, extracting useful and meaningful latent representations from large texts. Research on topic evolution, focusing on the representation of changes in topics over time, has begun to attract extensive attention in the fields of information retrieval and data mining. The dynamic topic model is a classical model for topic evolution. It assumes all topics exist throughout the entire time period, overlooking the fact that topics that were previously important are no longer considered, and new topics can also emerge. To address this issue, we propose a novel Bayesian sparse dynamic topic model, utilizing a spike-and-slab prior distribution to capture topic birth and death. The results demonstrate that our proposed model can effectively estimate both the topic distribution and topic sparsity at the same time. Furthermore, simulations and empirical studies on two real-world datasets demonstrate that our proposed model outperforms the classical dynamic topic model and provides rich semantic information on focused topics.
Financial fraud detection is essential to safeguard billions of dollars, yet the intertwined entities and fast-changing transaction behaviors in modern financial systems routinely defeat conventional machine learning models. Recent graph-based detectors make headway by representing transactions as networks, but they still overlook two fraud hallmarks rooted in time: (1) temporal motifs–recurring, telltale subgraphs that reveal suspicious money flows as they unfold–and (2) account-specific intervals of anomalous activity, when fraud surfaces only in short bursts unique to each entity. To exploit both signals, we introduce ATM-GAD, an adaptive graph neural network that leverages temporal motifs for financial anomaly detection. A Temporal Motif Extractor condenses each account's transaction history into the most informative motifs, preserving both topology and temporal patterns. These motifs are then analyzed by dual-attention blocks: IntraA reasons over interactions within a single motif, while InterA aggregates evidence across motifs to expose multi-step fraud schemes. In parallel, a differentiable Adaptive Time-Window Learner tailors the observation window for every node, allowing the model to focus precisely on the most revealing time slices. Experiments on four real-world datasets show that ATM-GAD consistently outperforms seven strong anomaly-detection baselines, uncovering fraud patterns missed by earlier methods.
Trajectory anomaly detection is critical in trajectory data mining. The objective is to identify abnormal movements of objects. Most existing trajectory anomaly detection methods focus on determining whether an entire trajectory is anomalous, lacking the ability to identify the exact anomalous sub-trajectories. Although recent research has started addressing anomalous sub-trajectories detection, these methods fail to extract the specific route pattern for the target trajectory. As a result, they struggle to identify anomalous sub-trajectories when the same sub-trajectory is regarded as normal in other routes. To overcome these limitations, we propose a Route-Enhanced Conditional Anomalous Sub-Trajectory detection model (RECAST). RECAST has two innovative components: (1) a Route Discovery Network (RDN) that extracts the normal route pattern of the given trajectory; (2) a Conditional Anomalous Sub-trajectory Detection (CASD) network that detects anomalies conditioned on the estimated route patterns. Our design enables RECAST to identify sub-trajectories as anomalous even if they are normal in other routes, as long as they are unlikely to occur in the route of the given trajectory. We evaluate the effectiveness and efficiency of RECAST using two real-world datasets. The results demonstrate that our method outperforms the state-of-the-art methods in detection accuracy with competitive runtime efficiency(1).
Massive trajectory data has originated from the development of positioning technology. Learning GPS trajectory representation to characterize a driver’s driving style is a challenging task with important applications in many areas, including autonomous driving, auto insurance, advanced driver assistance systems, urban computing, and the internet of things. Few studies have considered the interactions between different factors. In this study, we propose a novel trajectory representation method based on a multilevel attention mechanism (ATTraj2vec) and apply it to the task of driver identification. We use 1D CNN to summarize local features, including motion, spatial and temporal features. Then we utilize a multilevel attention mechanism to extract global features, aggregating the interactions of motion features with temporal and spatial features progressively. Additionally, we adopt multi-loss to optimize our model simultaneously, which consists of a softmax loss for driver classification and Siamese loss for making trajectories from the same driver more similar. Classification experimental results on a real-world automobile trajectory dataset demonstrate that our proposed model significantly outperforms existing baselines. Meanwhile, the proposed method provides significant gains in the trajectory clustering of unseen drivers. The code is available in the repository, https://github.com/limengyuan2021/ATTraj2vec,
The prevalence of multimodal data has become commonplace in e-commerce platforms. Both seller showcases (i.e., the seller’s show) and user-generated content (i.e., the buyer’s show) now incorporate diverse modalities, combining both textual and visual elements. In this work, we aim to unraveling the impact of seller’s show on buyer’s show through the anchoring effect. We narrow our research on the specific problem of review helpfulness prediction and further explore whether the anchoring effect can improve the prediction accuracy of review helpfulness. In pursuit of this goal, we develop the Multi-granularity Attention Network Model based on Anchoring Effect (MAN-AE). This model first extracts the multi-granularity features in both seller’s show and buyer’s show and then accounts for the anchoring effect through a cross-source transformer. Through extensive experiments on an Amazon dataset, we demonstrate the anchoring effect of seller’s show on buyer’s show in enhancing the review helpfulness prediction performance. In comparison with other state-of-the-art models, our model demonstrates significantly superior prediction performance.
Detecting knowledge emerging trends has received increasing attention. It can help researchers understand the history of the discipline and predict future research hotspots. Dynamic topic models can be used to identify knowledge emerging trends from academic papers. However, traditional dynamic topic models have some shortcomings, such as over-assumptions, insufficient topic distinction, and high computational cost. To address this problem, we propose a relevance-based dynamic thin topic model (RBDTTM). We model topic evolution with a Gaussian process and adopt a relevance-based mechanism on topic-word distributions. Under this assumption, only words relevant to a certain topic can be represented. This relevance-based mechanism can not only decrease the number of parameters to be estimated but also achieve more prominent and focused topics. We evaluate the estimation performance of RBDTTM using a series of experiments on synthetic data. Results show that RBDTTM has greater interpretability and generalization than its competitors. Finally, we take the statistics discipline as an example and apply RBDTTM to two corpora of journal articles and a Chinese graduation thesis to explore the emerging statistical knowledge trend in the past two decades.
This work addresses a key limitation in current federated learning approaches, which predominantly focus on homogeneous tasks, neglecting the task diversity on local devices. We propose a principled integration of multi-task learning using multi-output Gaussian processes (MOGP) at the local level and federated learning at the global level. MOGP handles correlated classification and regression tasks, offering a Bayesian non-parametric approach that naturally quantifies uncertainty. The central server aggregates the posteriors from local devices, updating a global MOGP prior redistributed for training local models until convergence. Challenges in performing posterior inference on local devices are addressed through the Polya-Gamma augmentation technique and mean-field variational inference, enhancing computational efficiency and convergence rate. Experimental results on both synthetic and real data demonstrate superior predictive performance, OOD detection, uncertainty calibration and convergence rate, highlighting the method's potential in diverse applications. Our code is publicly available at https://github.com/JunliangLv/task_diversity_BFL.
A Point-of-Interest (POI) refers to a specific place of potential interest in location-based system. Next POI recommendation based on Large Language Models (LLMs) transforms the recommendation task into a question-answering task to predict the next waypoint along a trajectory to enhance user experience. While existing research has made preliminary attempts, there are inherent limitations: (1) Without sufficient exploration of POI, user and temporal similarities between trajectories in the historical data, previous studies may fail to include the next POI in the prompt, inherently limiting the ability of LLMs to make accurate predictions; (2) Existing methods do not account for candidates derived from various travel semantics and employs a sample-based generation strategy, which provides duplicate outcomes. To address these issues, we propose an LLM-based model named STrajRAG. Our model introduces a supervised auxiliary task to facilitate the identification of trajectories including the next POI, integrating heterogeneous similarities such as user and time. We consider candidates based on spatial distances and overall transition frequencies, employing beam search to generate diverse outcomes. Extensive experiments on three datasets demonstrate that STrajRAG achieves a 3%–32% performance improvement across diverse metrics compared to existing state-of-the-art methods.
Identifying change points in dynamic text data is crucial for understanding the evolving nature of topics across various sources, such as news articles, scientific papers, and social media posts. While topic modeling has become a widely used technique for this purpose, capturing fine-grained shifts in individual topics over time remains a significant challenge. Traditional approaches typically use a two-stage process, separating topic modeling and change point detection. However, this separation can lead to information loss and inconsistency in capturing subtle changes in topic evolution. To address this issue, we propose TOPIC-PYP, a change point detection model specifically designed for fine-grained topic-level analysis, i.e., detecting change points for each individual topic. By leveraging the Pitman-Yor process, TOPIC-PYP effectively captures the dynamic evolution of topic meanings over time. Unlike traditional methods, TOPIC-PYP integrates topic modeling and change point detection into a unified framework, facilitating a more comprehensive understanding of the relationship between topic evolution and change points. Experimental evaluations on both synthetic and real-world datasets demonstrate the effectiveness of TOPIC-PYP in accurately detecting change points and generating high-quality topics.
Graph-based models have emerged as a powerful paradigm for modeling multimodal urban data and learning region representations for various downstream tasks. However, existing approaches face two major limitations. (1) They typically employ identical graph neural network architectures across all modalities, failing to capture modality-specific structures and characteristics. (2) During the fusion stage, they often neglect spatial heterogeneity by assuming that the aggregation weights of different modalities remain invariant across regions, resulting in suboptimal representations. To address these issues, we propose MTGRR, a modality-tailored graph modeling framework for urban region representation, built upon a multimodal dataset comprising point of interest (POI), taxi mobility, land use, road element, remote sensing, and street view images. (1) MTGRR categorizes modalities into two groups based on spatial density and data characteristics: aggregated-level and point-level modalities. For aggregated-level modalities, MTGRR employs a mixture-of-experts (MoE) graph architecture, where each modality is processed by a dedicated expert GNN to capture distinct modality-specific characteristics. For the point-level modality, a dual-level GNN is constructed to extract fine-grained visual semantic features. (2) To obtain effective region representations under spatial heterogeneity, a spatially-aware multimodal fusion mechanism is designed to dynamically infer region-specific modality fusion weights. Building on this graph modeling framework, MTGRR further employs a joint contrastive learning strategy that integrates region aggregated-level, point-level, and fusion-level objectives to optimize region representations. Experiments on two real-world datasets across six modalities and three tasks demonstrate that MTGRR consistently outperforms state-of-the-art baselines, validating its effectiveness.
Graph learning for urban region modeling has gained significant attention for leveraging multi-modal data to generate region representations for downstream task prediction. However, existing models face two key limitations: (1) they primarily adopt a global perspective, overlooking the joint modeling of both local and global aspects, and (2) they rely on redundant, low-information nodes, leading to suboptimal region representations. To address these challenges, we propose GraphJCL, a dual-perspective framework that models both local and global perspectives. Specifically, GraphJCL first constructs local graphs for individual regions and a global graph encompassing all regions, integrating POI, taxi flow, remote sensing, street view, and road network data. Additionally, GraphJCL employs specialized message-passing mechanisms to efficiently capture both local and global graph node representations. Furthermore, GraphJCL incorporates entropy-optimized graph node pruning, retaining only the most informative nodes to enhance final region representations. To ensure the effectiveness of the designed dual-perspective graph framework, GraphJCL introduces a joint contrastive learning approach, optimizing region representations through geography-driven, entropy-optimized, and mutual information-based optimization techniques. Extensive experiments on two real-world datasets across five modalities demonstrate that GraphJCL consistently outperforms state-of-the-art methods on three tasks, validating its flexibility and effectiveness.
The prevalent issue in urban trajectory data usage, notably in low-sample rate datasets, revolves around the accuracy of travel time estimations, traffic flow predictions, and trajectory similarity measurements. Conventional methods, often relying on simplistic mixes of static road networks and raw GPS data, fail to adequately integrate both network and trajectory dimensions. Addressing this, the innovative GRFTrajRec framework offers a graph-based solution for trajectory recovery. Its key feature is a trajectory-aware graph representation, enhancing the understanding of trajectory-road network interactions and facilitating the extraction of detailed embedding features for road segments. Additionally, GRFTrajRec's trajectory representation acutely captures spatiotemporal attributes of trajectory points. Central to this framework is a novel spatiotemporal interval-informed seq2seq model, integrating an attention-enhanced transformer and a feature differences-aware decoder. This model specifically excels in handling spatiotemporal intervals, crucial for restoring missing GPS points in low-sample datasets. Validated through extensive experiments on two large real-life trajectory datasets, GRFTrajRec has proven its efficacy in significantly boosting prediction accuracy and spatial consistency.
The Bayesian two-step change point detection method is popular for the Hawkes process due to its simplicity and intuitiveness. However, the non-conjugacy between the point process likelihood and the prior requires most existing Bayesian two-step change point detection methods to rely on non-conjugate inference methods. These methods lack analytical expressions, leading to low computational efficiency and impeding timely change point detection. To address this issue, this work employs data augmentation to propose a conjugate Bayesian two-step change point detection method for the Hawkes process, which proves to be more accurate and efficient. Extensive experiments on both synthetic and real data demonstrate the superior effectiveness and efficiency of our method compared to baseline methods. Additionally, we conduct ablation studies to explore the robustness of our method concerning various hyperparameters.
Aspect category sentiment analysis (ACSA) on user reviews is a fundamental and challenging task. It aims to identify all the aspect categories mentioned in the reviews and their corresponding sentiment polarities. In multimodal data, both text and image are highly associated with aspect sentiment. How to utilize detailed multimodal data to further enhance the effectiveness of fine-grained aspect-category sentiment analysis tasks is a highly worthy research issue. Most of existing works integrate text and visual information via attention mechanisms at a single level, neglecting the fact that, information contained in each modality can be divided into different levels. To address these issues, we propose a hierarchical classification modeling approach that jointly models attribute detection tasks and attribute sentiment classification tasks, and introduce a multimodal joint model (MJM) for aspect-category sentiment analysis. The MJM model utilizes text information from the word level and context level, as well as image information from the global level, scene level, and local level, to fully mine detailed features in multimodal scenes. The proposed model is evaluated on the MASAD dataset. It achieves higher performance compared to the baseline models across various evaluation indicators, as well as different ablation experiments.
Aspect-based sentiment analysis (ABSA) aims to conduct fine-grained sentiment analysis, necessitating the extraction of three key components: target entity, aspect category and sentiment polarity. These three components collectively form an integrated ABSA task known as TASD (Target-Aspect-Sentiment jointly Detection). Most of existing approaches on ABSA usually employ Recurrent neural networks(RNNs), Convolutional neural networks(CNNs) or pre-training models such as Bidirectional Encoder Representations from Transformers(BERT). However, they have some common weaknesses. First, most of the existing methods focus on one or two sub-tasks instead of triplet detection, thus they do not establish an end -to -end (training a complex learning system represented by a single model) ABSA model and cannot utilize the relevance of multiple ABSA sub-tasks during training. Second, they cannot achieve accuracy and efficiency simultaneously due to the coupling of context and given aspects. Third, they are poor in recognizing implicit targets. To tackle these limitations, this paper proposes a novel method named the Twin Towers End to End model (TTEE) to solve TASD task. It transforms complex TASD task into a simple end -to -end multi -task framework, simultaneously conducting target and aspect-sentiment detection. It builds twin towers system based on BERT or its updated versions to decouple context and given aspects, which can reduce redundant calculation to improve computational efficiency significantly. It offers great advantage to identify implicit target entity and its associated aspect-sentiment in the context without introducing extra model architecture. Experiments on three public datasets in different domains demonstrate that our approach not only achieves better performance on various evaluation metrics, but also has high efficiency in both training and inference phases, over a wide range of sample size and number of aspect categories.