Affective Image Manipulation (AIM) aims to evoke specific emotions through targeted editing. Current image editing benchmarks primarily focus on object-level modifications in general scenarios, lacking the fine-grained granularity to capture affective dimensions. To bridge this gap, we introduce the first benchmark designed for AIM termed AIM-Bench. This benchmark is built upon a dual-path affective modeling scheme that integrates the Mikels emotion taxonomy with the Valence-Arousal-Dominance framework, enabling high-level semantic and fine-grained continuous manipulation. Through a hierarchical human-in-the-loop workflow, we finally curate 800 high-quality samples covering 8 emotional categories and 5 editing types. To effectively assess performance, we also design a composite evaluation suite combining rule-based and model-based metrics to holistically assess instruction consistency, aesthetics, and emotional expressiveness. Extensive evaluations reveal that current editing models face significant challenges, most notably a prevalent positivity bias, which stemming from inherent imbalances in training data distribution. To tackle this, we propose a scalable data engine utilizing an inverse repainting strategy to construct AIM-40k, a balanced instruction-tuning dataset comprising 40k samples. Concretely, we enhance raw affective images via generative redrawing to establish high-fidelity ground truths, and synthesize input images with divergent emotions and paired precise instructions. Fine-tuning a baseline model on AIM-40k yields a 9.15
Existing Multimodal Aspect-Based Sentiment Analysis techniques primarily focus on associating visual and textual content but often overlook the critical issue of visual expression deficiency, where images fail to provide complete aspect terms and sufficient sentiment signals. To address this limitation, we propose a Multimodal Aspect-Based Sentiment Analysis network that leverages Text-Dominant Speech Enhancement (TDSEN), aiming to alleviate the deficiency in visual expression by synthesizing speech and employing a text-dominant approach. Specifically, we introduce a Text-Driven Speech Enhancement Layer that generates speech with stable timbre to identify all aspect terms, compensate for the lacking parts of visual expression, and provide additional aspect term information and emotional cues. Meanwhile, we design a semantic distance mask matrix to enhance the capability of capturing key information from the textual modality. Furthermore, a text-driven multimodal feature fusion module is incorporated to strengthen the dominant role of text and facilitate multimodal feature interaction and integration for the extraction of the term of the aspect and sentiment recognition. Comprehensive evaluations on the Twitter-2015 and Twitter-2017 benchmarks demonstrate TDSEN's superiority, achieving absolute improvements of 2.6% and 1.7% over state-of-the-art baselines, with ablation studies confirming the necessity of each component.
The widespread dissemination and misleading impact of fake news on the web have become a significant concern for the public and the government. Discovering fake news is crucial for ensuring that users receive authentic information and maintaining social harmony. However, most existing entity-based fake news detection methods have two issues: i) methods for acquiring additional information through entities lack flexibility and real-time capabilities. ii) approaches using entities to capture news semantics have not adequately revealed the interactions between words in the text. To address these issues, we propose a Multi-graph Semantic-aware Adaptive Graph Convolutional Network (MgSAN), which comprehensively captures the semantic information of news texts by constructing multiple semantic graphs and learns the features from these graph structures using an adaptive graph convolutional network (SwiGCN). Specifically, we design a global semantic interaction graph to capture the complex interactions between words, generating a comprehensive textual semantic representation. We also employ an entity-noun relationship graph to mine deep semantic associations, enhancing the model's understanding of fine-grained textual deep meanings. Additionally, we develop an adaptive graph convolutional network to effectively extract and aggregate feature information from different graph structures. Finally, we introduce a fusion module to integrate both global and local fine-grained semantic information, forming a rich composite semantic representation, thereby improving the effectiveness of fake news detection. Extensive experimental results on three public benchmark datasets verify the effectiveness and superior performance of MgSAN, outperforming state-of-the-art detection models.
The Sparse Subgraph Finding (SGF) problem addresses the challenge of identifying sub-graphs with weak social interactions and sparse connections within a graph, which can be effectively modeled as discovering sparse subsystems in intelligent sensor networks. Traditional methods often rely on manually designed heuristics, which are computationally expensive and lack scalability, especially when dealing with complex sensor network systems. In this paper, we propose RL-SGF, a novel framework that integrates deep reinforcement learning and graph embedding through joint optimization to overcome these limitations. By simultaneously optimizing subsystem sparsity and representation learning within a unified framework, RL-SGF enhances both the effectiveness and robustness of the model in sensor network applications. Experimental results on synthetic and real-world datasets, including social networks, citation networks, and sensor network simulations, demonstrate that RL-SGF outperforms existing algorithms in terms of efficiency and solution quality, making it highly applicable to real-world sparse subsystem discovery scenarios in intelligent sensor networks.
Facial expression image editing requires fine-grained control to strictly preserve human identity and background while precisely manipulating expression. However, existing editing benchmarks primarily focus on general scenarios, lacking high-quality facial images and corresponding editing instructions. Furthermore, current evaluation metrics exhibit systemic biases in this task, often favoring lazy editing or overfit editing. To bridge these gaps, we propose FED-Bench, a comprehensive benchmark featuring rigorous testing and an accurate evaluation suite. First, we carefully construct a benchmark of 747 triplets through a cascaded and scalable pipeline, each comprising an original image, an editing instruction, and a ground-truth image for precise evaluation. Second, we introduce FED-Score, a cross-granularity evaluation protocol that disentangles assessment into three dimensions: Alignment for verifying instruction following, Fidelity for testing image quality and identity preservation, and Relative Expression Gain for quantifying the magnitude of expression changes, effectively mitigating the aforementioned evaluation biases. Third, we benchmark 18 image editing models, revealing that current approaches struggle to simultaneously achieve high fidelity and accurate expression manipulation, with fine-grained instruction following identified as the primary bottleneck. Finally, leveraging the scalable characteristic of introduced benchmark engine, we provide a 20k+ in-the-wild facial training set and demonstrate its effectiveness by fine-tuning a baseline model that achieves significant performance gains. Our benchmark and related code will be made publicly open soon.
Semi-supervised node classification reduces reliance on labeled data. While GNNs (e.g., GCN, GAT) exploit local structures, they often miss global topology; contrastive methods (e.g., MVGRL) face rigid embeddings or poor view design. We propose the Collaborative Graph Contrastive Network (CGCN) to jointly capture local and global semantics via a dual-encoder framework with cross-view alignment. A fine-grained collaborative constraint preserves node similarity by softly regularizing cross-view correlation matrices, avoiding over-compression. Additionally, a category-aware contrastive loss enhances class compactness and separation with minimal labels. Experiments on six benchmarks show CGCN outperforms state-of-the-art methods. Ablation studies validate each component and demonstrate robustness under complex topologies.
In real-world scenarios, a large amount of noise in user historical behaviors obstructs the reflection of their genuine interests. The long-tail distribution of user-item interactions also makes it difficult to capture interest evolution patterns from historical sequences. Moreover, as user behavior sequences continue to grow, solely relying on conventional sequence models is insufficient to extract user interest information and learn accurate sequence representations, thus limiting recommendation accuracy. To address these issues, we propose a self-supervised graph neural sequential recommendation model called LS4SRec, which disentangles users’ long and short-term interests. Specifically, LS4SRec constructs two independent interest encoders to extract users’ long and short-term interests. By utilizing the global user behavior sequence graph WITG to provide additional collaborative signals for each interaction sequence, we alleviate the issue of data sparsity. Subsequently, contrastive learning is applied to WITG to remove noise information and enhance the sequence representation. Further, interest allocation matrices and sequence models are utilized to model users’ interest evolution patterns. Finally, we introduce sequence graph data augmentation methods and long and short-term interest pseudo-label construction methods to generate unsupervised signals that assist in model training. Extensive experiments conducted on real-world data validate the effectiveness of our proposed model. Our model implementation codes are available at the link https://github.com/jiubaoyibao/LS4SRec.
BackgroundThe COVID-19 pandemic took a toll on everyone’s health and mental health professionals were no exception. This study examined the trajectory of the relationship between levels of physical fatigue and each of depression and anxiety in mental health professionals (MHPs) recovering from COVID-19.MethodsA national survey of 9,858 MHPs who had recovered from COVID-19 was conducted between January and February 2023. The nine-item Patient Health Questionnaire (PHQ-9), the 7-item Generalized Anxiety Disorder (GAD-7) scale, and a numerical rating scale were used to measure depression, anxiety and physical fatigue, respectively. Logistic regression with restricted cubic spline (RCS) models were created to examine the association of physical fatigue with depression and anxiety.ResultsThe prevalence of depression and anxiety in MHPs who recovered from COVID-19 infection were 47.0% (95%CI: 46.0-48.0%) and 28.9% (95%CI: 28.0-29.8%) respectively. The prevalence of moderate to severe physical fatigue was 44.2% (95%CI: 43.2-45.2%). The RCS models revealed a significant nonlinear relationship between physical fatigue and both depression and anxiety, with an inflection point at a fatigue score of 4. Above this threshold, the risk of both conditions increased significantly. Participants with poor perceived health and lower socioeconomic status had a significantly greater increase in depression and anxiety when fatigue levels were higher.ConclusionsModerate to severe physical fatigue was associated with depression and anxiety in MHPs recovering from COVID-19. Interventions aimed at alleviating fatigue may play a critical role in improving mental health outcomes in this vulnerable population.
Nowadays, short-form videos (SVs) are essential to web information acquisition and sharing in our daily life. The prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA) towards SVs. Considering the lack of SVs emotion data, we introduce a large-scale dataset named eMotions, comprising 27,996 videos. Meanwhile, we alleviate the impact of subjectivities on labeling quality by emphasizing better personnel allocations and multi-stage annotations. In addition, we provide the category-balanced and test-oriented variants through targeted data sampling. Some commonly used videos, such as facial expressions, have been well studied. However, it is still challenging to analysis the emotions in SVs. Since the broader content diversity brings more distinct semantic gaps and difficulties in learning emotion-related features, and there exists local biases and collective information gaps caused by the emotion inconsistence under the prevalently audio-visual co-expressions. To tackle these challenges, we present an end-to-end audio-visual baseline AV-CANet which employs the video transformer to better learn semantically relevant representations. We further design the Local-Global Fusion Module to progressively capture the correlations of audio-visual features. The EP-CE Loss is then introduced to guide model optimization. Extensive experimental results on seven datasets demonstrate the effectiveness of AV-CANet, while providing broad insights for future works. Besides, we investigate the key components of AV-CANet by ablation studies. Datasets and code will be fully open soon.
Multimodal Aspect-Based Sentiment Analysis (MABSA) plays a pivotal role in the advancement of sentiment analysis technology. Although current methods strive to integrate multimodal information to enhance the performance of sentiment analysis, they still face two critical challenges when dealing with multi-aspect and multi-sentiment data: i) the importance of aspect terms within multimodal data is often overlooked, and ii) models fail to accurately associate specific aspect terms with corresponding sentiment words in multi-aspect and multi-sentiment sentences. To tackle these problems, we propose a novel multimodal aspect-based sentiment analysis method that combines Aspect Enhancement and Text Simplification (AETS). Specifically, we develop an aspect enhancement module that boosts the ability of model to discern relevant aspect terms. Concurrently, we employ text simplification module to simplify and restructure multi-aspect and multi-sentiment texts, accurately capturing aspects and their corresponding sentiments while reducing irrelevant information. Leveraging this method, we perform three tasks including multimodal aspect term extraction, multimodal aspect sentiment classification, and joint multimodal aspect-based sentiment analysis. Experimental results indicate that our proposed AETS model achieved state-of-the-art performance on two benchmark datasets.
Advances in Generative AI have made video-level deepfake detection increasingly challenging, exposing the limitations of current detection techniques. In this paper, we present HOLA, our solution to the Video-Level Deepfake Detection track of 2025 1M-Deepfakes Detection Challenge. Inspired by the success of large-scale pre-training in the general domain, we first scale audio-visual self-supervised pre-training in the multimodal video-level deepfake detection, which leverages our self-built dataset of 1.81M samples, thereby leading to a unified two-stage framework. To be specific, HOLA features an iterative-aware cross-modal learning module for selective audio-visual interactions, hierarchical contextual modeling with gated aggregations under the local-global perspective, and a pyramid-like refiner for scale-aware cross-grained semantic enhancements. Moreover, we propose the pseudo supervised singal injection strategy to further boost model performance. Extensive experiments across expert models and MLLMs impressivly demonstrate the effectiveness of our proposed HOLA. We also conduct a series of ablation studies to explore the crucial design factors of our introduced components. Remarkably, our HOLA ranks 1st, outperforming the second by 0.0476 AUC on the TestA set.
Multivariate time series anomaly detection (MTSAD) aims to accurately identify and localize complex abnormal patterns in the large-scale industrial control systems. While existing approaches excel in recognizing the distinct patterns under the low-dimensional scenarios, they often fail to robustly capture long-range spatiotemporal dependencies when learning representations from the high-dimensional noisy time series. To address these limitations, we propose DARTs, a robust long short-term dual-path framework with window-aware spatiotemporal soft fusion mechanism, which can be primarily decomposed into three complementary components. Specifically, in the short-term path, we introduce a Multi-View Sparse Graph Learner and a Diffusion Multi-Relation Graph Unit that collaborate to adaptively capture hierarchical discriminative short-term spatiotemporal patterns in the high-noise time series. While in the long-term path, we design a Multi-Scale Spatiotemporal Graph Constructor to model salient long-term dynamics within the high-dimensional representation space. Finally, a window-aware spatiotemporal soft-fusion mechanism is introduced to filter the residual noise while seamlessly integrating anomalous patterns. Extensive qualitative and quantitative experimental results across mainstream datasets demonstrate the superiority and robustness of our proposed DARTs. A series of ablation studies are also conducted to explore the crucial design factors of our proposed components. Our code and model will be made publicly open soon.
INTRODUCTION:Based on the Chinese Longitudinal Healthy Longevity Survey (CLHLS), this study aimed to examine the prevalence and correlates of depression, and its network structure and association with quality of life (QOL) in older adults with hypertension. METHODS:Depression and QOL were measured using the 10-item Center for Epidemiologic Studies Short Depression Scale (CESD-10) and the World Health Organization Quality of Life-brief version, respectively. Univariable and multivariable analyses were performed. Network analysis was used to explore the interconnections between depressive symptoms. The flow function was used to identify depressive symptoms that were directly associated with QOL. RESULTS:A total of 5032 older adults with hypertension were included. The prevalence of depression (CESD-10 total score ≥ 10) was 28.3% (95% confidence interval: 27.08%-29.59%), which was significantly associated with poor QOL (P < 0.001). Participants who were male (P < 0.001), resided in urban areas (P = 0.006), lived with their family (P < 0.001), had perceived fair or good economic status (P < 0.001), and higher level of instrumental activities of daily living (P < 0.001) had lower risk of depression. In the network model of depression, CESD3 'Feeling blue/depressed', CESD4 'Everything was an effort' and CESD8 'Loneliness' were the most central symptoms. CESD10 'Sleep disturbances' had the highest negative association with QOL, followed by CESD5 'Hopelessness', and CESD7 'Lack of happiness'. CONCLUSION:Depression was common among older adults with hypertension and significantly associated with poor QOL. To prevent and reduce the negative impact of depression in this population, appropriate interventions should target both central symptoms and the depressive symptoms associated with QOL.
ObjectivesCognitive impairment is a major health concern in older adults with hypertension, and both depression and abnormal sleep duration are recognized as potential contributing factors. This study aimed to explore the nonlinear association of depression and sleep duration with cognitive impairment among older adults with hypertension.MethodsThis cross-sectional study was based on the 2017–2018 wave of Chinese Longitudinal Healthy Longevity Survey. Depression and cognitive function were measured using the 10-item Center for Epidemiological Studies Short Depression Scale and Mini Mental State Examination, respectively. Univariate, binary logistic regression, and restricted cubic spline regression analyses were used to examine the associations between depression, sleep duration and cognitive impairment.ResultsA total of 3,989 older adults with hypertension were included. The prevalence of depression and cognitive impairment were 28.1% (95%CI = 26.7–29.5%) and 10.1% (95%CI = 9.2–11.1%), respectively. After adjusting for confounding factors, a significant linear association (nonlinear p = 0.814) between depression and cognitive impairment risk was found, while a U-shaped nonlinear association was identified between sleep duration and cognitive impairment risk (p = 0.040). Both shorter (<6.6 h) and longer (>7.7 h) sleep duration per day were associated with higher cognitive impairment risk, with an inflection point at 7.3 h. The effect of sleep duration on cognitive impairment risk was more significant for participants with a higher (≥ 6 years) education level.ConclusionThis study highlights the importance of managing depression and optimizing sleep duration in addressing the risk of cognitive decline in older adults with hypertension.
In recent years, multimodal agent AI (MAA) has emerged as a pivotal area of research, holding promise for transforming human-machine interaction. Agent AI systems, capable of perceiving and responding to inputs from multiple modalities (e.g., language, vision, audio), have demonstrated remarkable progress in understanding complex environments and executing intricate tasks. This survey comprehensively reviews the state-of-the-art developments in MAA and examines its fundamental concepts, key techniques, and applications across diverse domains. We first introduce the basics of agent AI and its multimodal interaction capabilities. We then delve into the core technologies that enable agents to perform task planning, decision-making, and multi-sensory fusion. Furthermore, we focus on exploring various applications of MAA in robotics, healthcare, gaming, and beyond. Additionally, we mainly focus on analyzing the challenges and limitations of current systems and propose promising research directions for future improvements, including human-AI collaboration, online learning method improvement. By reviewing existing work and highlighting open questions, this survey aims to provide a comprehensive roadmap for researchers and practitioners in the field of MAA.
Multimodal aspect-based sentiment analysis aims to extract aspects from different data sources and recognize the corresponding sentiments. While current research has broadly focused on syntax relation-driven semantic comprehension, the impact of the importance of different syntactic relations on semantic understanding has not been adequately investigated. To address this issue, we propose a Sentiment-enhanced Multi-hop Connected Graph Attention Network (MCG), aiming to enhance the discriminative capability of model for sentiments and to delve into the syntactic relationships within the text. Firstly, we design a contrastive sentiment-enhanced pre-training task that expands the diversity and complexity of training samples to improve the recognition of multiple sentiments. Secondly, we construct a multi-hop connected syntactic dependency graph to deeply explore the rich syntactic dependencies in the text and to reveal the differences among syntactic relations. Moreover, we develop a multi-hop connected graph attention mechanism that enables the model to focus on the key syntactic relations within the syntactic structure, thereby enhancing the comprehension and predictive capabilities of model in multimodal sentiment analysis. Experimental results on two benchmark datasets demonstrate that our method outperforms state-of-the-art methods. The source code is provided in the supplementary materials.
Contrastive deep graph clustering is a graph data clustering method that combines deep learning and contrastive learning, aiming to realize accurate clustering of graph nodes. Although existing hard sample mining-based methods show positive results, they face the following challenges: (1) Clustering methods based on data similarity may lose important intrinsic groupings and association patterns, thereby affecting a comprehensive understanding and interpretation of the data. (2) In the measurement of hard samples, ignoring hard boundary samples may exacerbate clustering bias. To address these issues, we propose a new contrastive deep graph clustering method called Hard Boundary Sample Aware Network (HBSAN), which introduces attribute and structure enhanced encoding and generalized dynamic hard boundary sample weighting modulation strategy. Specifically, we optimize the similarity computation among samples through adaptive attribute embedding and multiview structure embedding techniques to deeply explore the intrinsic connections among samples, thereby aiding in the measurement of hard boundary samples. Furthermore, we leverage the unreliable confidence information obtained from initial clustering analysis to design an innovative hard boundary sample weight modulation function. This function first identifies the hard boundary samples and then dynamically reduces their weights, effectively enhancing the discriminative capability of network in ambiguous classification scenarios. Combining extensive experimental evaluations and in-depth analysis, our approach achieves state-of-the-art performance and establishes superior results in handling complex network clustering tasks.
Given the heavy responsibilities placed on inflight security officers (IFSO) to ensure passenger safety and eliminate inflight hazards, they often turn to Internet use to cope with their work pressure. This study examined the prevalence of internet addiction (IA) among IFSO in China, and its associated factors, relationship with quality of life (QOL), and network structure. This was a cross-sectional study based on a national survey. Expected influence (EI) was used to identify the most central nodes within the network model. Among 3,475 IFSO included in this study across 10 airlines, the prevalence of IA (IAT-20 total score of ≥ 50) was 13.1
Qinbao Song (宋擒豹)合作论文数Faculty of Electronic and Information Engineering, Xi'an Jiaotong University13