
Osteoporosis has emerged as a significant global public health challenge, affecting more than 200 million people worldwide. Precise screening through bone mineral density (BMD) assessment is essential for timely intervention. Although dual-energy X-ray absorptiometry (DXA) serves as the gold standard for BMD evaluation, its widespread application is hampered by high costs and limited accessibility. In recent years, deep learning has demonstrated great potential in osteoporosis classification and predicting BMD from X-ray and CT scans. However, existing public datasets have notable limitations, particularly the absence of paired multimodal imaging datasets for cross-modal modeling and the lack of accurate BMD annotations required for osteoporosis-specific research and clinical model optimization. To address these issues, we introduce the LUmbar Multimodal Osteoporosis Screening dataset (LUMOS), the first multimodal dataset specifically designed for lumbar osteoporosis screening. LUMOS integrates clinical data from 803 patients, including 1,620 anteroposterior/lateral lumbar X-rays with BMD values and T-scores, comprehensive demographic information, and 280 lumbar CT scans. The advent of LUMOS is expected to propel forward research on automated osteoporosis classification, BMD prediction, and other related tasks. Its standardized and multimodal nature fills critical gaps in lumbar osteoporosis data, providing high-quality data support for the development and validation of medical AI algorithms in the early detection of osteoporosis. The dataset is available at https://keyueshi.github.io/LUMOS/.
Amidst the swift advancement of 3D vision technology, Multi-view Compression (MVC) has become a crucial technique, widely applied in fields such as virtual reality, augmented reality, autonomous driving, telemedicine, and security surveillance. The technology effectively handles views from multiple cameras, utilizing the inter-view correlations to compress data efficiently. It substantially decreases the data transmission and storage requirements, enabling a richer and more realistic visual experience within the same bandwidth constraints. To further enhance compression performance, new methods continue to emerge. However, the absence of a unified benchmark testing library capable of effectively evaluating existing algorithms poses significant challenges to the further development of the field and the practical deployment of algorithms. To address this issue, we introduce OpenMVC, an Open-Source Library for Learning-based Multi-view Compression. We provide a comprehensive description and analysis of the performance advantages of existing algorithms. Furthermore, we conduct extensive and comprehensive benchmark testing of nine typical algorithms in the last five years, evaluating them in a consistent environment across various metrics. The open-source library for OpenMVC is available at https://openi.pcl.ac.cn/OpenAICoding/OpenMVC.
Emotional Support Chatbots could unlock potential by providing scalable, low-cost, and personal emotional support, overcoming critical accessibility barriers inherent in traditional counseling. However, current text-based Chatbots fall short in conveying the multimodal empathy crucial in counseling. Humans naturally prefer face-to-face communication with peers to share feelings, encompassing spoken tone, micro-expressions, and body language to convey empathy. To bridge this gap, we propose EMO-Avatar, an LLMagent-orchestrated framework that integrates emotional reasoning capabilities and multimodal expression in counseling. Our approach introduces two innovations: (1) a Multimodal Emotional Support Agent. EMO-Avatar can follow adaptive instruction across TTS, pose, micro-expressions, and body actions, leading to the generation of highly expressive human animations. (2) a Comforting-Exploration-Action support strategy; EMO-Avatar systematically integrates Hill's three-stage counseling theory into its emotional reasoning capability. Guided by the LLM's reasoning, this strategy informs response generation and displays stage-specific preferences for speech, body language, and expressions. EMO-Avatar can provide deeper emotional support and therapeutic human-like interactions. Experimental validation on the AvaMERG Challenge demonstrates EMO-Avatar's superior performance, achieving top-2 ranking among 20 participants across response appropriateness, multimodal consistency, naturalness, and emotional expressiveness metrics. Our demo is available at https://ai4ai.anonymous-demo.fun/.
As immersive 360 degrees video experiences through head-mounted displays (HMDs) gain widespread adoption, the need for real-time, fine-grained assessment of Quality of Experience (QoE) becomes increasingly critical for optimising user engagement and system performance. This paper introduces RCQoEA-360VR, a novel multi-modal dataset designed for continuous QoE evaluation in virtual reality (VR) environments. In a controlled study (N=32), participants watched five selected 360 degrees video sequences across eight different video quality configurations (from the VQEG database) using a Vive Pro Eye while providing continuous QoE annotations via a touchpad-based input method, enhanced by the DotMorph peripheral visualisation technique. The dataset also includes synchronised physiological signals (electrocardiogram and galvanic skin response), behavioural data (eye and head movements) and post-viewing QoE ratings gathered through a within-VR interface. RCQoEA-360VR addresses a critical gap in existing public datasets by providing a fine-grained, synchronised multimodal data for immersive QoE analysis. It offers a unique and valuable resource for the research community, supporting a wide range of research applications, including QoE prediction, behavioural modelling, adaptive streaming, and implicit perceptual analysis.
The rise of Multimodal Large Language Models (MLLMs) offers new opportunities for Micro-Expression (ME) analysis. This paper introduces Micro-Expression Visual Question Answering (ME-VQA), a novel task reformulating ME annotations (e.g., emotion categories, action units) into QA pairs. To address key challenges-hardware limitations, context inconsistency, and compositional reasoning gaps-we propose a Relationship-Aware Hierarchical VQA Framework. Our approach leverages mined emotion correlations (e.g., coarse-to-fine label dependencies) and employs a two-stage process: 1) Coarse-grained anchoring for broad emotion categories, and 2) Fine-grained reasoning constrained by coarse outputs and statistical rules. We further optimize efficiency via a dual-phase video sampling strategy: during training, keyframes (onset/apex/offset) and random non-expression frames are used; uniform sampling is applied at inference. Experiments demonstrate significant improvements in answer consistency and accuracy.
Classifying anatomical regions in endoscopic ENT (ear, nose, and throat) images is challenging due to strong inter-class similarities, bilateral symmetry, and the scarcity of annotated datasets. To overcome these issues, we present HyMoENet. This novel hybrid deep learning architecture combines convolutional neural networks (CNNs) for localized feature extraction with Vision Transformers to represent global context. Furthermore, it leverages a sparse Mixture-of-Experts (MoE) technique to improve multi-perspective specialization. Our architecture makes use of parallel CNN-Transformer encoders, which are incorporated into a dynamic MoE layer that adaptively routes representations to the most appropriate experts. Concurrently, a semantic-preserving skip connection preserves global coherence. When tested on a clinically annotated ENT endoscopy dataset from Thong Nhat Hospital in Vietnam, HyMoENet outperformed both single-stream and conventional hybrid models with an accuracy of 97.50%. These results demonstrate that integrating local-global representation learning and expert modularization enhances classification accuracy for anatomically similar structures. HyMoENet sets a new benchmark for automated ENT image processing, setting the groundwork for intelligent diagnostic systems in clinical endoscopy.
In this companion paper, we reproduce the experiments presented in our work titled "NIF: A Fast Implicit Image Compression with Bottleneck Layers and Modulated Sinusoidal Activations" [2], presented at ACM Multimedia 2023. In this study, we present the architecture and the technical details of our implementation and provide instructions to reproduce the main results, the ablation study, the plots and the figures presented in the paper. All the material described in this paper is released on GitHub [3], featuring the full results, a reference software implementation and a generic environment setup that works on any system, even without a GPU.
Video Question Answering (Video QA) has emerged as a central task in multimodal learning. This tutorial provides a comprehensive overview of VideoQA research and highlights new frontiers. We begin with an introduction to VideoQA preliminaries, tracing how methods have adapted from third-person view short videos to capture egocentric and long-ranged spatial-temporal dynamics. We then focus on the impact of large multimodal models. Next, we expand the scope to spatial understanding beyond videos. Each topic is discussed through the lens of tasks, datasets, methods, and evaluation protocols. Finally, we conclude with future directions, including fine-grained and long-ranged video understanding, robustness and trustworthiness, Egocentric and embodied assistance, and omnimodal integration. This tutorial aims to equip participants with both a historical perspective and a forward-looking roadmap for advancing Video QA in the LLM era.
Recent text-to-image generative models facilitate creating vivid images with arbitrary contents that are indistinguishable from authentic ones by naked eyes. Despite progress in synthetic image detection, detecting the image from new generators remains challenging. Because advanced generators leave fewer visible forgery traces, while different generative frameworks produce varied forgery patterns. We notice that generative models consistently struggle with fine-detailed content generation, creating abnormal spatial dependencies among neighboring pixels in complex texture regions. In this paper, we propose a methodology of gazing local detail of forgery (GLDF) for generator agnostic synthetic image detection, which identifies prominent spatial dependencies to capture subtle forgery. Concretely, we design frequency-aware correlation discovering (FACD) module to learn dynamic filters by instance-adaptive frequency masking block for identifying prominent spatial deficiencies, which distributed in different spatial positions with various patterns. Furthermore, we introduce the spatial forgery clue distilling module (SFCD) to iteratively aggregate and refine spatial dependencies from different positions by spatial aggregating and prototype global interacting blocks. Extensive experiments demonstrate that GLDF outperforms state-of-the-art methods on detecting synthetic images from different generators.
In many real-world applications, labeling an image "a man riding a horse" fails to satisfy demands for the who, when, where, and why. Although LVLMs excel at describing visual content, isolated images often lack the event context; users thus rely on related news articles or social posts to enrich them, but cropping or resizing complicates tracking back to their source. In this paper, we propose ENRIC, an innovative end-to-end system for the EVENTA Challenge Track 1, leveraging the OpenEvents-V1 dataset, comprising over 200,000 news articles paired with more than 400,000 images. Our system includes three components: (1) semantic retrieval filters candidate article images via vision-language embeddings, (2) uncertainty-guided re-ranking flags ambiguous queries using three confidence heuristics and re-ranks candidates by combining visual similarity with texture similarity, and (3) event-aware caption generation employs chain-of-thought prompting that aggregates five inputs from article, image, and CIDEr-derived contexts to guide the LLM in incorporating all necessary elements. ENRIC achieved the highest combined evaluation score of 0.5501, ranking first and outperforming other solutions across nearly all metrics. By combining semantic retrieval, uncertainty-guided re-ranking, and event-aware caption generation, ENRIC demonstrates the efficiency of its approach for event-enriched image analysis. GitHub repository: https://github.com/NamQuanProject/EVENTA25-ENRIC
Multimodal interleaved reasoning, which requires models to understand interleaved image-text sequences and multiple images, is a critical challenge in contemporary AI. This paper proposes a parameter-efficient fine-tuning framework based on Large Vision-Language Models, with Qwen2.5-VL as the backbone and Low-Rank Adaptation for task-specific adaptation. The framework integrates four stages: multimodal input preprocessing to align with pre-training distributions, visual feature extraction via a modified Vision Transformer, cross-modal fusion via attention mechanisms, and response generation via an autoregressive decoder. By freezing pre-trained weights and fine-tuning low-rank adapters in both visual and language modules, it balances preserving general multimodal knowledge with optimizing target tasks, achieving high performance with low computational overhead. On the MIRAGE Challenge Track A Dataset, it performs strongly across subtasks, achieving an aggregate score of 0.7857 and securing second place in the challenge. Ablation studies confirm that joint LoRA fine-tuning of visual and language modules yields optimal results; limitations in fine-grained visual difference tasks indicate future directions in enhancing subtle feature capture and adaptive cross-modal alignment.
LAVA Challenge 2025 aims to improve the ability of large visual language models to accurately understand complex visual information such as data flow diagrams and Gantt charts contained in Japanese government and business documents. For this challenge, we adopted a two-stage approach consisting of retrieval and reading comprehension. Specifically, in the retrieval step, we select pages relevant to the question from multi-page PDF documents, and in the reading comprehension step, we perform question answering by referring to the top.. images selected in the retrieval step. For the retrieval step, we employ ColQwen2, which performs visual information retrieval using the multilingual Qwen2-VL. For the reading comprehension step, we propose a method that performs question answering, including voting, using the multilingual visual language model Qwen2.5VL under different prompts, model sizes, and image qualities. In LAVA Challenge 2025, we clarify the importance of the two-step process of retrieval search and reading comprehension in visual question answering, and verify the effectiveness of a method for determining the multiple results of reading comprehension steps through voting.
The goal of this workshop is to showcase the latest advancements in generative AI (GAI) for creating, editing, restoring, and compressing rich media data, including images, videos, and 3D content. GAI models such as VAEs, GANs, and diffusion models have demonstrated remarkable impact in both academic research and industrial applications. For example, GAI enables users to design and generate synthetic yet realistic content without requiring professional artistic or technical expertise, driving significant market growth in gaming and entertainment. Beyond creative applications, GAI also provides crucial simulated data for training embodied AI agents. When applied to media restoration and synthesis, GAI techniques can further alleviate transmission challenges by offloading computation to client devices. To advance this field, the workshop will host four competition tracks using novel industry-level data, solicit high-quality paper submissions, and invite leading speakers from academia and industry to foster collaboration and innovation. In particular, the competition focuses on media generation and transmission with GAI. The first three tracks address reducing computation and transmission costs for efficient media delivery, while the fourth track focuses on controlled novel content creation. To support these challenges, a large-scale multi-modality, multi-view dataset named (MVIR)-V-3 is provided. This dataset comprises a diverse collection of videos simulated using the UE5 Unreal Engine, with carefully matched content serving as ground truth for the competition tasks.
Understanding fine-grained sentiment dynamics in human conversations is a central goal for next-generation artificial intelligence, especially in scenarios where interactions are rich in both modalities and context. To advance research in this area, we organize the Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) challenge to the community of aspect-based sentiment analysis. The MCABSA challenge introduces two novel subtasks: 1) Panoptic Sentiment Sextuple Extraction, panoramically recognizing holder, target, aspect, opinion, sentiment, and rationale from multi-turn, multi-party multimodal dialogue; and 2) Sentiment Flipping Analysis, detecting the dynamic sentiment transformation throughout the conversation along with the causal reasons. To support these tasks, we present the PanoSent dataset, a high-quality, large-scale benchmark featuring multi-turn, multi-party dialogues annotated with both explicit and implicit sentiment elements across text, image, audio, and video modalities. PanoSent offers extensive real-world scenario coverage, providing a comprehensive testbed for multimodal conversational sentiment analysis. The challenge has attracted widespread participation from both academia and industry, with over 30 teams registered and more than 100 successful submissions. In this paper, we introduce the task, dataset, and evaluation settings, summarize the systems of the top teams, and discuss the findings of the participants. Further details of the challenge can be found at https://panosent.github.io/MM25-challenge.
Micro-action refers to subtle, low-intensity non-verbal behaviors that can provide insights into an individual's underlying emotions and intentions. Due to its brief duration and significant overlap, identifying these micro-actions poses a challenge for current models. In response to these challenges, this paper proposes a novel multi-feature fusion framework, which extracts coarse-grained body features and fine-grained action features separately. Specifically, we present Temporal Contextualization for fine-grained learning, a cross-frame injection mechanism designed to capture essential spatio-temporal information and introduce a 3D-ResNet Adapter for coarse-grained learning, which aggregates temporal data and facilitates parameter-efficient fine-tuning. In consideration of the task dataset distribution's long-tail nature, the implementation of Feature Decoupling is undertaken, adopting a two-stage training strategy. By conducting experiments, the aforementioned hierarchical multi-feature extraction and aggregation approach has been demonstrated to yield substantial enhancement in Micro-Action Recognition. Our method attains an F1-mean score of 77.75% on the MA-52 dataset, ranking 1st in the 2nd Micro-Action Analysis Grand Challenge in Conjunction with ACM MM'25.
This paper describes RoboSax Melody Slot Machine, which automatically plays fingering melodies selected using a dial-like controller with stave notation on a tablet screen. All that are required for the saxophone to output the melody selected by the tablet operator are the blowing and tonguing of the saxophonist. RoboSax Melody Slot Machine makes the task of selecting melodies, which was previously possible only for composers, possible for the general people and enables saxophone playing by those who cannot quickly read the note names on the staff.
Although current work of text-to-Image generation can preliminarily generate images from the descriptions of human-object interactions, it fails to consider the emotions involved in human-object interactions. While people often experience emotions when using objects or interacting with them. Therefore, in this paper, we propose Emotional Interaction Generation task, a novel image generation task, which generates emotionally expressive human-object interaction images from given prompts, human-object interaction (HOI) region, and emotions. First, we construct a new emotional interaction dataset, called EmotionHOI, which including 47,776 images with content prompt, emotions and human-object interaction bounding box. Second, we propose an emotion-aware text-to-image diffusion model, named EmIT, for emotional interaction generation. Specifically, EmIT consists of three components: (1) an emotion interaction tokenizer that encodes subject, object, action, and emotion into structured tokens; (2) an Emo-Interaction Self-Attention that preliminarily guides the latent space to conduct hybrid learning with emotional interaction tokens; and (3) a Hierarchical Emotion-Visual Cross-Attention that further focus on grounding affect-such as pose, gaze, or interaction intensity-into specific spatial regions and capture subtle emotional variations. These components jointly model interaction semantics and emotional context, enabling EmIT to generate images that are both behaviorally coherent and emotionally expressive. Experimental results on the EmotionHOI dataset demonstrate the superiority of the proposed model.
With the rapid advancement of AIGC technologies, audio deepfakes have become increasingly realistic, posing serious threats to information security and biometric authentication. Therefore, audio deepfake detection (ADD) has emerged as a critical and fast-evolving research area, particularly requiring superior generalization in out-of-domain scenarios. However, existing ADD methods suffer from constrained generalization and limited access to target data. To address these challenges, we propose Risk-Aware Style Alignment (RASA), a novel generalizable ADD framework that projects the style of any input feature into a shared style space through similarity-based projection. This alignment reduces both inter-domain and intra-source discrepancies without requiring target data during training. In addition, we adopt Structural Empirical Risk Minimization (SERM) in the Poincare ball model to capture the hierarchical structure of the data and further minimize source risk. By jointly optimizing RASA and SERM, the proposed method effectively tightens the theoretical upper bound of target risk across three key dimensions: source risk, inter-domain divergence, and intra-source discrepancy. Extensive experiments demonstrate that our approach achieves superior generalization and outperforms existing state-of-the-art methods.