3D object detection aims to accurately localize and recognize objects in 3D space. It serves as a fundamental task for reliable perception in intelligent transportation systems, enabling the monitoring of diverse traffic participants such as vehicles, pedestrians, cyclists, and public transport. Recently, transformer-based methods have gained significant attention in multi-view 3D object detection due to their strong global reasoning capabilities. However, their limited capacity to model spatial positional information hinders accurate object localization, especially in complex and large-scale scenes. To address this limitation, SpaceFormer is proposed as a novel transformer-based multi-view 3D object detector. Specifically, a Contextual Visual Prompts Learning strategy is proposed to enhance the perception of small and sparse traffic participants by incorporating contextual priors. To further suppress background interference, a Semantics-guided Depth Estimation method is proposed to refine depth representations using high-level semantic information. Furthermore, a Spatial Position Embedding mechanism is proposed to improve the spatial localization capability of the transformer by integrating geometric position and polar spatial embedding. Extensive experiments on the nuScenes benchmark demonstrate that SpaceFormer achieves state-of-the-art performance with 55.5% mAP and 62.9% NDS. These improvements indicate not only methodological advances but also practical benefits for intelligent transportation systems, enhancing safety, reliability, and efficiency in real-world deployments.
The introduction of a Digital Twin enhanced framework, which integrates Internet of Things (IoT) sensing, artificial intelligence (AI), and virtual representations of users and environments, opens up a world of possibilities for delivering personalized context-aware activity recommendations. Real-time environmen-tal and behavioral data are processed using a hybrid wide-and-deep learning model that combines historical feature co-occurrence with deep contextual generalization. The system simulates activity scenarios such as studying in a cafe or exercising in a park, using Digital Twins to predict potential impacts on mood, comfort, and productivity before actions are taken. A continuous feedback loop between physical and digital entities refines predictions over time. An initial evaluation using realworld Ottawa sensor data confirms the model's effectiveness in ranking points of interest based solely on environmental parameters. Unlike conventional recommender systems that rely on static profiles, this framework incorporates real-time contextual variables, such as air quality, noise levels, and UV exposure, to enhance the relevance of recommendations. The convergence of IoT infrastructure, AI modeling, and Digital Twin simulation not only enables scalable, adaptive personalization in smart environments but also paves the way for more informed decisionmaking and improved individual well-being.
Recent real-time detection transformers (DETRs) have gained popularity due to their simplicity and efficiency. However, these detectors do not explicitly model object rotation, especially in remote sensing imagery where objects appear at arbitrary angles, leading to challenges in angle representation, matching cost, and training stability. In this article, we propose a real-time oriented object DETR, the first real-time end-to-end oriented object detector to the best of our knowledge, that addresses the above issues. Specifically, angle distribution refinement (ADR) is proposed to reformulate angle regression as an iterative refinement of probability distributions, thereby capturing the uncertainty of object rotation and providing a more fine-grained angle representation. Then, we incorporate a Chamfer distance cost into bipartite matching, measuring box distance via vertex sets, enabling more accurate geometric alignment and eliminating ambiguous matches. Moreover, we propose oriented contrastive denoising (OCD) to stabilize training and analyze four noise modes. We observe that a ground truth can be assigned to different index queries across different decoder layers, and analyze this issue using the proposed instability metric. We design a series of model variants and experiments to validate the proposed method. Notably, our O-2-DFINE-L, O-2-RTDETR-R50 and O-2-DEIM-R50 achieve 77.73%/78.45%/80.15% AP(50) on DOTA1.0 and 132/119/119 FPS on the 2080ti GPU.
Dear colleagues and friends, We would like to begin by sincerely thanking the SIGMM community for the trust you have placed in us. We are honored to serve as Chairs alongside a talented and dedicated team. A special thanks goes to the previous Executive Committee for their outstanding work during challenging times and for laying down a solid foundation for the future.
Cross-modal federated learning is constrained by bandwidth and on-device compute. We present Mobiflip: a minimalist strategy that freezes a lightweight backbone and communicates only a channel-wise \(1\times 1\) scaling adapter appended to the image branch. Guided by the Information Bottleneck, we prove that under common distributional and linear-encoder surrogates, per-channel scaling attains the linear optimum; coupled with the directional geometry of (Mobile)CLIP, the adapter is, in first-order approximation, an optimal preconditioner of the cosine-similarity space—preserving discriminative directions while compressing redundancy and suppressing inter-client drift. We adopt MobileCLIP as a mobile-friendly backbone to jointly minimize compute and communication. On CIFAR-10/100 and medical imaging, a single aggregation already yields stable Bacc; each round transmits only about 0.7% of backbone parameters with \(>\!\!92\%\) reduction in communication. Compared with recent federated multimodal/large-model methods, Mobiflip maintains—or even improves—accuracy under ultra-low communication.
We introduce segmentation-guided spatial indexing for generalizable and explainable deepfake detection. The key idea reverses the standard design order: rather than pooling all facial tokens and classifying afterward, we first select semantically meaningful patch tokens, then pool only those. A frozen FaRL parser assigns each DINOv3 ViT-L/16 patch token a semantic label; non-target tokens are discarded; a linear probe classifies the retained region. This spatial indexing exploits DINOv3's patch-level spatial consistency, the same property that enables emergent segmentation, to present the probe with a purer regional subspace where manipulation-relevant evidence is less diluted by whole-face cues. Region attribution is structural: when the mouth model predicts fake, the decision used only mouth tokens, not an overlaid saliency map. On Celeb-DF v2, the mouth-indexed probe achieves AUC 0.905, outperforming LipForensics (+8.1 pp) and Xception (+16.9 pp), with no DINOv3 or FaRL fine-tuning and no target-domain data. Ablations isolate the mechanism: replacing regional selection with DINOv3's CLS token drops Celeb-DF v2 AUC by 26.4 pp; replacing DINOv3 with FaRL features drops it by 20.9 pp. Both DINOv3 representation and the spatial index are independently necessary; neither alone approaches the full system.
With the rapid advancement of Physical AI technologies, Physical AI labeling specifically refers to the specialized data annotation pipeline designed to support the training of AI systems that must operate and interact directly within the real physical world. In contrast to conventional labeling paradigms which primarily involve assigning categorical or semantic tags to 2D images or textual sequences, Physical AI labeling focuses on enriching raw, multi-modal demonstration data with high-fidelity annotations that capture essential physical, spatial and behavioral context. This includes multi-modality sensor streams (e.g., vision, and force), 3D spatial reasoning, material properties, interaction dynamics, etc.
This paper presents an empirical evaluation of the Matterport Pro3, a consumer-grade 3D scanning device, for large-scale environment reconstruction. We conduct detailed scanning (1,099 scanning points) of a six-floor building (17,567 square meters) and assess the device's effectiveness, limitations, and performance enhancements in diverse scenarios. Challenges encountered during the scanning are addressed through proposed solutions, while we also explore advanced methods to overcome them more effectively. Comparative analysis with another consumer-grade device (iPhone) highlights the Pro3's balance between cost-effectiveness and performance. The Matterport Pro3 achieves a denser point cloud with 1,877,324 points compared to the iPhone's 506,961 points and higher alignment accuracy with an RMSE of 0.0118 meters. The cloud-to-cloud (C2C) average distance error between the two point cloud models is 0.0408 meters, with a standard deviation of 0.0715 meters. The study demonstrates the Pro3's ability to generate high-quality 3D models suitable for large-scale applications, leveraging features such as LiDAR and advanced alignment techniques.
Despite the success of foundation models in language and vision, their expansion into embodied AI is bottlenecked by a lack of generalized touch sensing. This limitation is especially relevant to consumer electronics, where smartphones, wearables, VR controllers, home robots, and health monitoring devices require safe and adaptive physical interaction. Constrained by hardware heterogeneity and the necessity of active physical data collection, current haptic models remain rigidly task-specific. To overcome these limitations, this article explores the transformative potential and developmental trajectory of Haptic Foundation Models (HFMs). We detail the paradigm shift required to transition from passive Large Language Models and Vision Language Models into active HFMs across four core dimensions: action coupling, physical dynamical representation space, continuous time-series data granularity, and action-conditioned future state prediction. Furthermore, we synthesize existing large-scale tactile datasets and benchmark UniTouch, AnyTouch, T3, and Sparsh on TacBench for force estimation, slip detection, and relative pose estimation.
Unsupervised Graph Domain Adaptation (UGDA) aims to facilitate knowledge transfer from a labeled source graph to an unlabeled target graph by mitigating cross-domain distribution shifts. Existing methods primarily focus on node-level feature alignment in latent spaces, relying on the implicit assumption that all source nodes contribute positively to the alignment. However, this assumption often fails because a node's semantic information is intrinsically coupled with its topological graph structure. Due to structural shifts, source nodes with severe structural deviations (e.g., structural outliers) lack semantic counterparts in the target graph, and forcing alignment on them introduces severe noise and causes negative transfer. To bridge this gap, we argue that selective source node utilization is superior to full-graph training, thereby shifting the research paradigm from feature-level alignment to data-level refinement. To this end, we propose Source Node Influence Pruning (SNIP), a novel model-agnostic, data-centric refinement framework. Specifically, SNIP quantifies the structural discrepancy between individual source nodes and the target domain by integrating multiple centrality measures, assigning each source node an influence score. A rank-based normalization mechanism is further employed to eliminate scale variations across different measures, allowing SNIP to effectively identify and filter out structurally incompatible nodes with low influence scores. As a plug-and-play method, SNIP constructs a refined "sub-source" graph that is inherently more beneficial for subsequent alignment. Comprehensive experiments across eight transfer scenarios on five real-world datasets demonstrate that SNIP consistently outperforms competitive baselines and significantly enhances adaptation performance, validating the superiority of selective node utilization over full-graph training.
Poor posture contributes to musculoskeletal disorders that reduce productivity and quality of life. This paper presents the Haptic Posture Correction System (HPCS), an edge-assisted interactive platform that integrates augmented reality (AR) visualization with real-time vibrotactile feedback for ergonomic posture training. The HPCS combines a wearable vibrotactile jacket with anatomically informed actuator placement, a mobile AR interface for region selection, and a lightweight IoT communication layer that performs low-latency actuation using the Blynk relay platform. If a personalized or customized avatar is required, the trainer retrieves it from the cloud before the training session begins. This ensures that all subsequent sensing, control, and feedback operations occur locally with minimal delay. Trainer input on the AR interface directly triggers localized tactile cues on the trainee, while the trainee’s postural response is visually observed, forming an implicit perceptual feedback loop. Initial user studies show that participants experienced enhanced responsiveness, intuitiveness, and training effectiveness compared to conventional feedback methods. By emphasizing real-time multimodal feedback and edge-assisted control, the HPCS provides a sustainable and scalable framework for AI-augmented ergonomic training and remote posture correction applications.
Generalization under manipulation and dataset shift remains a core challenge in forged media detection for AI-driven edge sensing systems. Frozen vision foundation models with linear probes are strong baselines, but most pipelines use default backbone outputs without testing conditioning at the frozen feature interface. We present the first controlled probing study on DINOv3 ConvNeXt and show that, without task-specific fine-tuning, linear probing alone yields competitive forged-media detection performance, indicating that ViT-7B self-supervised distillation transfers to security-critical vision workloads at edge-compatible inference cost. Backbone, head, data, and optimization are fixed while conditioning is varied; LN-Affine, the default ConvNeXt head output, is the natural baseline. On FaceForensics++ c23, five conditioning variants are evaluated under in-distribution testing, leave-one-manipulation-out (LOMO), and cross-dataset transfer to Celeb-DF v2 and DeepFakeDetection. In ConvNeXt-Tiny, conditioning alone changes LOMO mean AUC by 6.1 points and reverses ID-vs-OOD ranking: LN-Affine is strongest on external datasets, while LayerNorm is strongest in-distribution. In ConvNeXt-Base replication, the OOD winner becomes protocol-dependent, and ID-optimal selection still fails as a robust deployment rule. Results show that feature conditioning is a first-order design variable and should be selected with robustness-oriented validation, not ID accuracy alone.
Extended Reality (XR) systems increasingly deliver high-fidelity visual and auditory experiences, yet tactile perception remains comparatively underutilized as a modality for enriching embodied interaction. This work presents a visual-to-haptic wearable glove and a feature-based visual-to-haptic mapping algorithm that translates spatial and temporal visual features from images and videos into distributed vibrotactile patterns. The proposed method extracts motion, edge, and brightness cues and fuses them into actuator-level intensity maps aligned with a 29-actuator glove arranged in a five-by-seven layout. The system is implemented through a modular four-layer architecture comprising the XR environment, media content handling, visual-to-haptic processing, and embedded haptic hardware. A within-subject user study (N = 20) compared visual-only interaction with visual-plus-haptic augmentation across texture-based and dynamic video scenarios. Results indicate that tactile augmentation significantly improves perceived realism in dynamic video scenarios and enhances immersion and visual-tactile correspondence across conditions, with stronger and more consistent effects observed for dynamic visual events. While the current implementation operates in a single-user, offline-synchronized configuration, the findings demonstrate that vision-driven tactile augmentation can function as a perceptual enhancement layer within multimodal XR systems. Such a layer may provide a foundation for future socially enriched XR environments where coherent multisensory grounding supports higher-level interaction and communication.
Open world object detection aims to identify unseen unknown categories during training and dynamically learn these new categories as their labels become available. However, the key challenge of this task lies in accurately distinguishing between known and unknown objects. Due to the absence of supervision for unknown classes during training, their feature representations often become entangled with those of known classes, inducing spurious correlations that lead to frequent misclassifications. To address this issue, we propose a Causal Unbiased Feature Intervention framework for open world object detection. Specifically, we construct a feature extraction causal graph to reveal spurious correlations between features of known and unknown classes. Then, we introduce a Causal Adjustment (CA) strategy to mitigate these confounding effects in the prediction distribution. Finally, we present an Unbiased Feature Intervention (UFI) module that further decouples known and unknown features at the feature level. Experimental results demonstrate that the proposed framework achieves superior performance across multiple open world object detection benchmarks, while improving the model's ability to recognize unknown classes.
The integration of artificial intelligence (AI) in mental health therapy has led to the emergence of AI-powered therapeutic assistants. However, existing systems primarily function as reactive conversational agents, lacking real-time emotional context awareness. This paper presents UbiMyTherapist, a digital-twin, ubiquitous, multimodal framework that integrates a Retrieval-Augmented Large Language Model (RAG-LLM) with near-real-time emotion detection to provide both reactive and proactive support on consumer electronics such as smartwatches and smartphones. The framework builds a continuously updated digital-twin of the user’s emotional state and medical history, leveraging affect recognition models that can process bio-signals, speech intonation, or text sentiment. This integration enables a more responsive AI therapy assistant. In this paper, we describe the system architecture, its components, and our prototype. The reactive mode was evaluated in a user study with 24 participants, where UbiMyTherapist outperformed baseline LLM setups on therapist-likeness and conversation quality.
Accurate and efficient modeling of indoor wireless signal propagation is crucial for the deployment of next-generation Wi-Fi. This article presents a digital twin-based measurement system that integrates real-world 3D environment reconstruction with deterministic ray tracing (RT) for physically grounded electromagnetic modeling. Building geometry is obtained through LiDAR scanning, followed by object segmentation and assignment of ITU-R standard material parameters. The propagation process is simulated with a GPU-accelerated ray-tracing engine that generates path-level channel attributes, including delay, power, angular dispersion, and Ricean K-factor. Under identical runtime constraints, the proposed system is evaluated against a commercial measurement simulator, demonstrating up to 21 dB higher path gain and consistently improved signal-to-interference-plus-noise ratio in line-of-sight (LOS) conditions. Additionally, experiments against onsite RSSI measurements confirms a high spatial correlation of 0.98 after calibration, proving the system's fidelity in real-world settings. Furthermore, coverage analysis across 2.4, 5, and 6 GHz bands demonstrates the capability of system to model frequency-dependent material attenuation for Wi-Fi 6E/7 networks. Finally, the system offers interactive 3D visualization and on-demand data extraction, highlighting its potential for digital twin-driven wireless system design and optimization.
Frozen self-supervised vision models can align parts of generic objects, but it remains unclear whether this correspondence extends to human faces, where global layout is shared while identity-specific appearance varies sharply. We test whether frozen DINOv3 features define a region-level facial coordinate system: a feature space in which eyes, brows, nose, mouth, skin, and hair remain distinguishable across people and across time without face-specific training. Using DINOv3 ViT-L/16 patch embeddings and FaRL only as a face-part labeling interface, we evaluate cross-identity nearest-neighbor matching and temporal label propagation on 200 CelebDF-v2 real videos. DINOv3 achieves 83.0
Digital twins can provide vital insights into agricultural products and processes. There have been a lot of documented attempts at digital twins in agriculture. However, majority of these attempts build synthetic models and ignore the temporal dimension of the plant growth. Therefore, the existing models fail to depict actual plant details and growth. Our work replicates the actual growth of a real plant in the digital world by acquiring 3D meshes of the plant at various instants. It focuses on the transition between those acquired meshes by approximating all the consecutive pairs into approximate mesh pairs that have a common topology. The quality of these common approximate mesh pairs is quantitatively measured by an Energy term, which is minimized during the optimization process. Later, the meshes with the common topology are interpolated (morphing) to build the final digital twin of the plant. Experimental results show that the proposed methodology to attain the final morph has the potential to be a vital module, which could be responsible for the visual updates in the digital replica of the digital twin of the plant.
Audio analysis is used in many real-world applications, but most methods assume a fixed data distribution. In practice, data can shift over time due to changes in distribution or the appearance of new classes, reducing model performance. Traditional deep learning models fail to adapt, making continual learning (CL) crucial. While some studies explore CL for audio, they mainly focus on a class-incremental scenario and largely ignore a domain-incremental scenario. In applications such as anomaly detection and monitoring, domain incremental scenarios are prevalent. To the best of our knowledge, no existing comprehensive benchmark for audio data supports both domain-incremental and class-incremental learning scenarios. To bridge this gap, we have curated a benchmark using the DCASE dataset (2020–2023) that properly considers both scenarios. Previous works have primarily explored anomaly detection approaches where new classes are treated as anomalies. Our work complements these approaches by examining anomaly detection within established classes, which is crucial for audio applications involving rare or unexpected sounds within known classes. We use the proposed dataset to compare regularization-based (EWC, LwF, and SI) and rehearsal-based (GEM, A-GEM, GDumb, Replay, DER++, and Co ^2 L) CL approaches, as well as non-CL approaches (Naive, Cumulative, and Joint training). We evaluate multiple metrics such as accuracy, forward transfer, and backward transfer using two popular backbones: ResNet50 and ViT. The key findings show that Replay outperforms other CL methods, achieving up to 76.53
With the growing integration of human-computer interaction into everyday life, advances in machine learning have enabled systems to better perceive and respond to users' emotional states. Most existing affect recognition datasets focus on static environments, limiting their applicability to immersive multimedia contexts such as Virtual Reality (VR). In this paper, we introduce WARM-VR, a novel publicly available multimodal dataset designed to support affect recognition in immersive, multisensory environments using wearable sensing instrumentation. Data were collected from 31 participants aged 19-37 using wearable sensors: a wristband measuring Blood Volume Pulse (BVP), EDA, skin Temperature, three-axis Acceleration, and a chest strap recording ECG signals. Participants engaged in immersive VR experiences designed to elicit relaxation through a calming beach environment following stress induction via an arithmetic task. These sessions incorporated synchronized multimedia stimuli: visual, auditory, and olfactory. Affective states were assessed subjectively through validated self-report questionnaires and objectively through the analysis of physiological measurements. Statistical analysis of the questionnaires confirmed that VR relaxation significantly reduced negative affect, particularly with olfactory enhancement. Furthermore, we established a benchmark on the dataset using widely recognized machine learning algorithms. The best performance for binary classification from BVP data of valence, was obtained with a CNN and a CNN-Bi-GRU model, both achieving an average F1-score of 0.63 and an AUC of 0.69. For arousal, a lightweight Transformer architecture provided the most balanced results (F1-0 0.54 and F1-1 0.63), outperforming recurrent hybrids. In the relaxation task, a CNN-Bi-GRU model reached the highest overall performance (average F1-score 0.64, AUC 0.69).
Shervin Shirmohammadi合作论文数University of Ottawa;School of Information Technology and Engineering (SITE)32
Pradeep Atrey合作论文数The University of Winnipeg;Department of Applied Computer Science31