
This paper presents a camera-based optical transmission system and architecture of secure access to Internet of Things (IoT) devices in ubiquitous computing environments. Devices emit light tokens, referred to as local identifiers (LIDs), which are captured and decoded by a nearby camera. A LID initiates a local discovery process, where a bootstrapping service links the LID to an IoT device within a local network, allowing interaction without physical contact or prior pairing. The system supports device setup and control in ad-hoc scenarios using standard commodity hardware. By connecting a temporary physical signal to a network device identity, it enables access without continuous connectivity or complex setup. The light signal serves as a temporary authentication mechanism, while the camera verifies physical proximity and enables positional discovery of nearby devices. It provides a simple, secure way to interact briefly with nearby devices and is well suited for deployment in smart environments, augmented reality interfaces, and other ubiquitous computing contexts.
Understanding and improving student performance is a central concern in education, and predictive models can provide valuable insights-provided their decisions are transparent and explainable. However, many machine learning (ML) models used for this purpose lack interpretability, limiting their practical utility. In this work, we apply two complementary eXplainable AI (XAI) methods: SHapley Additive exPlanations (SHAP) and the Bayesian Counterfactual Generator (BayCon), to explain the predictions of an ML model trained on multimodal data to forecast student exam outcomes. SHAP is used to identify and visualise the most influential features contributing to each prediction, while BayCon generates actionable counterfactuals. These counterfactuals are then converted into natural language using a large language model (LLM), making them easier to understand. Our approach is designed to support both students and educators by offering clear, personalised insights into the factors affecting academic performance. We present the explanation generation pipeline and discuss its interpretability and potential benefits for educational contexts.
Understanding mobile phone behavior has attracted significant attention in HCI and mobile health research. While experience sampling methods (ESM) are widely used, they primarily offer discrete snapshots, making it difficult to observe behavioral flow, transitions, and differences between phone use associated with varying intents and levels of habitualness. We introduce Pulse (Phone Use Labeling via ScrEenshot), a research application that integrates screenshot logging, passive sensor data collection, and a screenshot-based labeling interface designed to support efficient labeling workflows. The app also supports micro-ESM delivery and a dedicated researcher mode allows for flexible configuration and testing of study parameters. We report a case study deployment of Pulse, through which we collected approximately 1.65 million labeled screenshots from 25 participants. Preliminary analysis reveals associations between users' phone use intents and their subsequent evaluations of time use, demonstrating the utility of screenshotbased labeling for understanding mobile behavior in context.
Dan-Sco is a wearable soft sensor prototype designed for live performance contexts. This paper focuses on the design and fabrication of two custom garments that integrate textile-based stretch sensors to capture full-body movement. Through an iterative prototyping process-including embroidery, knitting, conductive painting, and custom tailoring-we explored how different e-textile techniques and materials respond to dynamic motion. Each garment embeds soft sensors at targeted body zones to ensure both comfort and reliable signal capture during high-mobility activities. We detail the construction process, sensor material selection, garment integration strategies, and design decisions that balance wearability, durability, and signal responsiveness. This work offers practical insights into building expressive, movement-sensitive wearables using soft, textile-based materials.
Wearable devices can collect sensitive biometric and behavioral information, necessitating authentication. Common smartphone biometrics include fingerprint and facial recognition; their application to wearable devices is challenging due to varied attachment and limited input. This study investigates active acoustic sensinga technique that analyzes responses from transmitted acoustic signals-as a method for user identification on wearable devices. We evaluated its performance across eight attachment sites: finger, wrist, neck, nose, earlobe, ear bone, teeth, and sole. Data were collected from seven participants over one week at three times of day (morning, midday, evening). Three evaluations were carried out to assess intra-day performance, generalization across days, and long-term stability. The results showed that the acquired biometric signals were unstable over time, leading to poor authentication performance at all sites, thus highlighting a significant challenge for wearable acoustic biometrics that motivates further research.
Service robots often engage in liquid-related tasks such as pouring, which require accurate liquid identification. Vision-language models (VLM) exhibit strong performance in general object recognition, they struggle to reliably distinguish between visually similar liquids, limiting their effectiveness in such scenarios. As a promising solution, millimeter-wave (mmWave) radar combined with neural networks has been widely adopted for material recognition and liquid identification. However, training a radar-based liquid classifier demands a large volume of labeled radar-liquid data pairs and often faces challenges in generalizing to real-world environments. To address this challenge, we propose FuseLID, a VLM-based camera-radar late-fusion system designed to improve liquid identification performance. FuseLID leverages the weak penetration capability of mmWave signals through liquids and combines this with the robot's precise motion control to capture distinctive radar signatures of various liquids. Subsequently, a VLM-based late-fusion module is designed to combine camera and radar outputs for achieving enhanced liquid identification with limited radar-liquid data. Preliminary experiments show that FuseLID improves liquid identification accuracy from 51.6% to 95.7% when classifying several commonly consumed beverages, including visually similar ones.
Achieving sustainable urban transportation requires not only urban planning but also technological innovation, particularly for vulnerable road users (VRUs) such as cyclists, pedestrians, and e-scooter riders. While active mobility (i.e., cycling and walking) offers environmental and health benefits, VRUs face a higher risk of serious and fatal accidents than vehicles. This full-day workshop will examine how ubiquitous computing, sensing, and communication technologies can improve VRU safety, promote active mobility, and support the development of a sustainable and resilient transportation system while addressing socio-technical aspects such as legal, ethical, psychological, and normative considerations. Therefore, we will invite two keynote speakers, one from technical and one from socio-technical science, and prepare an open discussion session to foster interdisciplinary exchange. The discussion will encourage participants to explore technical innovations and socio-technical implications for enhancing VRU safety and promoting sustainable mobility.
We present our preliminary findings from an ongoing study aimed at developing an Artificial Intelligence (AI)-powered system that incorporates artists' input to facilitate meaningful, conversational interactions about artworks, going beyond traditional labels and fostering connections between exhibition visitors and artists. Findings from the formative evaluation of the proposed system with artists (N=5) who hold a neutral attitude toward AI suggests that the system has the potential to expand access to and expression of art, act as an agent on behalf of artists, and fit well within their current exhibition practices. These findings may also inform the integration of AI into other artistic practices, including performance art. Based on these findings we discuss the need for educational efforts to expand artists' understanding of AI capabilities beyond generative tools and implications for the future design of the proposed system in terms of preserving artists' authentic voice and supporting multilingual accessibility.
Current medical LLM evaluation prioritizes accuracy over safety, creating a positive-negative capability divide" where models excel at selecting correct answers but struggle to detect misinformation. We develop an evaluation framework using Traditional Chinese Medicine (TCM) examinations, transforming over 3,000 questions into four testing paradigms: Standard multi-choice questions, Wrong Options Test, Misleading Guidance Test, and Fabricated Entity Test. Evaluating six models reveals striking disparities: while some achieve clinical thresholds (>60%) in standard evaluation, all experience dramatic degradation in hallucination detection (20-39%) and near-zero performance in detecting fabricated medical concepts. These findings challenge standard accuracy metrics and establish safety-prioritized evaluation standards.
With the increasing use of mobile devices in public spaces, users are becoming increasingly vulnerable to visual privacy attacks, commonly known as shoulder surfing. These attacks allowbystanders to visually extract sensitive information from screens without permission. In this work, we introduce a novel context-aware framework that leverages mobile device camera sensors and bystander positions to detect and prevent visual privacy attacks. Specifically, we present: a reactive system that detects nearby intruders using the front-facing camera and alerts users in real-time, and a proactive text reader that dynamically adjusts content visibility to thwart unauthorized viewing. Our experimental evaluation demonstrates the robustness of our system in varied real-world scenarios. This work lays the groundwork for intelligent, user-friendly visual privacy protection on mobile platforms.
This tutorial provides an overview of location-based augmented reality and the spatial web, their use cases, the ecosystem, and challenges related to the lack of interoperability between proprietary solutions. The Open AR Cloud Association promotes an open spatial web and develops open-source components to achieve this goal. Participants will learn about these components, their current state, future steps, and gain access to the toolset and proof-of-concept use cases to to build their own location-based AR experiences anchored to the real world.
This paper presents LODYSEI, an ongoing project aiming at generating intelligible, intuitive and easy to memorize object position descriptions in natural language for blind people. We present preliminary results concerning the development and testing of an indoor localization system using Ultra-Wideband (UWB) technology, the evaluation of a voice recognition interface and the use of the localization information by a Large Language Model (LLM) in order to generate position descriptions in French.
Counterfactual explanations (CFEs) offer a promising approach for understanding personalized psychological processes captured through Ecological Momentary Assessment (EMA). In particular, generating CFEs for time points associated with mental health deterioration can help identify alternative changes that might prevent such outcomes. To assess the quality of generated CFEs on timeseries data, we propose a 3-level framework: per explanation (based on feature changes, proximity, and model confidence), per time point (based on structure and diversity through clustering), and per individual temporal dynamics. Applying this framework to a real-world time-series EMA dataset, we demonstrate how we assess the validity and interpretability of large volumes of CFEs.
Brain-computer interfaces (BCI) based on augmented reality steady-state visual evoked potentials (AR-SSVEP) face critical challenges in mobile environments, including low signal-to-noise ratio (SNR) from dry electrodes and limited computational resources on mobile embedded platforms. To optimize the AR-SSVEP system performance, this study comprehensively considers the stimulusresponse coupling mechanism integrating visual optical principles with deep learning-based classification. First, we designed an optimal AR visual stimulation configuration scheme capable of adaptively adjusting key parameters. Second, to address the time-varying non-stationary characteristics of SSVEP and interelectrode quality variations in dry electrode systems, we propose CBAM-FNet-a lightweight SSVEP detection algorithm that incorporates the Convolutional Block Attention Module (CBAM) with multi-band fusion. The algorithm achieves classification accuracies of 93.84% on benchmark datasets and 74.43% on our self-collected AR-SSVEP dataset, representing performance improvements of up to 22.96% over state-of-the-art methods. Real-time implementation on an embedded unmanned vehicle platform demonstrates 65% control accuracy with an information transfer rate of 50.35 bits/min, validating the practical value of CBAM-FNet in embedded BCI applications and overcoming hardware-imposed performance limitations.
Existing counseling chatbots typically determine flow transitions based on explicit user profiling or emotion recognition. However, in real-world mental health contexts, users often struggle to accurately recognize or verbalize their internal states, making per-turn inference both costly and unreliable. We propose a selective modeling strategy that monitors deviations in the chatbot's own persona (so-called persona drift) as an indirect signal of user state change. When such deviation is detected by the heuristic module, signaling a subtle misalignment between the user's implicit state and the chatbot's consistent persona, the user state is reassessed and the counseling flow is conditionally adjusted. To reflect actual therapeutic processes, we implement a modular counseling system grounded in the Transtheoretical Model for Change (TTM), with five chatbots tailored according to user's behavior readiness stage. Together, the TTM-based architecture with persona drift module form a lightweight yet adaptive framework for tracking user state and guiding conversation flow in mental health chatbots.
Understanding when and where traffic calming measures are implemented is essential to assess their impact and plan future safety interventions for vulnerable road users. Yet, such records are often incomplete or unavailable. To support accurate, automated, and large-scale efforts to address these critical data gaps, we propose a computer vision-based framework to detect such measures from historical street view imagery that captures real-world urban complexity. We share key preliminary results demonstrating the effectiveness of our framework in overcoming visual challenges within these images, providing a solid foundation as we continue to improve and progress toward full implementation.
Localization technology relying on Radio Frequency (RF) signals holds immense promise across intelligent environmental monitoring, security surveillance, and navigation domains. However, existing technologies often overlook the tilt and rotation states of targets during real-world movement. This limitation in accurately capturing target movement dynamics restricts the precision of localization. The existing data-driven methods for localization are highly dependent on the quantity and quality of data. But in practical scenarios, collecting sufficient data can be challenging due to costs or other constraints. This paper proposes a data enhancement strategy that generates a large amount of data suitable for complex scenarios by simulating RF signals under various environmental conditions with embedded physical constraints. Then, this paper designs a deep learning-based RF signal processing method that can learn the complex mapping relationship between RF signals and target states. By training the neural network, we can extract the tilt and rotation information of the target from the RF signals, thereby achieving high-precision estimation of the target's location. The experimental results show that the method proposed in this paper can achieve high-precision localization of moving targets, even in the case of changes in the tag's direction and interference from environmental noise.
The convergence of Artificial Intelligence (AI) and the Internet of Things (IoT) is one of the main driving forces of the modern era of smart, autonomous systems. However, deploying AI models on resource-constrained IoT devices poses significant challenges in computation, memory, and power consumption. Model compression and sparsity have been a highly effective approach in the field of optimizing models while maintaining performance. But compared to traditional model compression techniques, its progress in Gen AI and IoT appears relatively slower. To address this research gap and provide insights into ways to promote the integration of model compression and sparsity of Gen AI models in IoT, we explore practical strategies including quantization, pruning, knowledge distillation, low rank factorization (LoRA, QLoRA), KV caching, etc. We aim to provide a road map that traces the evolution of model compression and sparsity from traditional small-scale models to the realm of Gen AI for efficient edge deployment across real-world IoT applications that includes the focus on the industrial applications and the research challenges for deploying intelligence in the IoT eco-system, particularly to enable edge intelligence to take care of the constraints of latency, privacy, and reliability.
Cooperative collision avoidance systems promise collision prevention by exchanging movement information between road users to alert them of imminent collisions. Precise global bicycle positioning is one key requirement in these systems. Existing technologies, such as GNSS, fall short of consistently delivering the required positioning accuracy in the real world. Therefore, this paper presents UWBike, a novel ubiquitous precise bicycle positioning system based on smartphone sensors that allows accurate and low-cost global positioning and position tracking. UWBike uses the smartphone's camera in combination with visual-inertial odometry (VIO) for relative position tracking. We show that UWBike is a feasible solution for short-term dead reckoning, suitable for the dynamic motion of bicycles, achieving an error rate of 1.3 cm/m in real-world experiments. To achieve a global position, we introduce a novel algorithm to convert the local coordinate system, derived from relative position tracking, into a global reference frame, taking advantage of a single ultra-wideband reference point. Our results demonstrate high accuracy, with a mean position error of only 36 cm, validating the real-life feasibility of our system for cooperative collision avoidance systems.
In the 2nd WEAR Dataset Challenge, held with HASCA 2025, the Sumitou Atsuya team aimed to achieve high performance with only simple preprocessing and rTsfNet, which is a distinctive state-of-the-art deep neural network (DNN) based human activity recognition (HAR) model proposed in 2024. The challenge was to predict 19 behaviors using acceleration data from one randomly selected sensor out of four sensors attached to each subject's limbs. The simple preprocessing we applied was axis inversion and the integration of left and right sensors separately for the arms and for the legs. Using the preprocessed data, we constructed arm and leg models with rTsfNet and optimized their hyperparameters with Optuna, respectively. In leave-one-subject-out cross-validation-based evaluation, the arm model achieved a macro-F1 score of 0.6213, and the leg model achieved a macro-F1 score of 0.5718. This suggests that modern models, when combined with simple preprocessing, demonstrate both effectiveness and practicality.