TouchFusion is a wristband that enables touch interactions on nearby surfaces without any additional instrumentation or computer vision. TouchFusion combines surface electromyography (sEMG), bioimpedance, inertial, and optical sensing to capture multiple facets of hand activity during touch interactions. Through a combination of early and late fusion, TouchFusion enables stateful touch detection on both environmental and body surfaces, simple surface gestures, and tracking functionality for contextually adaptive interfaces as well as basic trackpad-like interactions. We validate our approach on a dataset of 100 participants, significantly exceeding the population size of typical wearable sensing studies to capture a wider variance of wrist anatomies, skin conductivities, and behavioral patterns. We show that TouchFusion can enable several common touch interaction tasks. Using TouchFusion, a wearer can summon a trackpad on any surface, control contextually adaptive interfaces based on where they tap, or use their palm as an always-available touch surface. When paired with smart glasses or augmented reality devices, TouchFusion enables a ubiquitous, contextually adaptive interaction model.
We showcase a demo using ElectroRing 2.0, a multimodal, finger-worn input device for seamless, everyday AR/VR/AI interactions. Building on the fundamental sensing principle of ElectroRing [1], it integrates galvanic, inertial, and optical sensors to detect a rich set of micro-gestures, including scrolling, dial rotation, and more. ElectroRing 2.0 achieves over 95% precision and sub-40 ms latency, enabling robust, high-bandwidth input for wearable systems. We demonstrate ElectroRing 2.0’s real-time performance and discuss its potential for next-generation AR/VR/AI interfaces. This is an interactive demo in which users perform a variety of mid-air micro-gestures to control an AR interface in real time. The demo highlights ElectroRing 2.0’s responsiveness, versatility, and ease of calibration, illustrating its potential for intuitive, high-bandwidth wearable input in practical scenarios.
Radar is more resilient to adverse weather and lighting conditions than visual and Lidar simultaneous localization and mapping (SLAM). However, most radar SLAM pipelines still rely heavily on frame-to-frame odometry, which leads to substantial drift. While loop closure can correct long-term errors, it requires revisiting places and relies on robust place recognition. In contrast, visual odometry methods typically leverage bundle adjustment (BA) to jointly optimize poses and map within a local window. However, an equivalent BA formulation for radar has remained largely unexplored. We present the first radar BA framework enabled by Gaussian Splatting (GS), a dense and differentiable scene representation. Our method jointly optimizes radar sensor poses and scene geometry using full range-azimuth-Doppler data, bringing the benefits of multi-frame BA to radar for the first time. When integrated with an existing radar-inertial odometry frontend, our approach significantly reduces pose drift and improves robustness. Across multiple indoor scenes, our radar BA achieves substantial gains over the prior radar-inertial odometry, reducing average absolute translational and rotational errors by 90
Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To address these, we first present the EgoMAN dataset, a large-scale egocentric dataset for interaction stage-aware 3D hand trajectory prediction with 219K 6DoF trajectories and 3M structured QA pairs for semantic, spatial, and motion reasoning. We then introduce the EgoMAN model, a reasoning-to-motion framework that links vision-language reasoning and motion generation via a trajectory-token interface. Trained progressively to align reasoning with motion dynamics, our approach yields accurate and stage-aware trajectories with generalization across real-world scenes.
Knowing when a user picks up an object plays a vital role in many context-aware applications. For example, tracking water consumption, counting calories consumed, or reminding you to bring your keys are all context-centered scenarios involving picking up objects. In this project, we propose Contextra, a wrist-worn system that uses sensor fusion to recognize when a user grasps objects. Sensor fusion allows all parts of the grasp to be sensed in ways single channels cannot alone. In our wristband, we fuse EMG and IMU data with video captured from three low-power IR cameras. These cameras maintain privacy by using an active-illumination technique to only capture features close to the sensors. Beyond grasps alone, we see Contextra as playing a foundational role in providing continuous awareness of context triggers to extend the functionality of existing AI devices that cannot run continuously due to power and privacy concerns.
Smartwatches have firmly established themselves as a popular wearable form factor. The potential expansion of their interaction space to nearby surfaces offers a promising avenue for enhancing input accuracy and usability beyond the confines of a small screen. However, a key challenge is in detecting continuous contact states with the surface to inform the start and end of stateful interactions. In this paper, we introduce SoundScroll, enabling a rapid and precise determination of contact state and fingertip speed of sliding finger. We leverage vibrations from friction between a moving finger and a surface. Our proof-of-concept wristband captures a dual-channel vibration signal for robust sensing, considering both on-skin and inair components. Our software predicts a finger sliding state as fast as 20 ms with an accuracy of 93.3%. Augmenting prior approaches detecting tap events, SoundScroll can be a robust, low-latency, and precise contact and motion sensing technique.
Smart rings for subtle, reliable finger input offer an attractive path for ubiquitous interaction with wearable computing platforms. However, compared to ordinary rings worn for cultural or fashion reasons, smart rings are much bulkier and less comfortable, largely due to the space required for a battery, which also limits the space available for sensors. This paper presents picoRing, a flexible sensing architecture that enables a variety of battery-free smart rings paired with a wristband. By inductively connecting a wristband-based sensitive reader coil with a ring-based fully-passive sensor coil, picoRing enables the wristband to stably detect the passive response from the ring via a weak inductive coupling. We demonstrate four different rings that support thumb-to-finger interactions like pressing, sliding, or scrolling. When users perform these interactions, the corresponding ring converts each input into a unique passive response through a network of passive switches. Combining the coil-based sensitive readout with the fully-passive ring design enables a tiny ring that weighs as little as 1.5 g and achieves a 13 cm stable readout despite finger bending, and proximity to metal.
We present OptiRing, a ring-based wearable device that enables subtle single-handed micro-interactions using low-resolution camera-based sensing. We demonstrate that our approach can work with ultra-low image resolutions (e.g., 5 × 5 pixels), which is instrumental in addressing privacy concerns and reducing computational needs. Using a miniature camera, OptiRing supports thumb-to-index finger gestures, such as stateful pinch and left/right swipes, as well as continuous 1-DOF input. We present a modeling approach that uses heuristic-based methods to identify interactions and machine learning for input gating, generalizing recognition across users and sessions. We assess this technique’s capabilities, accuracy, and limitations through a user study with 15 participants. OptiRing achieved 93.1% accuracy in gesture recognition, 99.8% accuracy for stateful pinch gestures, and a minimal number of false positives. Further, we validate OptiRing’s ability to handle continuous 1D input in a Fitts’ law study. We discuss these findings, the tradeoff between resolution and interaction accuracy, and the potential of this technology, which can help address challenges associated with optical sensing for input.
Wearable computing platforms, such as smartwatches and head-mounted mixed reality displays, demand new input devices for high-fidelity interaction. We present AuraRing, a wearable magnetic tracking system designed for tracking fine-grained finger movement. The hardware consists of a ring with an embedded electromagnetic transmitter coil and a wristband with multiple sensor coils. By measuring the magnetic fields at different points around the wrist, AuraRing estimates the five degree-of-freedom pose of the ring. AuraRing is trained only on simulated data and requires no runtime supervised training, ensuring user and session independence. It has a dynamic accuracy of 4.4 mm, as measured through a user evaluation with optical ground truth. The ring is completely self-contained and consumes just 2.3 mW of power
Wearable computing platforms, such as smartwatches and head-mounted mixed reality displays, demand new input devices for high-fidelity interaction. We present AuraRing, a wearable magnetic tracking system designed for tracking fine-grained finger movement. The hardware consists of a ring with an embedded electromagnetic transmitter coil and a wristband with multiple sensor coils. By measuring the magnetic fields at different points around the wrist, AuraRing estimates the five degree-of-freedom pose of the ring. AuraRing is trained only on simulated data and requires no runtime supervised training, ensuring user and session independence. It has a resolution of 0.1 mm and a dynamic accuracy of 4.4 mm, as measured through a user evaluation with optical ground truth. The ring is completely self-contained and consumes just 2.3 mW of power.
On-body IMU-based pose tracking systems have gained prevalence over their external tracking counterparts due to their mobility, ease of installation and use. However, even in these systems, an IMU sensor placed on a particular joint can only estimate the pose of that particular limb. In contrast, activity recognition systems contain insights into the whole body's motion dynamics. In this work, we present ActivityPoser, which uses the activity context as a conditional input to estimate the pose of limbs for which we do not have any direct sensor data. ActivityPoser compensates for impoverished sensing paradigms by reducing the overall pose error by up to 17%, compared to a model bereft of activity context. This highlights a pathway to high-fidelity full-body digitization with minimal user instrumentation.
We introduce RotoWrist, an infrared (IR) light based solution for continuously and reliably tracking 2-degree-of-freedom (DoF) relative angle of the wrist with respect to the forearm using a wristband. The tracking system consists of eight time-of-flight (ToF) IR light modules distributed around a wristband. We developed a computationally simple tracking approach to reconstruct the orientation of the wrist without any runtime training, ensuring user independence. An evaluation study demonstrated that RotoWrist achieves a cross-user median tracking error of 5.9° in flexion/extension and 6.8° in radial and ulnar deviation with no calibration required as measured with optical ground truth. We further demonstrate the performance of RotoWrist for a pointing task and compare it against ground truth tracking.
We present ElectroRing, a wearable ring-based input device that reliably detects both onset and release of a subtle finger pinch, and more generally, contact of the fingertip with the user's skin. ElectroRing addresses a common problem in ubiquitous touch interfaces, where subtle touch gestures with little movement or force are not detected by a wearable camera or IMU. ElectroRing's active electrical sensing approach provides a step-function-like change in the raw signal, for both touch and release events, which can be easily detected using only basic signal processing techniques. Notably, ElectroRing requires no second point of instrumentation, but only the ring itself, which sets it apart from existing electrical touch detection methods. We built three demo applications to highlight the effectiveness of our approach when combined with a simple IMU-based 2D tracking system.
Gaze tracking is an essential component of next generation displays for virtual reality and augmented reality applications. Traditional camera-based gaze trackers used in next generation displays are known to be lacking in one or multiple of the following metrics: power consumption, cost, computational complexity, estimation accuracy, latency, and form-factor. We propose the use of discrete photodiodes and light-emitting diodes (LEDs) as an alternative to traditional camera-based gaze tracking approaches while taking all of these metrics into consideration. We begin by developing a rendering-based simulation framework for understanding the relationship between light sources and a virtual model eyeball. Findings from this framework are used for the placement of LEDs and photodiodes. Our first prototype uses a neural network to obtain an average error rate of 2.67° at 400 Hz while demanding only 16 mW. By simplifying the implementation to using only LEDs, duplexed as light transceivers, and more minimal machine learning model, namely a light-weight supervised Gaussian process regression algorithm, we show that our second prototype is capable of an average error rate of 1.57° at 250 Hz using 800 mW.
The ability to track handheld controllers in 3D space is critical for interaction with head-mounted displays, such as those used in virtual and augmented reality systems. Today's systems commonly rely on dedicated infrastructure to track the controller or only provide inertial-based rotational tracking, which severely limits the user experience. Optical inside-out systems offer mobility but require line-of-sight and bulky tracking rings, which limit the ubiquity of these devices. In this work, we present Aura, an inside-out electromagnetic 6-DoF tracking system for handheld controllers. The tracking system consists of three coils embedded in a head-mounted display and a set of orthogonal receiver coils embedded in a handheld controller. We propose a novel closed-form and computationally simple tracking approach to reconstruct position and orientation in real time. Our handheld controller is small enough to fit in a pocket and consumes 45 mW of power, allowing it to operate for multiple days on a typical battery. An evaluation study demonstrates that Aura achieves a median tracking error of 5.5 mm and 0.8 degrees in 3D space within arm's reach.
We present AuraRing, a wearable electromagnetic tracking system for fine-grained finger movement. The hardware consists of a ring with an embedded electromagnetic transmitter coil and a wristband with multiple sensor coils. By measuring the magnetic fields at different points around the wrist, AuraRing estimates the five degree-of-freedom pose of the finger. AuraRing is trained only on simulated data and requires no runtime supervised training, ensuring user and session independence. AuraRing has a resolution of 0.1 mm and a dynamic accuracy of 4.4 mm, as measured through a user evaluation with optical ground truth. The ring is completely self-contained and consumes just 2.3 mW of power.
Augmented reality (AR) promises to revolutionize the way people interact with their surroundings by seamlessly overlaying virtual information onto the physical world. To improve the quality of such information, AR systems need to identify the object with which the user is interacting. AR systems today heavily rely on computer vision for object identification; however, state-of-the-art computer vision systems can only identify the general object categories, rather than their precise identity. In this work, we propose IDCam, a system that fuses RFID and computer vision for precise item identification in AR object-oriented interactions. IDCam simultaneously tracks users’ hands using a depth camera and generates motion traces for RFID-tagged objects. The system then correlates traces from vision and RFID to match item identities with user interactions. We tested our system through a simulated retail scenario where 5 participants interacted with a clothing rack simultaneously. In our evaluation study deployed in a lab environment, IDCam identified item interactions with an accuracy of 82.0% within 2 seconds.
We present AuraRing, a wearable electromagnetic tracking system for fine-grained finger movement. The hardware consists of a ring with an embedded electromagnetic transmitter coil and a wristband with multiple sensor coils. By measuring the magnetic fields at different points around the wrist, AuraRing estimates the five degree-of-freedom pose of the finger. AuraRing is trained only on simulated data and requires no runtime supervised training, ensuring user and session independence. AuraRing has a resolution of 0.1 mm and a dynamic accuracy of 4.4 mm, as measured through a user evaluation with optical ground truth. The ring is completely self-contained and consumes just 2.3 mW of power.
Christian Holz合作论文数Department of Computer Science, Eidgenössische Technische Hochschule Zürich;Sensing, Interaction & Perception Lab, Eidgenössische Technische Hochschule Zürich5