
This article presents the design and fabrication of PolarVisor markers. PolarVisor [1] is a system design featuring cost-effective, electronics-free fiducial markers for millimeter-wave (mmWave) radars. MmWave radars are widely used sensing platforms in automotive, robotic, and augmented reality applications. They transmit radio-frequency (RF) signals, then infer the range and velocity information of objects in a certain environment by collecting and analyzing the reflected signals. Because wireless signals can still operate reliably under occlusions such as mist or smoke, mmWave radars have seen increasing deployment in robots and vehicles; this ensures robust operations when traditional optics-based sensing fails.
When you ask me about a single piece of work or a single paper that I am most proud of, it is difficult to choose. There are two papers that significantly influenced my own way of thinking and my career, as well as the community as a whole. The first one is ''On the Lifetime of Sensor Networks'' published in 2009 in ACM Transactions on Sensor Networks (TOSN). This was a quite exciting endeavor, collaborating with Isabel Wagner. Working in the field of sensor networks, we figured that the results in many papers are just not comparable due to the lack of a unifying performance metric. We spent months analyzing used metrics and derived a system to integrate measures such as energy, packet loss, delay, and throughput. It was also a long journey in publishing the results (you know, many revisions and the typical Reviewer #2). But we were rewarded, as the research community picked up the ideas and our work helped shape the world of Internet of Things as we see it today. The second paper is ''Bidirectionally Coupled Network and Road Traffic Simulation for Improved IVC Analysis'' published in 2011 in IEEE Transactions on Mobile Computing. I had been working on vehicular networks at this time, and we realized that simulations always used either mobility traces or simple random mobility models. Both are either turning wrong the moment the application successfully impacts mobility or are unrealistic in the first place. Together with Christoph Sommer, I developed a simulator, Veins, which uses realistic vehicular mobility modeling coupled with accurate wireless communication simulation. The simulator still (after 15 years!) is the de facto standard in the community, certainly now being a community-driven open source project.
Providing fully immersive volumetric videos on mobile devices requires photo-realistic, full-scene rendering with smooth playback experiences. Since traditional 3D representations such as point clouds struggle with visual quality, 3D Gaussian Splatting (3DGS) [4] has emerged as an effective way to represent high-quality 3D scenes. However, existing approaches for 3DGS-based video streaming incur significant rendering overhead, making realtime playback challenging. This article presents Vega [5], a 3DGS-based system that enables fully immersive volumetric video streaming on smartphones. Vega bridges the gap between highquality 3DGS rendering and the constraints of mobile devices—specifically, limited network bandwidth and compute power. The core idea is object-level selective computation, which optimizes both data size and rendering overhead by prioritizing visually significant objects. To realize this idea, Vega introduces a mobile-friendly encoding scheme on the server, in which key frames store full scene data while residual frames encode only dynamic object information to reduce redundancy. On the client side, Vega employs a view-adaptive rendering pipeline that selectively renders visually important objects to meet strict real-time deadlines while effectively utilizing the smartphone's CPU, GPU, and NPU.
As Gaussian Splatting (GS) is increasingly integrated into products and production pipelines across industry, organizations and researchers have begun developing standardization practices to formally accommodate this emerging technology. Originally regarded primarily as a 3D rendering technique for novel view synthesis, 3DGS can be applied more broadly to represent a wide range of 3D computer graphics content. However, as GS is a relatively recent technology, much of its development remains at an experimental stage, resulting in a sparse standardization landscape. As 3D modeling becomes increasingly comprehensive and the applications and implementations of GS rapidly diversify, many organizations have recognized the need for standardization, introducing new formats such as SPZ or adapting GS to existing 3D asset standards like glTF. In this article, we review the initial standardization efforts led by organizations such as Metaverse Standards Forum, Niantic, Khronos, and MPEG, which aim to unify baseline testing conditions, performance metrics, and development practices for GS. We begin by outlining the implementation and evolution of GS, highlighting how it advances beyond previous 3D scene representation approaches. We then provide an overview of emerging splatting standards. Finally, we conclude by discussing future research directions for GS and examining its current and potential applications.
Ambient IoT (A-IoT), aiming to connect hundreds of billions of ultra-low-power and battery-free devices, has been included in the agenda for 5G-A and 6G standardization by 3GPP [1-2]. Ultra-low-power means the capability of maintenance-free operation for years. 3GPP technical report provides A-IoT's application scenarios such as automated warehousing, electronic shelf labels, forest fire monitoring, and smart farms. These crucial applications require key performance indicators (KPIs) including: (1) The power consumption shall be less than 1μW for passive devices (battery-free) and 1~100μW for semi-passive devices (no battery replenishment) to minimize maintenance cost. (2) The communication range of both downlink and uplink will be tens to hundreds of meters to cover most indoor and outdoor applications. (3) The network will support node density ranging from 1 to 15 per square meter which translates to hundreds to thousands of devices within a cell's coverage area.
Today, digital media is constantly produced and consumed in enormous volumes. We rely heavily on smartphone images and videos from daily social sharing and entertainment to critical tasks, such as verifying a new Uber driver's identity, online banking operations, or providing evidence in legal proceedings. However, continuous advances in digital media manipulation, especially with the introduction of generative AI, yield increasingly sophisticated deepfakes [14]. This poses a massive threat to society, facilitating the spread of fake news, misinformation, and personal slander that greatly endanger our perception of reality. Restoring trust in visual content has immense societal benefits, ensuring that organizations, institutions, and individuals can once again safely rely on the digital media they consume, restoring the principle of ''seeing is believing.'' A good solution must provide a reliable way to verify where, when, and how a piece of media was created, rather than relying solely on deepfake detection algorithms, which is unfortunately shaping up to be a never-ending arms race.
Edge AI integrates AI techniques with heterogeneous mobile devices to enable perception and actuation in real-world environments, facilitating applications such as smart sensing [1] and healthcare [2]. To reduce the burden of developing such applications, recent works [3, 4] build agentic systems based on Large Language Models (LLMs) to automatically translate user requirements into executable edge AI programs. However, these systems typically rely on handcrafted, predefined agentic workflows, and therefore often struggle to handle diverse mobile devices, heterogeneous runtime environments, and dynamic resource constraints in real-world scenarios. As a result, even well-engineered agents may struggle to accommodate such variability, suffering from inflexibility, limited adaptability, and high maintenance costs.
Human activity monitoring in the water is essential for pool management and drowning prevention. Existing camera-based solutions pose significant concerns about privacy and extra installation costs. Although sonars have been widely used for underwater sensing in open aquatic environments such as oceans and lakes, monitoring human activities with sonars in a pool setup is still challenging. In this work, we propose AquaScan, the first scanning sonar-based underwater sensing system for human activity monitoring. To overcome the low frame rate due to the sonar's physical limitation, we propose a novel scanning strategy and apply an image reconstruction method to accelerate the scanning speed without compromising the performance of motion detection. To overcome the dynamic interferences in the underwater scenario, we develop a novel signal processing pipeline based on a physical model to remove noises and localize human subjects. We further extract features like motion, time, and spatial information from sonar images and develop a state-transfer-based activity recognition system to recognize five common water activities, i.e., swimming, motionless, splashing, struggling, and drowning. We have deployed AquaScan on three public swimming pools for a total period of 94 hours. The evaluation results show that AquaScan can successfully recognize the five activities in the water with around 91.5%.
Recent advances in epidermal interfaces have enabled applications in biometric sensing, medical monitoring, and expressive interactions. While these systems demonstrate technical potential, their design and fabrication often rely on cleanrooms, photolithography, sputtering, or chemical etching - resources from microfabrication labs. As a result, the construction times of a single on-skin interface can take 3.5 to 11 hours [4, 5, 9], making tailor-made circuitry for individual bodies impractical for early-stage prototyping and iteration. Creating customized on-skin interfaces often involves designing precise layouts, routing conductive traces, and assembling components off-body before application. Yet the body's soft, nonplanar surface makes small layout changes difficult to accommodate or require full reconstruction [5, 6]. Repeated modifications further challenge skin conformability and long-term wearability [7, 8].
Immersive telepresence, a primary use case envisioned for 6G [1], holds the promise of revolutionizing remote communication by providing deeply engaging and interactive user experiences [2]. Immersive content that depicts 3D objects/scenes is typically represented by point clouds or meshes, allowing users to not only change view directions but also move freely in 3D space, known as six degrees of freedom (6DoF) motion. This capability has driven the adoption of immersive content across various domains, including healthcare, education, professional training, scientific data visualization, and entertainment [3].
It's common to use the face or fingerprint to unlock the smartphone or log into an app. While convenient, these methods can be fooled. Researchers have shown that things like high-quality photos, 3D-printed masks, or fake fingerprints can trick these systems [10]. To prevent this, developers often add extra security steps, like asking you to blink your eyes during a face scan [4]. Existing research has explored the integration of user authentication and liveness detection, such as speech-induced facial vibration [7]. However, those designs even require sophisticated hardware design or specialized sensors [10]. Even worse, among those using the onboard camera for collecting sensitive facial information as biometrics [2], users' privacy is inevitably compromised.
In the summer of 2025, Dan Halperin and his collaborators were awarded the SIGMOBILE Test of Time Award for their SIGCOMM 2010 paper entitled ''Predictable 802.11 Packet Delivery from Wireless Channel Measurements.''
The past few years have witnessed growing interest in millimeter wave (mmWave) based reconstruction in the mobile community [3, 5, 7, 14]. Unlike classical vision-based imaging systems, which are limited to line-of-sight, these mmWave-based systems can operate in through-occlusion scenarios, enabling them to sense objects in closed boxes and beneath clutter. This is because mmWave signals can traverse through many everyday occlusions (e.g., cardboard, fabric, etc.) [1, 11], and reflect off objects behind these occlusions, allowing them to produce images of the occluded objects. This capability, combined with the recent emergence of low-cost commercial mmWave radars, has the potential to enable many promising applications. For example, pick-and-place robots can leverage through-occlusion reconstructions to find and manipulate hidden objects, such as those beneath clutter or within a closed box. Similarly, Augmented Reality (AR) devices could leverage them to perceive occluded objects and display them to the user, truly augmenting our human perception. Smart home devices can use them for through-occlusion gesture recognition, to enable non-verbal commands even when users are hidden from view.
Passive Internet of Things (Passive IoT) has attracted widespread attention from both academia and industry due to its potential for ubiquitous deployment and round-theclock operation without the need for dedicated power sources [1]. As a key enabling technology for Passive IoT, backscatter communication has been an active research area for over a decade and has achieved remarkable progress [2, 3, 4]. We have witnessed the feasibility of implementing backscatter communication using various ambient energy sources [5, 6, 7, 8], with the communication range continuously extending from a few meters to reliably transmitting data over distances of more than one kilometer [9].
Mobile Augmented Reality (AR) applications demand high-quality, real-time visual prediction, including pixel-level depth and semantics, to enable immersive and context-aware user experiences. Recently, Vision Foundation Models (VFMs) have offered strong generalization capabilities on diverse and unseen data, supporting scalable mobile AR experiences. However, deploying VFMs on mobile devices is challenging due to computational limitations, particularly in maintaining both prediction accuracy and real-time performance. In this article, we present ARIA [3], the first system that enables on-device inference acceleration of a VFM. ARIA employs the heterogeneity of mobile processors through a parallel and selective inference scheme: full-frame prediction is periodically offloaded to a processor with high parallelism capability like GPU, while lowlatency updates on dynamic regions are conducted via a specialized accelerator like NPU. Implemented and evaluated using mobile devices, ARIA achieved significant improvements in accuracy and deadline success rate on real-world mobile AR scenarios.
When 5G first arrived, it came with big promises: ultra-high bandwidth, ultra-low latency, and a new generation of mobile experiences - from AR/VR to connected cars to early visions of the metaverse. Many of these applications were designed for the road, where consistency matters just as much as speed in maintaining high QoE. Yet the very design choices that give 5G its power - especially its use of higher-frequency spectrum - also make it fragile. Signals attenuate faster in higher frequency bands and coverage shrinks. And once you start moving, performance can vary wildly from one minute to the next.
Sound has connected people across distance for centuries, carrying voices, music, and emotion through the air. In the digital age, it connects people with their smartphones, bridging human communication and computational sensing. Today, smartphones use sound not only to transmit information but also to sense their surroundings. For example, voice assistants respond to spoken commands [10, 19, 21], while acoustic sensing techniques enable smartphones to detect hand gestures [1, 27], finger movements [13, 14], and even subtle physiological and behavioral signals such as respiration [3, 15], eye blinks [4, 16], and heartbeats [20, 26]. Together, these capabilities have transformed sound into a versatile sensing modality that allows smartphones to perceive and interpret the physical world.
The core principle of wireless sensing is that human activities and the environment physically alter the radio signals that travel through them. When a person moves or even breathes, they disturb the wireless signals, like Wi-Fi and mmWave, causing measurable changes in their properties [1]. By analyzing these subtle changes, our devices can learn to perceive the world without cameras, enabling applications from smart home control to healthcare monitoring. This transforms everyday devices into powerful privacy-preserving sensors.
Length is one of the seven fundamental physical quantities. For several centuries we have measured distances using calibrated physical objects, and more recently using light, sound, and radio waves. These measurements and the tools we use have enabled advances in several different domains, from the construction industry to space travel, from GPS localization to tracking of airplanes. With advances in electronics, clocks, miniaturization, and development of new algorithms, wireless distance measurement has now become possible. Measuring distances using wireless sensors offers the option of locating objects across rooms, through walls, and without visual line-of-sight. The measurement accuracy improves with larger bandwidth, which has resulted in the ultra-wideband (UWB) radio technology gaining significant traction. Seeing an opportunity in this capability, smartphone manufacturers such as Google and Apple have incorporated UWB radios in their offerings, setting the stage for future innovation using this versatile technology. [10] In this article we look beyond the UWB object-finding use-case and explore how a modest radio can transform mobile computing for decades to come.
Large-scale IoT deployment calls for inexpensive, low-power sensor nodes that still perform long-range, large-scale networking at the system level. In this paper, we propose a novel processor-sharing IoT architecture that converts the vast majority of sensor nodes from embedded computers to low-end RF peripherals. The conventional full-fledged sensor nodes are smashed into the air, and the scattered chips are scaled well with negligible overheads through a virtual I2C bus called RFBus.