Ambient intelligence, continuously understanding human presence, activity, and physiology in physical spaces, is fundamental to smart environments, health monitoring, and human-computer interaction. WiFi infrastructure provides a ubiquitous, always-on, privacy-preserving substrate for this capability across billions of IoT devices. Yet this potential remains largely untapped, as wireless sensing has typically relied on task-specific models that require substantial labeled data and limit practical deployment. We present AM-FM, the first foundation model for ambient intelligence and sensing through WiFi. AM-FM is pre-trained on 9.2 million unlabeled Channel State Information (CSI) samples collected over 439 days from 20 commercial device types deployed worldwide, learning general-purpose representations via contrastive learning, masked reconstruction, and physics-informed objectives tailored to wireless signals. Evaluated on public benchmarks spanning nine downstream tasks, AM-FM shows strong cross-task performance with improved data efficiency, demonstrating that foundation models can enable scalable ambient intelligence using existing wireless infrastructure.
In-house occupancy detection is vital for smart energy management, resource optimization, and home security. Traditional sensor-based solutions can be inaccurate and invasive. This paper presents a WiFi-based occupancy detection system that utilizes existing WiFi infrastructure and IoT devices. Our method employs a neural network with a shared CNN and a transformer block. Our preliminary evaluation, conducted using 18 unique IoT devices and data collected from 7 different homes over 42 days, demonstrates a detection accuracy of 94.88% with 8.12% false alarms in familiar environments, and 90.79% accuracy with 10.11% false alarms in new settings, significantly outperforming model-based methods.
WiFi-based home monitoring offers compelling advantages over traditional camera and sensor solutions by leveraging existing wireless infrastructure for contactless, privacy-preserving and through-the-wall detection. This paper presents insights from developing and deploying a WiFi-based human monitoring system across real-world residential environments, addressing the gap between academic research and practical deployment. Through a two-year study involving 280 edge devices across 15 homes in 11 U.S. states, collecting over 4 million motion samples, we identify and address four critical deployment challenges previously underexplored in academic settings: (1) false positives from non-human motion sources (pets, robots) that degrade system reliability, (2) hardware heterogeneity in commercial IoT devices causing inconsistent CSI quality, (3) signal interference in multi-user environments limiting individual tracking capabilities, and (4) computational and bandwidth constraints preventing real-time edge processing. The deployed system integrates a biomechanics-based classifier that reduces non-human false alarms from 63.1% to 8.4%, a multi-layer sensing quality metric validating device suitability without environment-specific calibration, proximity-based multi-user detection leveraging distributed IoT devices, and a hybrid edge-cloud architecture with ACF-based compression achieving 99.72% data reduction. The integrated system achieves 92.61% human motion detection accuracy across diverse uncontrolled home environments. We have successfully deployed home monitoring technology on millions of WiFi routers nationwide and smart IoT devices (e.g., bulbs, plugs) worldwide, demonstrating its viability for large-scale real-world applications. We share these findings to guide future research toward deployable WiFi sensing systems.
Child presence detection (CPD) is a vital technology for vehicles to prevent heat-related fatalities or injuries by detecting the presence of a child left unattended. Regulatory agencies around the world are planning to mandate CPD systems in the near future. However, existing solutions have limitations in terms of accuracy, coverage, and additional device requirements. While WiFi-based solutions can overcome the limitations, existing approaches struggle to reliably distinguish between adult and child presence, leading to frequent false alarms, and are often sensitive to environmental variations. In this paper, we present DeepCPD, a novel deep learning framework designed for accurate child presence detection in smart vehicles. DeepCPD utilizes an environment-independent feature-the auto-correlation function (ACF) derived from WiFi channel state information (CSI)-to capture human-related signatures while mitigating environmental distortions. A Transformer-based architecture, followed by a multilayer perceptron (MLP), is employed to differentiate adults from children by modeling motion patterns and subtle body size differences. To address the limited availability of in-vehicle child and adult data, we introduce a two-stage learning strategy that significantly enhances model generalization. Extensive experiments conducted across more than 25 car models and over 500 hours of data collection demonstrate that DeepCPD achieves an overall accuracy of 92.86 margin (79.55 children while maintaining a low false alarm rate of 6.14
Indoor intrusion detection systems (IDS) are crucial for securing residential and commercial spaces. However, existing solutions often require expensive, professional installations or suffer from high false alarm rates (FARs) due to non-human movements. We introduce the first indoor IDS that leverages ubiquitous WiFi signals to detect intrusions through walls while remaining robust against interference from non-human subjects. It comprises three key components: robust signal pre-processing to extract environment-independent statistics; a ResNet model for motion source identification; and an LSTM-based state machine to incorporate historical data. Notably, our system operates independently of device location, subject orientation, and environmental changes, enabling swift deployment in real-world environments. Implemented with single-pair commodity WiFi devices, it was extensively evaluated in two typical indoor settings with various interference sources such as pets, cleaning robots, and fans. Results demonstrate that our system achieves an intrusion detection accuracy of 96.67% with a false alarm rate of 1.92%. These findings highlight its robustness and significant potential for widespread deployment in indoor security applications.
Indoor tracking plays a critical role in a wide range of applications, yet existing solutions based on cameras, acoustics, or radar often face challenges related to privacy, deployment cost, and environmental sensitivity. WiFi-based methods offer a promising alternative by leveraging existing infrastructure, but most current approaches are active, requiring users to carry dedicated devices—limiting practicality in everyday scenarios. Passive WiFi tracking is more user-friendly, but existing solutions typically rely on complex feature engineering, require large training datasets, and struggle to generalize across different users and environments. In this work, we introduce a novel passive tracking system that requires only a single-shot training phase. By leveraging location signature based on statistical proximity metrics derived from CSI across multiple distributed WiFi devices, our method enables accurate, scalable, and training-efficient indoor tracking.
WiFi sensing has emerged as a compelling contactless modality for human activity monitoring by capturing fine-grained variations in Channel State Information (CSI). Its ability to operate continuously and non-intrusively while preserving user privacy makes it particularly suitable for health monitoring. However, existing WiFi sensing systems struggle to generalize in real-world settings, largely due to datasets collected in controlled environments with homogeneous hardware and fragmented, session-based recordings that fail to reflect continuous daily activity. We present CSI-Bench, a large-scale, in-the-wild benchmark dataset collected using commercial WiFi edge devices across 26 diverse indoor environments with 35 real users. Spanning over 461 hours of effective data, CSI-Bench captures realistic signal variability under natural conditions. It includes task-specific datasets for fall detection, breathing monitoring, localization, and motion source recognition, as well as a co-labeled multitask dataset with joint annotations for user identity, activity, and proximity. To support the development of robust and generalizable models, CSI-Bench provides standardized evaluation splits and baseline results for both single-task and multi-task learning. CSI-Bench offers a foundation for scalable, privacy-preserving WiFi sensing systems in health and broader human-centric applications.
Radio-frequency (RF)-based high-resolution human imaging is an emerging area of research fueled by the increasing availability of RF-radar devices. Even though existing works achieve accurate human body reconstruction for pose estimation purposes, human identification with imaging has not been feasible due to its limited resolution. In this work, we present high-resolution neural network (HRNet), a deep neural network based on conditional generative adversarial network architecture, to achieve high-resolution human silhouette images, which can be used for human identification. HRNet uses radar spatial spectrum generated using a modified multiple signal classification algorithm as input and is trained with Kinect images as ground truth. We tested our design using a commodity millimeter-wave radar device operating at 60 GHz. Experiments performed with 12 users in three different environments show that our proposed system can reconstruct human images with 4% mean silhouette difference when compared with Kinect images. Moreover, the system achieved an average classification accuracy of 90.6% for 12 users and 95.0% for seven users in unseen environments; thereby proving robustness to environment changes.
Indoor falls often lead to fatalities due to delayed assistance. Current approaches to detecting indoor falls, such as cameras and wearables, intrude on privacy and are inconvenient. Radar-based device-free sensing has a limited range and requires dense deployment, leading to overhead costs. WiFi-based solutions, while promising, are currently either environment-dependent or insufficiently tested. In this work, we propose a fusion approach that leverages signal processing techniques to extract environment-independent features from the Channel State Information (CSI) in commercial WiFi devices. We then use a neural network to detect differentiating patterns from these features. Our lightweight LSTM network, with just 21,000 parameters, has been tested on 2,400 fall events from over 25 volunteers in 5 environments. It has also undergone 21 months of false alarm testing in 6 diverse settings. The system achieves a 94.1% detection rate and fewer than 5 false alarms per month in single-person homes.
Passive indoor localization are critical for smart home intelligence services. Conventional methods using vision, acoustics, or radar face limitations in scalability and effectiveness due to their invasive nature. WiFi-based methods are emerging as a promising alternative because of the ubiquity of WiFi, its cost-efficiency, and non-intrusive manner. However, the performance of existing WiFi-based approaches in residential multi-room environments is often hampered by the limited bandwidth of standard WiFi devices. In this work, we introduce an innovative system that leverages commodity WiFi for precise human presence detection and room-level localization. Our system applies a novel multipath selection technique to focus a limited set of multipaths relevant to proximate motions to the device. We also introduce a spatial feature that capitalizes on multiple antennas, further improving spatial resolution and detection accuracy. Our method robustly integrates time, frequency, and spatial domain features for reliable room identification for human presence. The evaluations, considering real-world residential house and various walking patterns, validate the system’s robustness and high performance, achieving 87.62% test accuracy, an 88.03% test true positive rate, and an 88.74% test positive predictive value, surpassing state-of-the-art methods by over 25 %, showing its effectiveness in the real world.
Device-free indoor object detection and localization are essential for the success of smart homes. Traditional vision/acoustic/radar-based approaches face operational constraints that limit their effectiveness and scalability. WiFi-based approaches have recently been a promising candidate due to their ubiquity, cost-effectiveness, and privacy-preserving nature. However, most of them show inadequate performance in typical residential settings due to the limited WiFi bandwidth and the resulting low spatial resolution. In this article, we introduce a novel system using commodity WiFi that can accurately determine the specific room where the person is, i.e., room-level localization. The system employs a novel multipath selection technique to concentrate on a limited set of multipaths predominated by the proximate motions to the device. Based on the technique, a spatial feature leveraging multiple antennas to enhance the spatial resolution is proposed for more refined detection coverage. Combining the spatial feature with time- and frequency-domain features, the system is shown to achieve an overall test accuracy of 87.63%, a true positive rate of 89.47%, and a positive predictive value of 88.51%, outperforming state-of-the-art methods by >20% and showing its potential for real-world applications.
Indoor falls have proved fatal to many people due to a lack of timely assistance. Existing approaches for fall detection using cameras and wearable devices intrude on privacy and cause inconvenience. Passive sensing approaches using radar have limited coverage and demand dense deployment. Current solutions using commercial off-the-shelf (COTS) WiFi devices are either environment-dependent or lack extensive testing in real environments to confidently assess false alarm rates. In this work, we propose a fusion approach to detect falls with COTS WiFi, where we leverage signal processing techniques to extract environment-independent features, and use a neural network to detect differentiating patterns in those features. We designed a lightweight Long Short-Term Memory (LSTM)-based neural network with only 21 k parameters that can easily be deployed on edge devices. We further provide a framework to explain the network's behavior that supports a calibration-free design. Our proposed FallAware system's detection performance has been extensively tested on $\sim$ 2400 falls gathered from over 25 volunteers in 5 different environments. In addition, we conducted long-term false alarm testing in 6 diverse environments for a total duration of 21 months. The results show that FallAware can detect falls with an average detection rate of 94.1% in unseen environments with $< $ 5 false alarms per month in single-person occupancy homes.
Voice interfaces have become one of the most ubiquitous human-computer interaction methods in recent years. Voice activity detection (VAD) is typically the first building block of a complex voice interface, often relying on audio signals. Acoustics-based VAD systems do not perform well in noisy and interference-prone environments. Smart assistants mitigate this problem by using a dictionary-based detection system. However, this approach is limited in its applicability. For instance, users may still need to manually mute and unmute their microphones during online meetings to prevent detection of interfering users, and speech leakage. In order to automate voice detection in challenging environments without these limitations, we propose RadioVAD, a noise and interference-resilient VAD system that uses radio modality, which is already available in various smartphones and home assistants. RadioVAD works by detecting possible human presence in the Field of View of the device, extracting the vocal fold's vibration signal from the target speaker, and utilizing a time-domain neural network on raw radio signals to detect voice activity. Extensive experiments reveal that RadioVAD can detect voice activity in challenging environments with high accuracy and outperforms audio-based VAD when the audio signal has signal-to-noise ratio below 5 dB. Furthermore, RadioVAD reduces false alarm rate in interference-prone environments by 52%-72%, bringing significant improvements to VAD task. RadioVAD lays the foundation for future voice interfaces utilizing radio modality.
Indoor intelligent perception systems have gained significant attention in recent years. However, accurately detecting human presence can be challenging in the presence of non-human subjects such as pets, robots, and electrical appliances, limiting the practicality of these systems for widespread use. In this paper, we propose a novel system (“WI-MOID") that passively and unobtrusively distinguishes moving human and various non-human subjects using a single pair of commodity WiFi transceivers, without requiring any device on the subjects or restricting their movements. WI-MOID leverages a novel statistical electromagnetic wave theory-based multipath model to detect moving subjects, extracts physically and statistically explainable features of their motion, and accurately differentiates human and various non-human movements through walls, even in complex environments. In addition, WI-MOID is suitable for edge devices, requiring minimal computing resources and storage, and is environment-independent, making it easy to deploy in new environments with minimum effort. We evaluate the performance of WI-MOID in five distinct buildings with various moving subjects, including pets, vacuum robots, humans, and fans, and the results demonstrate that it achieves 97.34% accuracy and 1.75% false alarm rate for identification of human and non-human motion, and 95.98% accuracy in unseen environments without model tuning, demonstrating its robustness for ubiquitous use.
Respiration monitoring has been attracting substantial attention because of its potential for assessing sleep stages and quality. Traditional approaches for respiratory rate (RR) tracking require dedicated wearable devices, which can intrude upon and create an unwelcoming experience for users. To address this, researchers have proposed Wi-Fi-based respiration monitoring systems that capitalize on Wi-Fi's ubiquity, cost-effectiveness, and noncontact nature, effectively converting existing infrastructure into ubiquitous sensors. However, existing systems often struggle with the subtle signal-to-noise ratio of embedded breathing signals, leading to limited coverage and inflexible deployment in noisy environments. This article presents Wi-Fi respiration perception (WiResP), an ingenious and pragmatic Wi-Fi-based system for tracking respiration, employing the spectrum enhancement approach to boost respiration detection. By treating the spectrum of the breathing signal obtained from channel state information as an image, the system can leverage image processing techniques to greatly enhance the respiration signal trace, improving detectability and increasing sensing coverage for RR estimation. Moreover, an image-based continuity checker module is proposed to verify the signal trace continuity to reduce false alarms (FAs). We conducted extensive experiments and assessed WiResP under different settings. The experiments demonstrate that WiResP can reliably capture RR during sleep and achieve a detection rate of 92% and an FA rate <5%, leading to significantly improved sleep stage recognition under flexible device placements. The promising performance positions WiResP as a candidate for a real-world in-home respiration tracking system.
Addressing the pivotal challenge of discerning human and non-human activities in smart environments, in this demo, we present a system utilizing commercial WiFi transceivers for precise human and non-human motion differentiation through the walls. This system effectively filters non-human interference in smart home systems by extracting physically and statistically explainable features from ubiquitous WiFi signals. It passively recognizes moving subjects in real time without constraining their movement, even in complex environments. Tailored for edge computing, it ensures minimal resource consumption and generalizes well across various settings. Our long-term field tests confirm a high accuracy rate of 97.34% and a low false alarm rate of 1.75%, underscoring its robustness and readiness for practical deployment. Please find the companion video with the URL: https://youtu.be/6xkJZ_VvL9Q.
As WiFi becomes increasingly pervasive in communications, its role in sensing applications is likewise expanding. However, current WiFi-based sensing technologies often operate under the limit assumption that all detected motion originates from human activities, there by neglecting influences from non-human subjects. Being able to differentiate human motions from non-human ones is essential in many application use cases. This paper presents a deep learning framework that can accurately recognize human and various non-human moving subjects using single-pair WiFi devices, even through the walls. Utilizing environment-invariant features, the framework is tested across three settings with commodity WiFi devices and various deep neural network architectures for four-class recognition. Achieving an average validation accuracy of 95.57% and an average testing accuracy of 87.09% in unseen environments with a challenging dataset, our approach demonstrates its robustness and readiness for integration into intelligent IoT systems and applications.