3D Gaussian Splatting (3DGS) has achieved remarkable success in novel view synthesis; however, reconstructions under sparse views often exhibit noticeable artifacts. While recent video diffusion models provide strong spatio-temporal priors for 3DGS restoration, directly fine-tuning them for restoration is suboptimal, as they lack awareness of the underlying multi-camera geometry, resulting in multi-view inconsistencies. In this work, we propose a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction. Specifically, we construct a large-scale 3DGS video dataset to enable specialized fine-tuning. To bridge the gap between 2D video generation and 3D multi-view constraints, we introduce a camera-conditioned geometric prior. By using the first and last frames as boundary anchors and encoding the corresponding camera relationships, we explicitly inject spatial structure into the video generation pipeline. This boundary-anchored, camera-aware prior guides the network toward geometrically grounded restoration that remains coherent across viewpoints. Extensive experiments show that, among video-prior restoration methods, our approach attains the best pixel- and structure-level fidelity (PSNR/SSIM) and improves multi-view consistency, while remaining competitive in perceptual quality (LPIPS).
BACKGROUND:Older adults managing chronic illnesses, such as cancer and Alzheimer disease and related dementias (ADRD), often experience significant physical or cognitive impairments that hinder daily activities and increase caregiver burden. Smart Internet of Things (IoT) technologies offer promising solutions by enabling passive monitoring, timely reminders, and personalized support at home. However, these technologies must be carefully tailored to accommodate users' individualized needs and preferences. OBJECTIVE:This formative qualitative study aimed to explore stakeholder perspectives, including patients, caregivers, health care providers, and technical experts, on the use of smart home-based IoT systems to support chronic illness management. The goal was to inform the early development of the audio and radio connected (AURA) system, an IoT prototype integrating Wi-Fi sensing, wearable trackers, and voice-assistive features. METHODS:Semistructured interviews were conducted with 6 patients who underwent postostomy creation for colorectal or bladder cancer treatment and 5 patients with ADRD and their caregivers. Input from additional stakeholders, including 2 health care providers, 2 community health workers, and 2 computer scientists, was also included in the report. Stakeholders reviewed a demonstration video depicting the conceptual features of the AURA system. Interviews explored stakeholders' needs and preferences for using such systems. Thematic analysis was guided by the extended Unified Theory of Acceptance and Use of Technology 2 (UTAUT2) framework, with 5 adapted constructs: performance expectancy, effort expectancy, social influence, facilitating conditions, and hedonic motivation and habit. RESULTS:Stakeholders identified distinct yet complementary needs across populations. Patients with cancer emphasized physical health monitoring, integration with health care systems, and customization; ADRD stakeholders prioritized routine support, emotional engagement, and simplicity; caregivers and clinicians emerged as key influencers of adoption. Barriers included privacy concerns, technology literacy, and fatigue, while facilitators included perceived caregiving support, streamlined interfaces, and electronic health record integration. Patients with cancer focused on motivational cues for physical activity, while emotional engagement and habit were more prominent for ADRD users. CONCLUSIONS:Stakeholder insights underscore the importance of designing adaptable, user-centered IoT systems that reflect the varied capabilities and care needs of older adults with chronic illnesses. These findings informed the design of the AURA prototype and highlighted theoretical considerations for technology acceptance in health care. Future work will test AURA in real-world settings to evaluate usability, acceptability, and clinical relevance.
In-home health activity monitoring using environmental sensors enables continuous and privacy-preserving care; however, real-world deployments face two fundamental challenges: (1) unpredictable partial sensor failures and corrupted modalities can silently degrade recognition performance, and (2) fault-labeled training data are often scarce or entirely unavailable. To address these challenges, we propose SAFE, a unified Sensor-Aware Fusion Engine for fault-resilient multi-modal activity recognition during inference. The core idea of SAFE is a dynamic reliability-driven reweighting mechanism: instead of treating modalities equally, SAFE estimates per-modality health and applies a softmax-normalized weighting scheme to reallocate fusion importance, automatically suppressing degraded sensors and amplifying reliable ones. To operate in data-insufficient setting, SAFE further leverages modality-specific self-reconstruction signals trained only on normal data to infer degradation without requiring fault annotations. Focus on a limited set of environmental sensors, we evaluate SAFE across five fault types, including single-sensor and dual-sensor faults, under both data-sufficient and data-insufficient settings. SAFE consistently and substantially outperforms a traditional late-fusion baseline. In controlled validation experiment, SAFE achieves 97.58% and 94.68% accuracy under single-sensor faults in data-sufficient and data-insufficient settings, respectively, compared to 90.33% for the baseline. Under dual-sensor faults, SAFE maintained 96.54% and 86.69% accuracy, substantially higher than the baseline’s 83.18%. Similar robustness gains are observed in a restroom usage case study. These results demonstrate that SAFE transforms multi-modal fusion from failure-prone aggregation into reliability-aware adaptive integration, providing a practical and generalizable solution for fault-resilient in-home health activity monitoring.
Leveraging the widespread deployment of IoT devices as a substitute for dedicated sensing hardware is highly promising. However, a key challenge is that their limited storage and computational resources often cannot meet the requirements of sensing tasks. In this paper, we propose EMRC, an Efficient Channel Measurement scheme for Resource-Constrained IoT devices. First, we design a channel inversion mechanism to significantly reduce the execution time and memory overhead of channel measurement. Then, we develop a parallel channel measurement mechanism for multiple responders, which improves sensing performance while conserving valuable channel resources. Furthermore, we address practical deployment challenges in EMRC, including the optimization of deeply faded subcarriers and the peak-to-average power ratio (PAPR). Finally, we evaluate EMRC on software-defined radios and several commercial IoT modules. Experimental results under periodic HE-LTF sequence transmission show that the proposed scheme improves the sensing SNR by 29% when four devices perform simultaneous sensing. It also reduces the computation time by up to 56% on ESP32, 21.5% on ESP8266, and 15.6% on STM32, while saving approximately one CSI frame of memory for each sensing link.
Magnetic field sensing is essential for electrical safety and early hazard detection. While existing solutions are generally considered reliable and accurate, they are hindered by battery-dependent challenges, invasive detection, and poor scalability. To address these limitations, in this paper, we propose a novel μT-level magnetic field sensing method based on RFID tags, named RF-Gaussmeter. The design of RF-Gaussmeter involves embedding a Tunnel Magnetoresistance (TMR) sensor into the COTS RFID tags. We consider the magnetic field generated by current-carrying conductors as the sensing target. When the TMR sensor captures the magnetic field generated by a conductor, it can output the sensing voltage signal to the RFID tag, altering the chip’s impedance and thereby influencing the received signal. To amplify weak magnetic fields, we explore the interaction model between the MOSFET in RFID tags and differential voltage signals. Based on this model, we propose a differential signal amplification design based on RFID to enhance sensing sensitivity. To extract the magnetic-field-related sensing features, we develop a self-differential module based on RFID to remove multi-path interference. Experimental results show that RF-Gaussmeter achieves an absolute error of 0.0825A for current sensing and 1.32μT for magnetic field sensing within 80cm range, demonstrating excellent practicality and accuracy.
The Mixture of Experts (MoE) architecture has emerged as a key technique for scaling Large Language Models by activating only a subset of experts per query. Deploying MoE on consumer-grade edge hardware, however, is constrained by limited device memory, making dynamic expert offloading essential. Unlike prior work that treats offloading purely as a scheduling problem, we leverage expert importance to guide decisions, substituting low-importance activated experts with functionally similar ones already cached in GPU memory, thereby preserving accuracy. As a result, this design reduces memory usage and data transfer, while largely eliminating PCIe overhead. In addition, we introduce a scheduling policy that maximizes the reuse ratio of GPU-cached experts, further boosting efficiency. Extensive evaluations show that our approach delivers 48
Rotation is a fundamental form of motion and rotation speed measurement holds paramount importance for assessing the health and performance of machinery with rotating components. However, existing measurement systems often face challenges such as limited measurement distance, low accuracy, and complex installation or maintenance processes. In this paper, we propose RoLEX, a LoRa-based rotation speed measurement system for long-distance and contactless monitoring of rotating machinery in ubiquitous scenarios. RoLEX employs a novel Signal Selection method to eliminate chirp interference and adapt to varying rotation speeds, along with a Boost Sensing method to enhance sampling rates and an advanced feature processing algorithm for precise rotation speed estimation and tracking. Comprehensive experiments validate that RoLEX achieves a measurement distance of 50 m, approximately 17 times farther than the latest wireless rotation speed measurement systems. Moreover, RoLEX is robust to interference and obstructions (including through-wall scenarios) and achieves an average measurement error less than 0.69% across different rotation speeds (100-5100 Revolutions Per Minute). For tracking performance, RoLEX achieves a relative error less than 2.8% in 90% of cases. We also present a case study to highlight RoLEX's practical applicability in real-world scenarios.
Mixture-of-Experts architectures have become the standard for scaling large language models due to their superior parameter efficiency. To accommodate the growing number of experts in practice, modern inference systems commonly adopt expert parallelism to distribute experts across devices. However, the absence of explicit load balancing constraints during inference allows adversarial inputs to trigger severe routing concentration. We demonstrate that out-of-distribution prompts can manipulate the routing strategy such that all tokens are consistently routed to the same set of top-k experts, which creates computational bottlenecks on certain devices while forcing others to idle. This converts an efficiency mechanism into a denial-of-service attack vector, leading to violations of service-level agreements for time to first token. We propose RepetitionCurse, a low-cost black-box strategy to exploit this vulnerability. By identifying a universal flaw in MoE router behavior, RepetitionCurse constructs adversarial prompts using simple repetitive token patterns in a model-agnostic manner. On widely deployed MoE models like Mixtral-8x7B, our method increases end-to-end inference latency by 3.063x, degrading service availability significantly.
Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to interpret music sheets remains underexplored. To bridge this gap, we introduce MusiXQA, the first comprehensive dataset for evaluating and advancing MLLMs in music sheet understanding. MusiXQA features high-quality synthetic music sheets generated via MusiXTeX, with structured annotations covering note pitch and duration, chords, clefs, key/time signatures, and text, enabling diverse visual QA tasks. Through extensive evaluations, we reveal significant limitations of current state-of-the-art MLLMs in this domain. Beyond benchmarking, we developed Phi-3-MusiX, an MLLM fine-tuned on our dataset, achieving significant performance gains over GPT-based methods. The proposed dataset and model establish a foundation for future advances in MLLMs for music sheet understanding. Code, data, and model will be released upon acceptance.
In-network aggregation (INA) offloads gradient aggregation onto switches, and thus effectively reduces the aggregation latency and the volume of traffic. However, INA resources are limited due to the high cost of on-chip memory, which imposes distinct challenges to the effective scheduling of these resources in multi-job MLaaS scenarios. In this paper, we explore the scheduling of INA resources in spatial and temporal dimensions, specifically focusing on its impact on the average job completion time (JCT) and the efficiency of INA resources. We propose Mina,an innovative co-design of algorithm and system that intelligently assigns INA resources to each job and effectively schedules these resources among multiple jobs. Our experiments show that Minaattains an INA efficiency score of 0.9998, implying that almost all jobs run nearly as efficiently as they would with exclusive INA acceleration.
In-network aggregation (INA) offloads gradient aggregation onto switches, and thus effectively reduces the aggregation latency and the volume of traffic. However, INA resources are limited due to the high cost of on-chip memory on switches, which imposes distinct challenges to the effective scheduling of these resources in multi-job Machine Learning as a Service (MLaaS) scenarios. In this paper, we explore the scheduling of INA resources in spatial and temporal dimensions, specifically focusing on its impact on the average job completion time (JCT) and the efficiency of INA resources. We propose Mina, an innovative co-design of algorithm and system that intelligently assigns INA resources to each job and effectively schedules these resources among multiple jobs. Our experiments show that Mina attains an INA efficiency score of 0.9099 on average, 2.67x higher than the baseline, implying that almost all jobs run nearly as efficiently as they would with exclusive INA acceleration. Furthermore, Mina proves to be highly adaptable to varied cluster configurations and incurs only minimal additional overhead.
The emergence of LLMs has ignited a fresh surge of breakthroughs in NLP applications, particularly in domains such as question-answering systems and text generation. As the need for longer context grows, a significant bottleneck in model deployment emerges due to the linear expansion of the Key-Value (KV) cache with the context length. Existing methods primarily rely on various hypotheses, such as sorting the KV cache based on attention scores for replacement or eviction, to compress the KV cache and improve model throughput. However, heuristics used by these strategies may wrongly evict essential KV cache, which can significantly degrade model performance. In this paper, we propose QAQ, a Quality Adaptive Quantization scheme for the KV cache. We theoretically demonstrate that key cache and value cache exhibit distinct sensitivities to quantization, leading to the formulation of separate quantization strategies for their non-uniform quantization. Through the integration of dedicated outlier handling, as well as an improved attention-aware approach, QAQ achieves up to 10x the compression ratio of the KV cache size with a negligible impact on model performance. QAQ significantly reduces the practical hurdles of deploying LLMs, opening up new possibilities for longer-context applications. We make our code publicly available(1) to support reproducibility and promote broader awareness within the community.
The rapid development of large language models (LLMs) has significantly advanced code completion capabilities, giving rise to a new generation of LLM-based Code Completion Tools (LCCTs). Unlike general-purpose LLMs, these tools possess unique workflows, integrating multiple information sources as input and prioritizing code suggestions over natural language interaction, which introduces distinct security challenges. Additionally, LCCTs often rely on proprietary code datasets for training, raising concerns about the potential exposure of sensitive data. This paper exploits these distinct characteristics of LCCTs to develop targeted attack methodologies on two critical security risks: jailbreaking and training data extraction attacks. Our experimental results expose significant vulnerabilities within LCCTs, including a 99.4% success rate in jailbreaking attacks on GitHub Copilot and a 46.3% success rate on Amazon Q. Furthermore, We successfully extracted sensitive user data from GitHub Copilot, including 54 real email addresses and 314 physical addresses associated with GitHub usernames. Our study also demonstrates that these code-based attack methods are effective against general-purpose LLMs, highlighting a broader security misalignment in the handling of code by modern LLMs. These findings underscore critical security challenges associated with LCCTs and suggest essential directions for strengthening their security frameworks.
The remarkable performance of large language models (LLMs) in various language tasks has attracted considerable attention. However, the ever-increasing size of these models presents growing challenges for deployment and inference. Structured pruning, an effective model compression technique, is gaining increasing attention due to its ability to enhance inference efficiency. Nevertheless, most previous optimization-based structured pruning methods sacrifice the uniform structure across layers for greater flexibility to maintain performance. The heterogeneous structure hinders the effective utilization of off-the-shelf inference acceleration techniques and impedes efficient configuration for continued training. To address this issue, we propose a novel masking learning paradigm based on minimax optimization to obtain the uniform pruned structure by optimizing the masks under sparsity regularization. Extensive experimental results demonstrate that our method can maintain high performance while ensuring the uniformity of the pruned model structure, thereby outperforming existing SOTA methods.
With the rising of demands for novel Human-Computer Interaction (HCI) approaches in the 3D space, a number of intelligent approaches have been proposed to achieve the HCI by tracking the translation and rotation of the target devices. In this article, we propose to realize a light-weight, battery-free, 3D motion tracking solution by leveraging a spinning linearly polarized antenna to track a passive RFID tag array. Instead of using the fixed antennas, which can only receive stable signal in some specific environments due to the unpredictable multipath effect, we propose to mitigate the multipath effect and the ambient interference by continuously spinning a linearly polarized antenna, and then extract the most distinctive features based on the optimal reading conditions of the spinning antenna. In particular, because the phase variation around the matching direction is more stable while the RSSI variation around the mismatching direction is more distinctive, we leverage such matching/mismatching property of the linearly polarized antenna to extract the most distinctive features for motion tracking. To depict the property, we build a theoretical model to explain the RSSI and the phase variation of the RFID tag along with the spinning of the antenna, and further extend the model from a single RFID tag to an RFID tag array. Based on the model, we can extract the distinctive RSSI features for the rotation tracking and the stable phase features for the translation tracking. Moreover, to tackle the low rate of feature extraction due to the spinning of antenna, we further propose to enhance the unstable phase features based on the overall trend of other tags with interpolation, such that the sampling rate can be efficiently improved. Finally, we propose a LSTM (Long Short Term Memory)-based network to track the 3D motion based on the signal features extracted based on the polarization model. The experimental results show that our system can achieve an average error of 10.45 cm in the translation tracking, and an average error of $6.02(degrees) in the rotation tracking in the 3D space.
We propose UltraCLR, a new contrastive learning framework that fuses dual modulation ultrasonic sensing signals to enhance gesture representation. Most existing ultrasound-based gesture recognition tasks rely on a large amount of manually labeled samples to learn task-specific representations via end-to-end training.However, they cannot exploit unlabeled continuous gesture signals that are easy to collect. Inspired by recent self-supervised learning techniques, UltraCLR aims to autonomously learn a ubiquitous gesture signal representation that can benefit all tasks from low-cost unlabeled signals.We use the STFT heatmap as a secondary input and leverage the contrastive learning framework to improve the high-quality Channel Impulsive Response (CIR) heatmap input representations. The learned representations can better represent the spatial-position information and intermediate states of gesture movement. With the representation learned by UltraCLR, we can greatly reduce the complexity of downstream gesture recognition tasks so that they can be completed using a simple classifier trained with a small training set and a lower computational cost. Our experimental results show that UltraCLR outperforms state-of-the-art gesture recognition systems with only a few labeled samples, and achieves more than 85% reduction in computational complexity and over 9x improvement in inference speed.
This paper examines the application of WiFi signals for real-world monitoring of daily activities in home healthcare scenarios. While the state-of-the-art of WiFi-based activity recognition is promising in lab environments, challenges arise in real-world settings due to environmental, subject, and system configuration variables, affecting accuracy and adaptability. The research involves deploying systems in various settings and analyzing data shifts. It aims to guide realistic development of robust, context-aware WiFi sensing systems for elderly care. The findings suggest that a shift in WiFi data can come from various sources such as unseen environment and user, degrading the performance of WiFi-based activity sensing systems. While conventional domain shift techniques can partially mitigate data shift effects, further research is warranted to bridge the gap between academic research and practical applications.
Patients with Parkinson's disease (PD) often show gait impairments including shuffling gait, festination, and lack of arm and leg coordination. Quantitative gait analysis can provide valuable insights for PD diagnosis and monitoring. Prior work has utilized 3D motion capture, foot pressure sensors, IMUs, etc. to assess the severity of gait impairment in PD patients These sensors, despite their high precision, are often expensive and cumbersome to wear which makes them not the best option for long-term monitoring and naturalistic deployment settings. In this paper, we introduce mP-Gait, a millimeter-wave (mmWave) radar-based system designed to detect the gait features in PD patients and predict the severity of their gait impairment. Leveraging the high frequency and wide bandwidth of mmWave radar signals, mP-Gait is able to capture high-resolution reflected signals from different body parts during walking. We develop a pipeline to detect walking, extract gait features using signal analysis methods, and predict patients' UPDRS-III gait scores with a machine learning model. As gait features from PD patients with gait impairment are correctly and robustly extracted, mP-Gait is able to observe the fine-grained gait impairment severity fluctuation caused by medication response. To evaluate mP-Gait, we collected gait features from 144 participants (with UPDRS-III gait scores between 0 and 2) containing over 4000 gait cycles. Our results show that mP-Gait can achieve a mean absolute error of 0.379 points in predicting UPDRS-III gait scores.