
Drones serve as the primary camera-mounted platform to collect high-resolution images of buildings from various viewpoints for 3D neural radiance fields (NeRF) reconstruction. However, the flexibility of drones in lowering flight heights to supplement richer building details and thus improve the reconstruction quality of NeRF has not been explored. Given the limited power of a drone and the rapidly changing nature of outdoor scenes, it is necessary to quickly predict an optimal supplementary lower flight height after the drone finishes capturing at a default height. The main challenges involve the implicit relationship between supplementary flight heights and reconstruction quality, the complexity of prediction task with little prior information, and the tension between the need for real-time decision making and the resource constraints inherent to on-drone task execution. In this work, we develop an end-to-end system pipeline with an offline-online decoupling feature. We first design a model architecture that embeds supplementary flight heights to predict the reconstruction quality of NeRF using images captured at the default height. To enhance model generalization ability in a cost-effective manner, we offline train an ensemble of multiple models over the samples constructed from sandbox buildings using building-level cross validation. During the online serving phase on the drone, the most appropriate model is selected based on feature matching between real and sandbox buildings to determine the optimal supplementary flight height. We build a testbed and evaluate on 16 sandbox buildings and 8 real-world buildings. Evaluation results demonstrate that our pipeline accurately identifies the optimal supplementary lower height in a few minutes and improves reconstruction quality.
Machine learning has become a prominent technique used in wireless sensor networks (WSNs) in the past 10-15 years. Because of their high accuracy, adaptability, and potential to lesson computational overhead compared to existing mathematical algorithms, machine learning techniques are excellent for dynamic environments. Within machine learning, there are many different techniques to consider, such as supervised learning, unsupervised learning, reinforcement learning, and others. This paper will focus on Q-learning, and its advantages in WSN selection algorithms. We propose a Q-learning-based sensor selection framework to optimize node scheduling while providing k-coverage by learning policies that balance coverage and remaining energy. In particular, sleep scheduling algorithms are investigated and the Q-learning developed policy is compared with an existing mathematical algorithm via simulations. To conclude, the advantages and disadvantages of machine learning with resource constrained WSNs is discussed.
The fusion of multi-modal information, such as images and LiDAR scans, is instrumental to maximize the performance of many computer vision tasks in next generation systems and applications. However, supporting fusion necessitates considerable effort, challenging the availability of computing and communication resources in edge systems. This work addresses this challenge by maximizing resource efficiency in systems where mobile devices collect multi-modal sensor data and use dynamic multi-branched DNN models to adapt inference to the operating context. To tune the overall system response to the context (e.g., weather conditions), we propose a dual-scale control approach: centralized orchestration of spectrum resources, and distributed individual device-level control of the execution path of the dynamic DNN fusion models. The control agents are driven by a novel context-aware decision-making method combined with game theory, named Context-Aware Network Slicing Auction (CANSA), which optimizes DNN inference performance, network slicing, and energy consumption. The decision-making performs such optimization by: (i) selecting data and features that best fit the current context; (ii) deciding on the appropriate DNN model complexity, including the use of multi-modal sensor fusion techniques for better data integration, and (iii) deploying these models on the most appropriate nodes (local nodes or edge servers). Results, obtained using real-world multi-modal data, show that CANSA surpasses conventional allocation methods by up to 52.3% in terms of inference task success rate.
Door Access Control (DAC) plays a pivotal role in balancing security and convenience in modern infrastructures. However, current vision-based DAC systems exhibit limitations including privacy concerns (e.g., facial data leakage), performance degradation under suboptimal lighting conditions, high computational overhead and system costs. While RFID and Bluetooth-based alternatives exist, they exhibit vulnerabilities to attacks including replay attacks, signal cloning, and eavesdropping. Recent advances in visible light sensing and backscatter communication have enabled promising opportunities for secure, low-power access control systems with sub-dollar hardware costs. In this paper, we propose ViKey, the first visible light backscatter-based DAC system that utilizes polarized birefringence to generate 3D position-dependent color patterns as keys, enabling robust and contactless authentication. We design and implement a ViKey prototype using commercial off-the-shelf (COTS) components, with a tag cost of less than $0.2. Real-world experiments show that our current ViKey prototype can achieve an average authentication accuracy of 90.5% at 0.5m with our best patterns. These results demonstrate the effectiveness of our low-cost visible light backscatter technology for future smart DAC applications.
In WiFi and many other sensing applications, researchers have adopted the approach of representing each activity class with a small portion of data to address user variability across different environments, thereby enabling model generalization. This solution raises yet another challenge - limited training samples for each activity class, particularly for a considerable number of classes to be classified. In this paper, we first introduce a visionary architectural layering architecture that enables the end-to-end lifecycle of multi-sensor federated learning, paving the way for realizing federated AI-based sensing. We identify three challenges, e.g., data heterogeneity, scalability under non-IID conditions, and few-shot learning. According to the architecture, we showcase a lightweight, general workflow that utilizes a federated learning framework to address the few-shot learning challenge across various classifications in WiFi Channel State Information (CSI). We analyze this workflow and identify two key factors that most significantly affect its performance (i.e., the aggregation method and the model architecture). Based on the analysis, we propose three network models for CSI-based activity recognition: CSI-AlexNet, CSI-ActNet, and CSI-ResNet. Through extensive performance evaluations, our experimental results show that the proposed approach achieves an accuracy of 97.97% on CSI-ActNet while maintaining low computational demands, making it well-suited for real-time fine-tuning on edge devices with constrained resources.
Self-driving platooning trucks are becoming increasingly common due to the economic gains for companies that utilize them. However, high throughput, long range communication between truck platoons can be very difficult if they are not in areas with pre-built cellular infrastructures. To combat this problem, we propose PlaCoB (Platoon Collaborative Beamforming), which enables multiple platooning trucks to collaboratively beamform to maintain communication over long distances. We conduct extensive simulations under various settings to evaluate PlaCoB’s performance compared to single-truck methods. By leveraging multi-truck platoons, collaborative beamforming, and frequency shifting, (1) PlaCoB can transmit 1.23x more data compared to single vehicle transmission methods, and (2) PlaCoB can transmit data 2.44x further than single vehicle transmission methods. These results demonstrate the effectiveness of our proposed PlaCoB’s approach.
Object distance estimation is a critical task in autonomous driving and has received increasing attention in recent years. Existing methods rely on either geometric cues, which typically reduce structures to coarse attributes like 2D box dimensions, or visual cues that degrade under occlusion and truncation. In this work, we propose a geometry-enhanced framework that augments visual features with a more stable and descriptive geometric attribute—Patch-To-Center Geometry (PCG)—which encodes the spatial distance between local object patches and the 3D object center projection of the object to improve distance estimation. Specifically, each object is divided into multiple patches, and a dedicated attention module is employed to learn the pixel-wise spatial offsets from each patch to the projected center, serving as an auxiliary supervision signal. Additionally, a head token is introduced to aggregate global information from all patches for final distance prediction. Extensive experiments on the KITTI and nuScenes datasets demonstrate that our method outperforms existing approaches when limiting the minimum depth threshold.
This paper investigates the implementation of artificial noise (AN) in a combinatorial design based multiple-input single-output non-orthogonal multiple access (MISO-NOMA) communication system to enhance physical layer security. To address the risk of eavesdropping and enhance the physical layer security, we propose an artificial noise-aided transmission scheme designed to protect the confidential information of legitimate NOMA users by intentionally degrading the channel quality of potential eavesdroppers. Our proposed system leverages a combinatorial design-based approach to configure the MISO-NOMA framework, ensuring effective user pairing and beamforming strategies. The artificial noise is carefully injected into the null space of the legitimate users’ channels to minimize interference to them while significantly deteriorating the eavesdropper’s reception capability. Maximum likelihood detection is applied to detect the transmitted message of each user. We provide a comprehensive performance evaluation through numerical simulations, demonstrating the effectiveness of the proposed method in enhancing the secrecy rate and overall security of the MISO-NOMA system under various network conditions and channel scenarios.
The Internet of Underwater Things (IoUT) requires a fast and trustworthy communication framework in order to ensure the timely delivery of mission-critical data in dynamic environments. Named Data Networking (NDN), a data-centric architecture, offers inherent data authentication and in-network caching. However, secure data exchange across organizational or administrative domains becomes a challenge due to isolated trust schemas. This paper proposes and evaluates an Inter-Zone Trust framework within NDN that allows for controlled and authenticated communication between multiple trust zones in an IoUT environment. A custom simulation is implemented using ndnSIM, custom producer, consumer, and controller apps, and dynamic trust revocation. Three simulation scenarios, namely, baseline, inter-zone trust, and malicious attack of a trust revocation is used to quantify potential performance overhead. The results of application delay, packet count, and energy consumption show that despite the slight increase in overhead, the proposed inter-zone trust architecture successfully responds to cache poisoning attacks and provides a scalable foundation for secure inter-zone communication for multi-organizational IoUT deployments.
Panoptic perception models in autonomous driving use deep learning models to interpret their surroundings and make real-time decisions. However, these models are susceptible, carefully designed noise can fool models all while being imperceptible to humans. In this work, we investigate the impact of black-box adversarial noise attacks on three core perception tasks: drivable area recognition, lane line segmentation, and object detection. Unlike white-box attacks, black-box attacks assume no knowledge of the model’s internal parameters making them a more realistic and challenging threat scenario. Our goal is to evaluate how such an attack affects the model’s predictions and explore countermeasures towards such attacks. In response to our implemented attack, we have tested various defense methods. With each defense method, we have assessed the recovery on prediction accuracy. This research aims to provide valuable insights into the vulnerabilities of panoptic perception models and highlights strategies for enhancing their resilience against adversarial manipulation within real-world scenarios. All our attacks are performed against images from the BDD100K dataset.
LoRaWANs, a widely accepted IoT connectivity solution, adopt a simple (ALOHA-like) MAC layer, enabling low-power communication at the cost of scalability due to packet collisions. Hence, current studies on LoRaWAN conclude that the network does not support dense deployments. Several alternative MACs are proposed but they stumble upon well-known limitations: time division eliminates the asynchrony of LoRa nodes but requires feedback from the gateways; carrier-sensing-based protocols are heavily constrained by the reduced sensing ranges of the devices, thus creating a large number of hidden terminals, leading to collisions.To enhance LoRaWAN to cater to both low- and high-density deployments, in this paper, we propose Spreading Factor MAC (SFMAC), a novel, practical, distributed, and energy-efficient MAC protocol. SFMAC, a channel-sensing-based MAC, takes an unconventional approach to eliminate hidden terminals – by operating with pairs of SFs, wherein the higher SF is used for channel sensing and the lower for data transmission. Bleeps are transmitted in the higher SF as they can be sensed at longer ranges. SFMAC does not require any change in hardware or the LoRaWAN protocol. We demonstrate that the fundamental trade-off made by SFMAC – utilizing two SFs per data transmission instead of using all for data – works extremely well due to the elimination of hidden terminals. Through real-world experiments on 30 SX1261 devices and data-driven ns-3 simulations, we showcase that SFMAC increases goodput and channel utilization by manifolds over state-of-the-art protocols such as p-CARMA, np-CECADA, and LMAC.
Robust object detection in adverse weather conditions is critical for ensuring the safety and reliability of autonomous driving systems. In this work, we present a detailed study on the adversarial robustness of YOLO-based detectors using the RealDriveSim dataset, which includes foggy, rainy, and nighttime scenarios. We benchmark YOLOv9 and YOLOv10 under clean conditions and observe high performance, with YOLOv10 achieving a mean average precision (mAP) of 69.6%. To evaluate vulnerability, we introduce an adversarial patch optimized to suppress road object detections. After patch-based perturbation, mAP drops to 44.3%, highlighting the importance of a defense system. To counter this degradation, we propose a lightweight LiDAR-camera fusion framework that does not require model retraining or architectural changes. Our method projects 3D LiDAR point clouds into the 2D image plane using intrinsic and extrinsic calibration parameters and cross-validates each 2D detection by checking for supporting 3D LiDAR points within its bounding box with an inference time of only 7.2 ms. Our fusion strategy effectively filters adversarial false positives, leading to a recovery in mAP to 62.9%, without requiring model retraining or architectural changes. To the best of our knowledge, this is the first work to benchmark adversarial robustness and sensor-level fusion defense on the RealDriveSim dataset, setting a new standard for evaluating real-world physical attack resilience in autonomous perception.
The classification of ransomware remains a critical yet challenging task in cybersecurity. Motivated by the increasing sophistication and overlap in behaviors between ransomware and general malware, this work addresses the need for more precise differentiation methods to facilitate targeted mitigation efforts. Our study proposes an innovative approach for ransomware classification using sub-graph mining of Function Call Graphs (FCGs). We employ Cuckoo Sandbox™ to extract dynamic API calls and construct detailed FCGs. Through focused subgraph mining, we isolate critical API call patterns specifically relevant to ransomware behavior. These extracted patterns are then vectorized and classified using a Convolutional Neural Network (RansomNet-CNN), achieving high precision in distinguishing ransomware from general malware. Unlike full-graph or flat-sequence models, our subgraph-level approach precisely captures ransomware-relevant behaviors. The RansomNet-CNN model demonstrates superior performance, achieving a precision of 99% and a recall of 100%, thus underscoring its practical effectiveness in ransomware identification. The dataset and code are publicly available at our Zenodo Repository1.
Intrusion detection systems (IDS) primarily rely on signature-based approaches, which can fail to detect novel or sophisticated attacks. This paper addresses the underutilized potential of leveraging a multi-modal approach that combines packet capture (PCAP) and log data for anomaly detection. To enhance detection capabilities, we propose an interpretable hybrid neural network architecture, TransIDS, that integrates a packet-based transformer with an efficient transformer-based language model for log messages. The proposed framework extracts semantic vectors from raw log messages and concatenates them with packet embeddings. An attention-based classification model then detects anomalies by determining the importance of each log message and packet for the neural network’s decision. By fusing spatial features from PCAP data with temporal features from log data, TransIDS utilizes this multi-modal data fusion to identify anomalies that might be missed by conventional systems. This approach not only leverages the strengths of two distinct transformer-based architectures but also provides a more comprehensive analysis of network traffic, leading to more effective detection of previously undetected attacks and strengthening overall network security. We use a real testbed for our experiments to validate the effectiveness of our proposed approach.
Recent studies have demonstrated significant success in detecting attacks on the Controller Area Network (CAN) bus network using machine learning and deep learning models, including convolutional neural networks and transformer-based architectures. Building on this foundation, our work investigates the use of large language models (LLMs) not only for intrusion detection but also for providing interpretable explanations of their decisions. We fine-tuned three LLMs, i.e., SecureBERT, LLaMA-2, and LLaMA-3, for intrusion detection on CAN bus data. Among them, LLaMA-3 delivered the best results, achieving SOTA performance on the Car-Hacking dataset. Beyond attack classification, we evaluated LLaMA-3’s ability to generate reasoning for its decisions through zero-shot prompting. The model successfully articulated its rationale, particularly for Denial-of-Service (DoS) attacks, demonstrating strong potential for explain-ability in intrusion detection systems. These findings highlight the potential of LLMs to serve as a highly accurate intrusion detection system while simultaneously providing interpretable explanations, thereby enhancing the investigative capabilities of cybersecurity professionals.
Door Access Control (DAC) plays a pivotal role in balancing security and convenience in modern infrastructures. However, current vision-based DAC systems face limitations such as privacy risks, degraded performance, and high computational costs, while RFID and Bluetooth alternatives remain vulnerable to replay, cloning, and eavesdropping attacks. Recent advances in visible light sensing and backscatter communication have enabled promising opportunities for secure, low-power, and cost-effective access control systems. We propose ViKey, the first visible light backscatter-based DAC system that utilizes polarized birefringence to generate 3D position-dependent color patterns as keys, enabling robust and contactless authentication. The design methodology of ViKey highlights the potential of visible light backscatter technology for future smart DAC applications.
Depth estimation is a critical component of many computer vision applications, enabling accurate spatial awareness from visual inputs. It is particularly vital in autonomous driving and robotics, where precise depth information supports essential functions such as navigation and obstacle detection. Traditional methods, including sensor-based and monocular vision techniques, often face limitations such as high costs or reduced accuracy. In contrast, binocular depth estimation, which infers depth by analyzing disparities between stereo image pairs, offers a compelling balance of cost-effectiveness and precision. However, many existing models suffer from high computational demands and long inference times, limiting their suitability for real-time deployment. To address these challenges, we propose a streamlined convolutional neural network optimized for binocular depth estimation. Our model significantly reduces architectural complexity while maintaining strong performance. Additionally, we introduce a novel loss function that incorporates adaptive weighting and consistency constraints to enhance accuracy and stability during training. We evaluate our method on the KITTI 2015 benchmark and demonstrate that it achieves competitive accuracy while significantly reducing runtime compared to existing approaches. These improvements make our method a practical and efficient solution for real-time depth estimation in resource-constrained environments.
Efficient post-outage recovery in smart grids is critical for minimizing service disruption, reducing recovery costs, and maintaining system stability. This paper formulates the backup power scheduling problem during recovery process as a Markov Decision Process (MDP), and employs cooperative multi-agent Deep Q-Networks (DQN) to jointly coordinate the activation of backup power units and the prioritization of damaged power nodes for repair. Unlike traditional heuristic or static methods, the proposed method adapts in real time to evolving grid conditions and captures complex interdependencies between the power and communication networks. Simulation results on different test systems demonstrate that the proposed method significantly reduces recovery time, conserves limited backup power, and enhances grid resilience compared to baseline strategies.
LoRa is widely viewed as a promising wireless technology to support connections of IoT devices to the gateway. The downlink of LoRa carries traffic such as acknowledgments and faces challenges as the network size grows. In this paper, we propose 2-Pipe, a novel method that significantly enhances the downlink capacity of LoRa without modifying any nodes. With 2-Pipe, the gateway can transmit two downlink packets simultaneously with the same Spreading Factor (SF) to two distinct nodes. Simultaneous transmissions are achieved by modulating packets with intentional misalignment both in the time and frequency domains, so that a node can demodulate its own packet correctly without being affected by the other because the intentional misalignment reduces the impact of interference. Experiments with commodity LoRa devices show that the throughput of 2-Pipe is 1.62× that of the existing LoRa.
Professional networking at in-person events plays a crucial role in career growth. However, many professionals struggle to maintain and transition these connections to online platforms like LinkedIn. Missed opportunities often arise due to challenges in tracking interactions, forgetting to add contacts, and the absence of tools that facilitate seamless networking. This paper examines how we can track those missed opportunities to connect with people from events by developing a protocol to find and record proximity interactions using Bluetooth while maintaining security and privacy of individual and device data, and then implementing that protocol in a proximity networking application developed for mobile phones. By bridging the gap between offline encounters and online professional networks, we believe that our approach enhances networking efficiency, ensuring that valuable connections made at events are sustained beyond the physical space.