Contemporary advances in the Neural 3 Dimensional (3D) Reconstruction (N3DR) have enabled new possibilities for detailed scene understanding and immersive environment modeling. However, the quality of reconstruction remains heavily sensitive to viewpoint selection and flight trajectory, particularly in low-altitude Unmanned Aerial Vehicle (UAV) applications, where energy, time, and occlusions pose significant challenges. This article presents an intelligent and adaptive path planning framework that dynamically optimizes UAVs’ viewpoints to maximize neural reconstruction quality while respecting real-world constraints. Unlike static or heuristic approaches, our system integrates online feedback from the reconstruction engine to guide UAVs toward underexplored or uncertain regions in real time. We review key principles of view planning, identify limitations of current strategies, and introduce a modular framework capable of fusing scene priors, uncertainty maps, and flight constraints into trajectory optimization. Finally, we describe an experimental evaluation methodology that leverages a real-world testbed deployment and synthetic benchmarks to validate the effectiveness of dynamic planning on both reconstruction accuracy and operational efficiency. This work aims to bridge aerial robotics, Internet of Things (IoT)-enabled vision systems, and Artificial Intelligence (AI)-driven planning for robust 3 Dimensional (3D) mapping in complex environments.
Tracking maneuvering targets requires estimators that are both responsive and robust. Interacting Multiple Model (IMM) filters are a standard tracking approach, but fusing models via Gaussian mixtures can lag during maneuvers. Recent winnertakes-all (WTA) approaches react quickly but may produce discontinuities. We propose SAFE-IMM, a lightweight IMM variant for tracking on mobile and resource-limited platforms with a safe covariance-aware gate that permits WTA only when the implied jump from the mixture to the winner is provably bounded. In simulations and on nuScenes front-radar data, SAFE-IMM achieves high accuracy at real-time rates, reducing ID switches while maintaining competitive performance. The method is simple to integrate, numerically stable, and clutter-robust, offering a practical balance between responsiveness and smoothness.
Dynamic obstacle avoidance (DOA) for uncrewed aerial vehicles (UAVs) requires fast reaction under limited onboard resources. We introduce the distributionally robust acceleration control barrier function (DR-ACBF) as an efficient collision avoidance method maintaining safety regions. The method constructs a second-order control barrier function as linear half-space constraints on commanded acceleration. Latency, actuator limits, and obstacle accelerations are handled through an effective clearance that considers dynamics and delay. Uncertainty is mitigated using Cantelli tightening with per-obstacle risk. A DR-conditional value at risk (DR-CVaR) early trigger expands margins near violations to improve DOA. To meet real-time avoidance-control at 100 Hz, we use fixed-time Gauss-Southwell projections instead of quadratic programs (QP). Simulation results show similar avoidance performance with 31% lower computational load than QP and outperform the state-of-the-art baseline approaches. Experiments with Crazyflie UAVs demonstrate the feasibility of our approach.
Fast dynamic obstacle avoidance (DOA) on uncrewed aerial vehicles (UAVs) demands not only low-latency control and actuation but also reliable perception with sufficient sensing range for accurate obstacle detection and speed estimation. This letter presents, to the best of our knowledge, the first mmWave RADAR-based perception-and-control system for fast onboard DOA. We derive and analyze latency and spatial bounds that relate sensing range, relative speed, and control delay, yielding sufficient conditions for successful avoidance. Our system adopts a lightweight tracker based on interacting multiple models and a controller based on control-barrier functions that directly outputs evasive accelerations. It achieves position errors of less than 0.15 m, 0.93 m, and 0.87 m in x, y, and z directions for 300 experiments with three different object sizes and varying visibility (light and dark), and a similar spread for 90 experiments in smoke. An onboard implementation on a Raspberry Pi 4B demonstrates real-time feasibility with an end-to-end sensing-to-command latency of approximately 14 ms. Code and the full dataset of 390 throws are available (https://tinyurl.com/radardoagit).
Vision is the most widely used perception modality in resource-constrained drones. However, camera-based perception faces limitations in dynamic, cluttered, and dark environments. This paper investigates the application of millimeter-wave (mmWave) Radio Detection and Ranging (RADAR) as an Object Detection (OD) and Object Classification (OC) perception modality for small-scale drones due to its resilience to environmental conditions. We propose a lightweight sliding-window method for static and dynamic OC using sparse RADAR point clouds captured by a cost-effective mmWave RADAR sensor. We compare the performance of RADAR- and vision-based detection performance with dedicated test data for three scenarios. The proposed method achieves robust OD, desired for aerial platforms, effectively handling objects with small RADAR Cross-Section (RCS) of 1-10 cm(2). Comparative analyses against camera-based YOLOv5 show an average recall of 0.8 for three different scenarios. Moreover, our system results in a mean absolute position error between 0.11m and 0.16m for RCS values of 10 and 1 cm2 compared to OptiTrack ground truth data. The results show that the proposed solution excels in all scenarios, enabling practical RADAR-based drone perception.
Miniaturized Uncrewed Aerial Vehicles (UAVs) can access indoor and hard-to-reach spaces, but severe constraints on payload and autonomy have limited their use in demanding tasks such as high-quality 3D reconstruction. We introduce a novel system architecture that enables autonomous, high-fidelity 3D scanning of static objects with sub-100 gram UAVs. Our core innovation lies in a closed-loop active viewpoint selection framework specifically tailored for ultra-constrained micro-platforms, advancing beyond standard static or offline active reconstruction methods. The framework establishes a dual-reconstruction pipeline that creates a real-time (RT) feedback loop between data capture and flight control. A near-RT process uses Structure-from-Motion (SfM) to generate an instantaneous point-cloud of the object. A systematic trajectory adaptation algorithm analyzes the model quality on the fly and dynamically adapts the UAV's trajectory based on parameterized spatial partitioning to intelligently capture new images of poorly covered areas, ensuring comprehensive acquisition. For the final, high-fidelity output, a non-RT pipeline employs a Neural Radiance Fields (NeRF)-based Neural 3D Reconstruction (N3DR) approach, fusing SfM-derived camera poses with precise external location data, evaluated across both radio-based Ultra Wideband (UWB) and visual motion-capture setups, to correct sensor noise and achieve superior accuracy. We implemented and validated this architecture using Crazyflie 2.1 UAVs. Our experiments, conducted in both single- and multi-UAV configurations show that algorithmic dynamic trajectory adaptation consistently improves reconstruction quality over static flight paths. This work demonstrates a scalable and autonomous solution that unlocks the potential of miniaturized UAVs for fine-grained 3D reconstruction, a capability previously reserved for much larger platforms.
Maintaining formation integrity is a significant challenge for a multi-agent system while navigating in cluttered environments. This paper presents a multi-agent formation control, integrating a virtual leader strategy with formation switching (FS) and Safe Artificial Potential Field (SAPF) control. Our goal is to navigate an Unmanned Aerial Vehicle (UAV) formation collision-free around obstacles allowing smooth transitions among three formations based on the available space. Simulation results show that our FS-SAPF framework achieves a higher success rate by effectively handling local minima and reducing oscillations than traditional Artificial Potential Field (APF) approaches. Our FS-SAPF maintains a larger minimum distance between agents and obstacles, thus enhancing safety in complex environments.
Smart agriculture is an enabling technology addressing the increasing challenges of efficiency, sustainability, and quality of food production. It requires rich data from the farming area at high spatial and temporal resolution. Although remote sensing systems have become readily available recently, in-situ sensing is still required to capture important properties of soil, crops, and their close environment. Ubiquitous sensor networks (USNs) provide a seamless and real-time in-situ sensing infrastructure that could overcome some limitations of smart agriculture. Resource efficiency is essential for USNs due to (1) the expected long operation time in typically resource-constrained environments, (2) the vast amount of captured and processed data, and (3) the ever-increasing application requirements. This survey comprehensively analyzes resource optimization techniques for USNs along three USN layers: the sensing, the communication & connectivity, and the processing & analysis layer. It discusses the application of these techniques in the smart agriculture domain and identifies current challenges and open research issues.
Multi-human parsing algorithms have significant potential for real-time surveillance applications. By accurately segmenting humans and their body parts, such algorithms help to understand and better differentiate multiple human subjects in video frames. However, deploying such algorithms on resource-constrained embedded devices such as smart cameras presents challenges due to memory constraints and limited computational power. Therefore, this work investigates the limitations of existing multi-human parsing algorithms and proposes MHParsNet, a lightweight yet accurate model for human parsing on embedded devices. Compared to benchmark algorithms, MHParsNet performs competitive segmentation while requiring only 125 MB of memory. We deployed MHParsNet on a smart camera prototype using the Jetson Nano embedded board and achieved an average inference of 6 frames per second. These results demonstrate the effectiveness of MHParsNet and its suitability for real-time applications on resource-constrained embedded devices.
Occupancy refers to the presence of people in rooms and buildings. It is an essential input for IoT applications, including controlling lighting, heating, access, and monitoring space limitation policies. Occupancy information can also be used to improve users’ comfort and to reduce energy waste in buildings. This paper evaluates the performance and resource consumption of recent machine learning techniques for occupancy detection and measurement by exploiting data from distributed environmental sensors. This evaluation is founded on a dataset captured by our dedicated sensor network for indoor monitoring, comprising temperature, humidity, and carbon dioxide (CO 2 ) sensors. Using different sensor modalities and spatio-temporal data selections, we compare eight classification algorithms based on the accuracy achieved and the required runtimes. Binary classification for occupancy detection (OD) achieves accuracies over 90% for individual modalities and close to 100% for modality combinations. Multi-class classification for occupancy measurements (OM) shows as clear ranking of the sensor modalities, and gradient boosting algorithms are superior when combining sensor modalities and fusing data from multiple sensors.
Self-awareness is intelligent agents' capability to become aware of new experiences using their sensory data. This paper focuses on anomaly detection following principles of self-awareness and proposes coupled hierarchical dynamic Bayesian networks (DBN) as causal–temporal models to learn cooperative multi-robot behaviors from sensory data. These trained models allow anomaly detection whenever an observed behavior deviates from the learned behavior. We evaluated our approach with a two-drone leader–follower setup where the drones with GPS and LIDAR sensors conduct different maneuvers. Our simulation study shows that coupled DBNs trained with independent sensory data achieve better anomaly detection than DBNs trained with aggregated sensory data.
Accurate location information is essential for autonomous robots since inaccurate localization can impede many robotic tasks or even lead to collisions. We investigate Fisher information theory as tool for assessing the quality of location information and making robots self-aware about their own location uncertainty. We further propose a navigation framework that exploits this location uncertainty for the motion control of the robot and aims for improving the mission performance while reducing location uncertainty. The framework enhances an artificial potential field (APF) controller by an adaptive information-seeking (IS) force towards areas with low spatial uncertainty.
This work addresses the challenge of text-based person re-identification (re-ID) on resource-constrained embedded devices, a critical component in modern surveillance systems. Text-based person re-ID involves using textual descriptions to search persons across multiple camera views. Implementing such algorithms on embedded devices such as smart cameras is challenging due to limited memory and computational constraints. In this work, we propose TextReIDNet, a lightweight person re-ID model designed explicitly for embedded devices. Compared to state-of-the-art models, TextReIDNet aims at an optimal balance between person re-ID accuracy and computational efficiency, thus making it well-suited for low-resource devices. With the smallest model size of only 32.29 million parameters, TextReIDNet achieves a competitive 52.76% and 35.71% top-1 accuracy on the CUHK-PEDES and RSTPReid datasets, respectively. We implemented TextReIDNet on the Jetson Nano board to demonstrate its capability for embedded deployments. On average, TextReIDNet requires 1.13ms to process a text and 30.92ms for an image.
Text-based person re-identification (re-ID) is an emerging research domain in multi-camera surveillance systems. It involves identifying individuals across different camera views using textual descriptions as queries. Although text-based person re-ID holds significant potential for surveillance systems, its practical applications are limited by the high computational demands of existing algorithms. This is because most state-of-the-art algorithms prioritize identification accuracy over resource efficiency. Therefore, surveillance systems based on existing algorithms rely heavily on centralized architectures, where a central application server aggregates videos from camera nodes, generates galleries from the videos, and performs person identification. However, such centralized systems face challenges in large-scale deployments, including scalability issues, high bandwidth utilization, and processing bottlenecks. This paper presents a decentralized approach to text-based person re-ID to overcome these limitations. First, it introduces U-TextReIDNet, a resource-efficient model designed to identify persons in multi-person images on the NVIDIA Jetson Nano embedded board. U-TextReIDNet achieves Top-1 accuracies of 54.02% on the CUHK-PEDES dataset and 38.45% on the RSTPReid dataset. With only 38.77 million parameters, U-TextReIDNet is significantly smaller than most existing text-based person re-ID models. Using U-TextReIDNet, we implement a decentralized system that distributes the person identification task across camera nodes, transmitting only videos containing persons of interest to the command station. Additionally, we developed a prototype of this decentralized system, and conducted performance and usability tests using real human subjects. The prototype successfully performs real-time person re-ID, reduces bandwidth utilization, improves system scalability, and eliminates processing bottlenecks.
Tools for specifying and executing multidrone missions that go beyond pure orchestration of waypoints are rare. We present the EAMOS framework, which introduces a simple and intuitive text-based mission specification process to execute a multidrone mission onboard different heterogeneous drones. Key benefits of EAMOS are the easy handling of sequential and parallel drone actions and their automatic synchronization. A uniform drone-interface abstracts the handling of different drone types, and specialized mission control structures enable specifying high-level missions. Our EAMOS prototype has been completely implemented in Go and successfully demonstrated in combination with the Airsim multidrone simulation environment and the PX4 flight controller as a software-in-the-loop component. Synchronization among multiple drones wrt. their sequentially and concurrently performed actions as well as the correct application of mission control structures behave as expected.
The increasing use of low-cost sensors in monitoring the surrounding environment requires efficient handling of sensor drift and sensor errors. Therefore, there is a pressing need to develop lightweight methods to determine and calibrate the sensor’s readings accurately. This article focuses on the calibration of low-cost sensors using lightweight techniques to effectively detect and correct sensor drifts. The proposed approach combines clustering, which offloads computational burdens to cluster heads, with temporal and spatial estimation among neighboring sensors, such as inverse distance weighting (IDW). Additionally, autoregression (AR) and interquartile range (IQR) techniques are employed to monitor the stability of sensor readings based on the previous measurements. Through simulation experiments using a dataset from the Intel Berkeley Research Laboratory (IBRL), the effectiveness of the proposed method in reliably detecting and improving the sensor is demonstrated. These findings contribute to advancing sensor calibration techniques for low-cost sensors, enhancing their reliability and accuracy in environmental monitoring applications.
Existing software tools for specifying and executing multidrone missions are limited to route planning or tightly coupled to specific drone hardware. We introduce EAMOS (Execution of Aerial Multidrone Missions and Operations Specification Framework), which allows us to specify missions intuitively, text-based, and provides a mission compiler, a mission middle layer, and a distributed drone execution environment. The middle layer wraps the control of individual drone-specific capabilities, such as launch, fly to position, or perform a maneuver, into a public API that transparently utilizes the capabilities of numerous drone platforms. We exploit the Go programming language to implement critical components of the framework and provide an interface for ROS-based drone platforms. EAMOS automates the mission execution on real, virtual, and even hybrid robotic setups involving real and virtual drones. We demonstrate the successful deployment of EAMOS with four missions executed on Pixhawk/PX4-equipped quadcopters and virtual drones simulated with Airsim. We assess the performance of our proposed approach by analyzing the number of nodes and arcs of the mission graphs, which are an essential artifact of our mission compilation, the utilization of ROS service calls during mission execution, and the duration of compilation, deployment, and mission execution. Overall, our experiments showed that our drones correctly behaved during mission execution as expected and specified by their mission, the generated mission artifacts were efficiently manageable, and processing times allowed for a fluent workflow.
Planning the collision-free, simultaneous movement of multiple agents is a fundamental problem in multirobot systems and becomes particularly challenging in highly confined environments. This paper addresses this problem by first searching for collision-free paths in a discretized environment and then optimizing the agents' dynamically feasible trajectories along the discrete paths with respect to energy and time. Our approach extends the available movement options at each planning step of the enhanced conflict-based search and results in shorter path lengths, faster planning times, and reduced number of waiting events for agents at waypoints. We compare our new approach with the original algorithm in a simulation study and investigate the performance in confined spaces.
Real-time and accurate position estimation is critical for various multi-robot applications and serves as a prerequisite for location-based multi-sensor data analysis. However, it is often impeded by energy, sensing, and processing limitations. In this work, we study the problem of information-seeking in localization and navigation in multi-agent systems, which aims to navigate mobile agents while reducing position errors. We formalize information-seeking as reducing spatial uncertainty and introduce an efficient motion controller based on artificial potential fields superimposing attractive, repulsive, and information-seeking forces. We evaluate the effect of information-seeking on localization and mission planning in a simulation study with non-collaborative and collaborative localization approaches.
We consider a multi-robot patrolling scenario with intermittent connectivity constraints, ensuring that robots' data finally arrive at a base station. In particular, each robot traverses a closed tour periodically and meets with the robots on neighboring tours to exchange data. We model the problem as a variant of the min-max vertex cycle cover problem (MMCCP), which is the problem of covering all vertices with a given number of disjoint tours such that the longest tour length is minimal. In this work, we introduce the minimum idleness connectivity-constrained multi-robot patrolling problem, show that it is NP-hard, and model it as a mixed-integer linear program (MILP). The computational complexity of exactly solving this problem restrains practical applications, and therefore we develop approximate algorithms taking a solution for MMCCP as input. Our simulation experiments on 10 vertices and up to 3 robots compare the results of different solution approaches (including solving the MILP formulation) and show that our greedy algorithm can obtain an objective value close to the one of the MILP formulations but requires much less computation time. Experiments on instances with up to 100 vertices and 10 robots indicate that the greedy approximation algorithm tries to keep the length of the longest tour small by extending smaller tours for data exchange.