
The primary objective of this paper is to propose a multi-objective optimization process of an IoT network for cattle monitoring in open environments. The proposed approach utilizes a LoRa network simulation model and optimization algorithms focused on improving the choice of its configuration parameters. It simulates cattle's real displacement and compares the optimization performance of three algorithms, indicating the best of these. Results demonstrate that the approach used in the work successfully optimizes the parameters, reducing energy consumption and increasing communication between gateways in this application scenario.
Field-Programmable Gate Arrays (FPGAs) have emerged as a promising platform for accelerating AI applications in IoT devices. In this context, Spiking Neural Networks (SNNs) have demonstrated superior energy efficiency compared to traditional artificial neural networks when implemented on FPGAs. Although various frameworks have been proposed for mapping SNNs onto FPGAs, none of them have leveraged High-Level Synthesis (HLS), which facilitates faster development cycles and greater design transparency. This work introduces NeuroHLS, an HLS-based framework for implementing SNNs on FPGAs. NeuroHLS offers both resource- and latency-optimized implementation strategies, and supports fine-grained parameter customization at the layer level. To validate the framework, a 784128-10 SNN was implemented and evaluated on the N-MNIST dataset. Performance was benchmarked against existing state-of-the-art frameworks. Experimental results demonstrate that both implementation strategies offered by NeuroHLS significantly outperform existing approaches in terms of inference latency and energy efficiency. In particular, one configuration achieved an inference latency of 0.8us and an average energy consumption of 2.16uJ per inference.
The deployment of Deep Neural Networks (DNNs) in safety-critical systems necessitates resilience against hardware threats like the Bit-Flip Attack (BFA). This challenge becomes particularly nuanced for quantized networks. Unlike their floating-point counterparts where single flips can be catastrophic, the inherent robustness of quantized models—the exclusive focus of this work—makes identifying critical bits computationally prohibitive. We introduce the Machine Learning-Guided Framework for Efficient BFAs (MLBFA), a novel framework that leverages Machine Learning (ML) to execute highly efficient Bit-Flip Attacks (BFAs). MLBFA operates in three stages: (1) it builds a Vulnerability Model (VM) from a limited statistical fault injection campaign using a multidimensional feature set; (2) it uses the VM to predict and rank the vulnerability of all parameters across the network; and (3) it exploits this ranking to guide advanced search algorithms (Greedy, Evolutionary, and Progressive) in identifying the minimal set of bit-flips required to neutralize the DNN. On quantized ResNet architectures, crucial for embedded systems, MLBFA consistently outperforms state-of-the-art methods, neutralizing a ResNet-32 with only six bit-flips and reducing the required flips by nearly 50% on deeper models like the ResNet-56. This work demonstrates that a holistic, ML-guided approach provides a superior strategic map for executing more potent hardware-level attacks with fewer resources.
Deploying deep learning models in embedded agricultural systems requires balancing predictive performance with strict hardware constraints such as memory, power, and latency. This work uses quantization techniques, specifically post-training quantization (PTQ) and quantization-aware training (QAT), to optimize convolutional neural networks for real-time pest detection in smart traps. Using the Brevitas framework for quantization and FINN for hardware generation targeting FPGAs, we evaluate a range of weight and activation bit-width configurations. Our results show that QAT significantly outperforms PTQ, particularly in aggressive low-bit scenarios, achieving high accuracy while drastically reducing hardware resource utilization. Among the proposed solutions, one achieves 87.47% accuracy while using less than 10% of the LUTs required by its full-precision counterpart. Comparative analysis with standard models such as ResNet18 and MobileNet further validates the effectiveness of our approach. This study highlights the practicality of QAT-driven quantization for edge Artificial Intelligence applications in agriculture. It paves the way for future work, including power, latency, and throughput profiling, to support large-scale deployment.
Semantic segmentation of 3D point clouds plays a fundamental role in applications such as autonomous vehicles, robotics, and 3D mapping, enabling the precise identification of objects and relevant classes. However, performing this task in real time presents challenges related to computational efficiency and model optimization. This work presents the development and performance analysis of a pipeline for real-time Light Detection and Ranging (LiDAR ) data acquisition and inference, integrating machine learning models based on Convolutional Neural Networks (CNNs). The proposed approach combines Software-in-the-Loop techniques and the concept of a digital twin, covering the entire process from data capture and preprocessing to semantic segmentation, ensuring a continuous processing workflow. The experiments done with an adaptation of the RIU-Net architecture demonstrated the feasibility of the proposed solution, showing satisfactory results in terms of execution time, accuracy, and Intersection over Union (IoU), reinforcing its applicability to real-time 3D computer vision systems.
With the increasing adoption of advanced drone technologies across diverse fields, the Internet of Drones (IoD) has emerged as a novel mobility paradigm, particularly enhancing Intelligent Transportation Systems (ITS) in urban environments. Despite its significant potential, IoD faces substantial challenges due to inherent resource constraints such as limited computational power and energy capacity, which hinder the implementation of robust cybersecurity solutions. These limitations expose IoD networks to various security vulnerabilities and privacy threats, necessitating an exhaustive analysis and understanding of these risks.This survey presents a comprehensive examination of security issues within the IoD domain. We review potential threats, and propose an threat modeling approach. Our key contributions include: (i) a unified taxonomy of attacks (ii) an extended and refined IoD definition that encompasses drone-based network considerations This work aims to foster further academic inquiry and practical advancements, accelerating the secure and privacy-conscious development and deployment of IoD systems.
Safety-Critical V2X maneuvers—platooning, coordinated braking, distributed sensing—collapse when wireless delay jitter flares or one talkative node monopolizes the communication channel. Commodity Wi-Fi provides neither determinism nor per-station fairness. We address both issues with two lightweight Time-Aware Shaper (TAS) schedules that carve each cycle into eight fixed, non-overlapping slots, reserving gaps for best-effort traffic. The design relies only on coarse IEEE 802.1AS synchronization—eschewing full 802.1Qbv gate lists—so it fits embedded on-board units. We implement the shapers in OMNET++ 6.1/INET 4.5.2 and evaluate eight vehicles at motorway speed while deliberately stressing their clocks.The tighter-cycle variant (TAS2, Tcycle = 48 ms, Tslot = 6 ms) slashes the end-to-end delay coefficient-of-variation from 177% (EDCA) to 63%, holds mean delay at 22 ms, and enforces the desired 0.125 RPHY bandwidth cap per vehicle with only a 3.1% aggregate-throughput penalty versus an unconstrained baseline. These results show that simple static slotting—not complex per-packet scheduling—can deliver deterministic latency and predictable per-car capacity on off-the-shelf Wi-Fi, making wireless TSN a viable foundation for next-generation cooperative vehicular systems.
Unmanned aerial vehicles (UAVs) have been widely used in various Internet of Things (IoT) applications. However, given the stringent limitations of UAVs, strategies to maximize computational efficiency considering these devices’ processing and energy constraints have recently attracted significant interest. This work presents BatFed, a battery-aware strategy for federated machine learning based on transfer learning for remote monitoring using UAVs. The MobileNetV3Small architecture was employed as the foundation for developing a model capable of maximizing computational efficiency on edge devices with processing and energy constraints, while the EuroSAT_RGB dataset served as the evaluation basis. Our strategy employs dynamic client selection to reduce the communication rate between devices and the central server, promoting greater bandwidth efficiency. Furthermore, the experiments, made in a non-IID scenario using Dirichlet distribution to simulate a moderately heterogeneous scenario (α = 1.0) achieved an accuracy exceeding 80% in most cases, reaching over 90% accuracy with additional local training epochs. Compared to our centralized learning baseline, which achieved 95% accuracy, this is a positive result with other significant advantages in terms of time savings, energy consumption and communication savings, demonstrating its feasibility for distributed monitoring applications.
This paper presents a computational simulation focused on wildfire suppression through the use of drone swarms. The aim was to simulate the operational efficiency of the swarm in extinguishing fires in scenarios influenced by environmental factors such as wind and topography. Hansen’s method was employed to estimate the number of drones required. The A* algorithm was used for navigation, and a cell matrix-based model was implemented for fire propagation. A series of 30 tests were conducted on a device, where initial fire hotspots were generated at random points in an area of 10,000 m2, with drone deployment occurring 30 seconds after the fire started. Wind was simulated with random directions and speeds ranging from 10 to 30 km/h, mimicking different environmental conditions. Each test saved data on fire duration, burned area, number of drones, and wind variables. The results indicated that the use of drones reduced the fire duration by an average of 33% compared to fires without drones, as well as decreasing the burned area by approximately 2,167 m2. The drone simulation showed a significant reduction in fire suppression time, with duration ranging from 38 to 43 minutes, compared to 57 to 65 minutes without drones, and the burned area ranged from 7,742 m2 to 8,101 m2, compared to 9,442 m2 to 10,000 m2 without drones.
Integration between different layers in an IoT solution is challenging due to the heterogeneity of hardware and software components, communication infrastructure and protocols. In this paper, we explore the use of Function-as-a-Service to integrate and distribute execution across different layers in an architecture consisting of Edge and Fog devices. We also explore a circular economy solution for e-waste by reusing confiscated TVBoxes as nodes in a Computing Continuum model. The architecture uses a TVBox as a sensing layer, and we evaluate different solutions such as TVBoxes and Raspberry Pis as Edge devices. A x86 server is used as a Fog layer. To evaluate these different solutions, we implemented two case studies involving computer vision tasks: face counting and crowd counting. We analyze the trade-offs in terms of performance, energy consumption and accuracy executing concurrent requests to simulate different workloads. The results show that while Edge devices are effective for single requests, the Fog resources offer better scalability and can be more energy efficient at high concurrency.
TORVS is a lightweight, open-source, parameterizable, dual-issue, in-order superscalar RISC-V core optimized for FPGA-based systems. It implements the RV32I instruction set and comprises two five-stage symmetric pipelines: fetch, decode, execute, memory access, and write-back. Parameterization is achieved through two orthogonal tuning axes: (i) branch predictor selection and (ii) the subset of instructions allowed for dispatch to the second pipeline. Each combination of parameters defines a unique design point with its own area–performance trade-off, characterized by FPGA usage, cycles per instruction (CPI), and maximum operating frequency. All configurations were validated with two standard benchmarks: CoreMark (1.176–1.600 CoreMark/MHz) and Dhrystones (1.708–2.331 DMIPS/MHz). The core was synthesized and successfully tested on an Intel DE10-Standard FPGA board. Several comparisons with a variety of processors were made to understand the advantages and disadvantages of a lightweight core. The conclusion shows that the lightweight design of TORVS delivers competitive performance in basic operations, whereas its efficiency declines for more complex workloads.
Although Internet of Things (IoT) devices have enabled Deep Learning (DL) applications to operate closer to end-users, they struggle to scale with modern, computationally heavy DL algorithms due to their limited processing power. A promising solution is to offload computation to remote edge and cloud servers throughout the IoT-edge-cloud continuum. However, inconsistent network conditions, variable latency, and the distinct accuracy and precision requirements of different DL applications hinder the ability to uphold necessary performance standards. To address these issues, we propose AdaptSwitch, a dynamic framework that selects both the DL model and the most suitable execution location across the IoT-edge-cloud continuum, based on the application’s model accuracy and latency needs according to network conditions at a given moment. We evaluated AdaptSwitch using real IoT, edge, and cloud devices running state-of-the-art DL models for image classification while considering network delay data from real-world 5G/Edge/Cloud traffic datasets. Experimental results demonstrate that our approach improves Quality of Experience (QoE) by reducing inference latency while ensuring accuracy requirements are met through adaptive model switching.
Predictive Maintenance plays a vital role in improving cost-efficiency of commercial fleets while improving reliability and safety metrics. This work presents a lightweight, vision-based framework for vibration-based suspension health monitoring that leverages video-derived vertical acceleration signals to detect anomalies on-the-fly. By classifying driving events that naturally excite the suspension (e.g., traversing speed bumps and potholes), our approach triggers condition-aware assessment of suspension health that uses a deep Autoencoder trained solely on health behavior where the reconstruction error serves as a proxy for degradation severity. The proposed solution employs CARLA simulator with custom maps and suspension configuration to emulate a wide range of fault scenarios, overcoming the lack of high-fidelity, labeled datasets. Experimental results demonstrates the model’s ability to detect suspension degradation, with a mean detection error of 19.13%. The framework operates without the need for dedicated vibration sensors or labeled data, making it highly suitable for real-world deployment in fleet-scale predictive maintenance systems.
Sparse linear algebra plays a crucial role across various fields because it reduces computational workload and optimizes memory usage. However, the irregular structure of sparse data creates difficulties for traditional software and hardware systems. Although dedicated accelerators can boost performance, they often lack the flexibility of general-purpose processors and depend on processor communication, which can introduce bottlenecks. This work tackles these challenges by proposing a tiling method that enhances the use of vector registers and by extending the RISC-V Vector (RVV) instruction set with a custom merge instruction. The project experimented with gem5 demonstrates that the tiled vector version delivered average performance gains of up to 1.31× and 1.67× for 95% and 65% sparsity, respectively, and the version incorporating the proposed merge instruction achieved speedups of up to 1.80× and 5.31× compared to a standard baseline. Additionally, the conversions required to implement the instructions did not exceed 2.5% of the area overhead for the SIMD Unit of the Sargantana processor.
The growth of camera networks in various areas has increased the demand for scalable intelligent real-time image processing solutions in surveillance and analytics. Python, known for its extensive AI and machine learning ecosystem, is a popular choice for image analysis tasks, but its limitations, such as the Global Interpreter Lock (GIL), pose challenges for efficient parallel processing. This paper presents a distributed architecture that uses message-oriented middleware (MOM) to orchestrate scalable, asynchronous data pipelines and enable Python-based image analysis. Using standard protocols such as RTSP, the architecture ensures compatibility with existing camera systems and supports seamless integration. Performance is evaluated using M/M/c queuing theory metrics, showing significant scalability improvements and near-linear throughput gains. The proposed architecture demonstrates not only scalability improvements but also applicability across a range of high-throughput scenarios.
The use of drones in firefighting has increased in recent years. In such applications, monitoring external temperatures is essential to ensure aircraft safety, particularly to prevent operation above a predefined threshold (70°C). This is critical when the drone approaches the fire, since the measured temperature reflects a past position, not the current one. Additionally, the drone takes some time to execute the maneuver and move away from the fire. While this delay depends on mechanical factors, the sensing latency is influenced by the embedded system architecture and is the focus of this paper. We present a method to estimate the temperature variation (ΔT) during this interval, based on an AADL model of the system. This model includes the drone’s temperature sensors and supports calculating their respective ΔT, helping to define when failsafe mechanisms should be triggered.
This paper presents a comparative analysis of three control strategies—PID, Fuzzy Logic, and Neuro-Fuzzy—for thermal regulation in a greenhouse cultivating tomatoes, peppers in Guaslán Grande, Ecuador. The experimental setup utilizes the Libelium Wasp-mote Pro embedded platform equipped with Smart Agriculture Xtreme sensors, enabling real-time monitoring of four critical environmental variables. A first-principles thermal model was developed and validated to guide the design and tuning of each controller. The PID controller was tuned using a Cohen-Coon method, the Fuzzy controller implemented 45 adaptive Mamdani-type rules, and the Neuro-Fuzzy controller combined a three-layer neural network with Takagi-Sugeno inference, trained via backpropagation. Experimental results showed the Neuro-Fuzzy controller achieved the lowest mean absolute error (0.51 ± 0.03°C), though these differences were not statistically significant (p > 0.05). Neuro-Fuzzy also led to a 63% reduction in integral squared error and 57% faster settling time compared to PID, with nearly identical average power consumption (0.2% difference). The system maintained internal temperature within ± 0.5°C, operated autonomously on lithium-ion batteries with OTA updates, and demonstrated that low-cost embedded systems (0.15 $/m2) are viable, scalable alternatives to industrial platforms (1.2$/m2) for precision agriculture in developing regions. The integration of physical modeling, intelligent control, and advanced sensing via Wasp-mote Pro enables efficient and adaptable climate management, contributing to sustainability and resource optimization in agricultural systems.
This paper introduces Go-Fast, a novel approach that accelerates the simulation and optimization of LUT-based FPGA circuits using GPUs. Unlike previous GPU-based simulators that target general digital circuits with event-driven approaches, Go-Fast employs batch simulation techniques specifically optimized for approximate computing scenarios. The system utilizes both data parallelism to execute testbenches across numerous threads and structural parallelism for simultaneous simulation and logic pruning. Go-Fast fully exploits GPU architectural features by maximizing register usage, minimizing memory access, and allocating more work per thread. This domain-specific approach generates efficient source-to-source CUDA code that achieves significant performance improvements: five orders of magnitude over the Verilator simulator and two to three orders over optimized multi-core implementations.
In this work, we explore the generation of synthetic data for Human Activity Recognition (HAR) using wearable sensors through a conditional version of SeriesGAN. We propose a few-shot adaptation mechanism that adjusts conditional embeddings or the final layers of the generator using a small number of genuine windows, enabling the generated data to be personalized for specific scenarios and potentially for new users in the future. We evaluate the approach using a real dataset developed by us, measuring both synthetic quality and classification performance. The results show that the conditional SeriesGAN improves classifier generalization, even in data-scarce settings, and that few-shot adaptation significantly increases accuracy when handling new scenarios.
Machine learning (ML) increases the demand for real-time inference and efficient hardware-based acceleration. This paper evaluates decision tree ensemble implementations on FPGA platforms coupled with High Bandwidth Memory (HBM). While traditional implementations often deploy hundreds of trees ("jungles") to maximize accuracy, our work demonstrates that carefully designed compact ensembles ("copses") can achieve comparable accuracy while better aligning with the architectural constraints of modern FPGA-HBM systems. We present a flexible wrapper for the automatic implementation of ML models, particularly tree-based models, on hardware. In addition, we developed a custom tree-based architecture with a variable pipeline aimed at maximizing throughput. By distributing smaller tree ensembles across parallel processing modules with direct HBM access, we eliminate the performance bottlenecks associated with large monolithic implementations that necessitate complex interconnection networks. We implemented the accelerator on the Xilinx Alveo FPGA platform, using 32 HBM channels, and performed the experiments in a realistic high-performance HPC cluster. Our results demonstrate that the accelerator processes 17 billion samples and low logic resource utilization while leveraging the high memory bandwidth of HBM.