
Human Activity Recognition (HAR) plays a crucial role in applications such as health monitoring, physical training, and therapeutic support. Smartphones equipped with embedded sensors provide a convenient platform for these applications. However, running deep learning models continuously can drain device energy autonomy. Despite recent advances, few studies have explored how executing deep learning-based HAR models across different computational environments affects energy consumption, especially when combined with strategies that leverage application-level contextual information. This work develops and evaluates optimized deep learning models, including variants with early exits, in both local and remote execution scenarios. It also incorporates an adaptive sampling strategy that dynamically adjusts sensor reading frequency according to movement intensity (contextual information). Results show that the proposed approach reduces energy consumption and inference time while preserving model accuracy.
Detecting civil construction defects, such as material deterioration and cracks on masonry or structural members, is crucial for building safety and durability. Among all defects, cracks – whether visible or hidden – are critical issues that can compromise the integrity of buildings, bridges, roads, and other infrastructure elements. Artificial intelligence (AI) approaches, such as deep learning algorithms, can assist in the early identification and characterization of cracks, facilitating preventative actions to avoid future problems. In this study, we explore the application of deep learning with image segmentation techniques for crack detection and characterization in civil construction elements. We implemented a residual neural network capable of detecting cracks either in isolation or by mapping their distribution across surfaces such as concrete, bricks, steel, and wood. Additionally, we integrated the segmentation model SAM to improve the precision of crack segmentation in images. Through simulations and comparative analysis, we evaluated the performance of the models in accurately identifying and delineating cracks in civil infrastructure. The proposed model achieved an accuracy of 100\% and an intersection over union of 0.95. Despite these high-performance metrics, there remains room for error analysis to further refine the approach, particularly in complex or edge-case scenarios. These results demonstrate the efficacy of the proposed approach in achieving accurate crack detection.
The reliable operation of photovoltaic (PV) systems depends on the early detection of failures such as hotspots, which are localized overheating regions typically caused by partial shading, cell mismatch, soiling, or internal defects. In these regions, affected cells operate in reverse bias and dissipate energy as heat instead of generating electricity, leading to efficiency losses, accelerated material degradation, and, in severe cases, irreversible damage or fire hazards. Conventional inspection methods, including visual evaluation and electrical testing, are time-consuming, labor-intensive, and unsuitable for large-scale solar plants. The use of Unmanned Aerial Vehicles (UAVs) equipped with thermal cameras has emerged as a promising alternative for autonomous inspection, enhancing operational efficiency and reducing human risks. In this work, we propose a hotspot detection approach based on a U-Net segmentation network applied to thermal infrared imagery of PV modules. The model was trained on a dataset that combined publicly available and laboratory-acquired images, utilizing preprocessing and augmentation strategies to enhance robustness. Experimental results demonstrate that the proposed method effectively identifies hotspot regions with precise delineation of their morphology and spatial distribution, outperforming bounding-box-based approaches such as YOLO in terms of fine-grained localization. This level of detail is crucial for predictive maintenance, as it enables the accurate measurement of hotspot size and shape. The findings highlight the potential of integrating UAV-based thermal inspection with deep learning segmentation models as a scalable and reliable solution for autonomous PV system monitoring and maintenance.
The main objective of this work is to investigate the efficiency of data weighting techniques for calculating weighted principal components in gender and facial expression experiments, considering classification and reconstruction problems. Specifically, the methodology consists of generating spatial weights (statistical maps) to weight the pixels of the input images. Then, the weighted data are used as input for the principal component analysis (P CA) algorithm. We will consider the following techniques for calculating spatial weights: (a) Student’s t-test; (b) Hyperplanes computed using SVM (Support Vector Machine); (c) Shannon entropy calculated pixel-by-pixel; (d) Inverse Shannon entropy calculated pixel-by-pixel; (e) Jensen-Shannon divergence. The application of techniques (a), (c), (d), (e) for calculating spatial weights in face recognition is the main contribution of this work. The evaluation of the efficiency of the different principal components obtained with each technique is done by analyzing the results of image reconstruction and classification over FEI and FERET databases. For classification, the following classifiers were used: K-Nearest Neighbors and Mahalanobis Distance. A quadtree-based methodology was also used to reduce small local variations in the statistical maps, to test its effect on reconstruction and classification.
Congestive heart failure (CHF) is a serious medical condition associated with high mortality rates. To improve prognosis and treatment, exploring new strategies is essential. This study investigates the use of electronic medical records to train machine learning models in predicting survival and hospitalization time for patients with CHF. Using data from 299 patients collected in Faisalabad, Pakistan, a suite of algorithms was evaluated, including MLP, logistic regression, Random Forests, decision tree, k-Nearest Neighbors, Naive Bayes, and Gradient Boosting. The methodology employed the SMOTE technique for class balancing and a rigorous 10-fold stratified cross-validation for performance evaluation. The Random Forest model emerged as the top performer, achieving a mean accuracy of 0.80 (±0.05) and an F1-Score of 0.80 (±0.05) in predicting patient survival. Statistical significance tests confirmed that the superiority of the Random Forest over the worst-performing model (MLP) is statistically significant (p < 0.05). For the prediction of hospitalization time, an error rate of 26% was observed. These findings underscore the statistically validated potential of machine learning models in predicting clinical outcomes in patients with CHF, representing an innovative approach to improve diagnostic efficiency and reduce the impact of CHF on public health.
This paper proposes automatically estimating postural balance test scores and their subsystems using stabilometry results. These results are obtained using a force platform that captures the variation in the person's center of pressure position over time. On the other hand, postural balance test scores are conducted by qualified health professionals and comprise the assessment of various subsystems that can provide specific diagnoses separately and, when combined, can form a comprehensive assessment method called the MINI-BESTest. Early diagnosis makes it possible to identify the risk of falls due to advancing age or related limiting illnesses. Using a hierarchical approach, which makes it possible to automate the process and reduce the subjectivity of the assessment, it is possible to estimate the MINI-BESTest score and its Reactive Postural Control subsystem with an accuracy of, respectively, 17% and 20% higher than the state-of-the-art.
The degradation of bearings is a critical issue in mechanical systems, leading to performance deterioration and potential failures. Classifying degradation stages in bearing systems is essential for effective maintenance and the prevention of unexpected downtime. Based on this, the study presents a methodology for detecting degradation stages in bearing systems by employing multiple techniques, such as Fourier transform, principal component analysis, and decision trees. The results demonstrate that this approach enables the detection of the current degradation stage using techniques that generally require lower computational effort compared to methods previously used in the literature. Additionally, the comparative analysis reveals superior performance in terms of bearing lifespan after fault detection in advanced stages, surpassing results from the literature. Moreover, interpreting the results through SHAP plots identifies value ranges within crucial features for fault classification, such as mean absolute deviation and linear predictive cepstral coefficients. The dependency analysis between these features provides valuable insights into failure patterns. The achieved classification accuracy exceeded 90%, demonstrating the effectiveness and efficiency of the proposed method in enhancing predictive maintenance for rolling bearings.
This study presents an analysis of the scalability and dispersion of results in Federated Learning (FL) using two algorithms: EnBaSe, based on entropy, and Random, a random selection approach. The Random algorithm ensures that each member of the population has an equal probability of inclusion. At the same time, EnBaSe calculates the information gain and selects the most informative samples for the neural network. Both algorithms were applied in federated learning scenarios with data distributed non-independently and non-identically (Non-IID). The MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100 datasets were used for the evaluation, representing different levels of computer vision classification. The results show that the EnBaSe algorithm achieves high accuracy while halving computational and energy costs compared to training with all samples from the datasets. In addition, EnBaSe demonstrated greater resilience to variability, showing low variance and a more stable distribution, especially in Internet of things (IoT) environments with limited computational resources.
Telecommunications improvements in the last decades have led to an increase in the available data rate and reliability, and to reduce latency. Massive MIMO (mMIMO) has emerged as a promising solution to replace simple antenna setups. This arrangement involves a large number of transmitter antennas compared to the served terminals in the covered area, enabling higher data rates through the key characteristics of mMIMO. Despite its advantages, the utilization of such large arrays is energy-intensive. With fewer antennas, it is possible to achieve high data rates with reduced power consumption. Antenna selection is crucial for energy efficiency while maintaining computational efficiency. Additionally, power allocation in the downlink for each served terminal is an important step to consider. This work aims to propose a joint algorithm for selecting antennas and distributing energy among terminals, simplifying the process through low-complexity algorithms. The results show the penalty in sum-rate capacity for each proposed method. The two proposed algorithms demonstrate an adequate approach to the target value asymptotically as the number of transmitter antennas increases. This implies the possibility of achieving the joint process with significantly lower computational costs compared to perform antenna selection and power allocation separately.
This article provides a survey of open-world learning in the context of image segmentation and a subset of associated tasks that exhibit desirable characteristics for real-world applications, such as autonomous driving, industrial inspection, medical diagnosis, remote sensing, among others. The objective is to identify the main approaches, challenges, and gaps in this research field. Through a rigorous procedure, 39 articles published between 2013 and 2023 were selected for analysis. Then, a review was conducted to examine these documents, and three main research questions were posed. After an extensive analysis, results indicate that open-world learning for image segmentation has been explored only recently, and it seems to be a promising field of research for the upcoming years. Established on this reviewed literature, we provide the potential directions and open gaps for future works on this topic.
Automated early recognition of enemy radar emissions are essential for the survival of a warship. Radar signals intercepted by a passive digital Electronic Support Measure (ESM) receiver can be classified based on the type of intrapulse modulation. The modulation classification is typically based on features extracted from the preprocessed radar’s signal. Low Probability of Intercept (LPI) radars can use phase or frequency modulated signals to make radar emissions difficult for the enemy to detect. This paper proposes the use of a new feature, the symmetry measured in a Time-Frequency (TF) matrix, to improve intrapulse radar modulation classification. The analysis of the symmetry in a STFT matrix is characterized by an image processing problem, where the matrix is interpreted as a grayscale image. This paper also proposes the use of Weightless Neural Network WiSARD in identifying symmetry patterns in the Short Time Fourier Transform (STFT) matrix, a characteristic present in signals with Barker and Polytime phase modulations and which can be used by classifiers to discriminate them from signals with other types of modulations such as polyphase modulations (P1, P2, P3, P4 and Frank), linear frequency modulation (LFM), and Non Linear Frequency Modulation.
This paper presents an adaptive immune fuzzy quasi-sliding mode kinematic control integrated with a PD dynamic control for the trajectory tracking and the leader-follower formation control by nonholonomic differential-drive wheeled mobile robots under incidence of uncertainties and disturbances. An immune regulation mechanism bio-inspired approach with reaction effect established by novel fuzzy rules set to adjust the control effort adaptively is designed, also using a fuzzy boundary layer and introducing an adaptation law for the immune portion gain online adjustment in such a way that they can also avert parameter drift, dealing with the drawbacks of a classic first-order sliding mode control, suppressing chattering and still maintaining the robustness with no a priori knowledge of the bounds of the disturbances. An obstacle avoidance strategy with a reactive method and variable avoidance radius is also proposed. The stability analysis is performed based on the Lyapunov theory. Simulation results demonstrate the proposed control effectiveness.
Time Series (TS), including stock prices, temperatures, and health markers are monitored all the time and everywhere. Time-dependent events, and TS are frequently studied in the scope of forecasting, where past values of the TS are used to preview future ones. On the other hand, Time Series Classification (TSC) aims at creating models that label TS instances. Open Set Recognition (OSR) techniques, designed to classify known samples and detect unknowns simultaneously, have been less applied in TSC compared to other fields like image classification. Existing methods have limitations, such as not using known-unknown in the training stage, limited transfer learning across Neural Network models, and experiments with benchmark datasets with narrow coverage. This study introduces a novel OSR approach for TSC by addressing the mentioned issues and using a blending ensemble of Neural Networks with an OpenMax layer. The results vouch for the model’s performance and potential superiority as an alternative to existing methods in open-set recognition tasks.
Point clouds generated in simulators avoid collection time costs and provide a high and organized amount of point clouds, an ideal scenario for deep learning networks. However, these networks have limitations when applied to real point clouds. This work proposes a multilayer perceptron-based method to classify 3D objects based on real point clouds obtained using LiDAR sensors. The method includes a pre-processing step that normalizes and adjusts the point clouds in the 3D Cartesian plane to overcome discrepancies in the point distribution. Furthermore, we created a dataset and used the ModelNet dataset for comparison purposes. The proposed neural network, Lidar3DNetV2, achieved 98.47% and 125 μs in accuracy and test time with real data, respectively. The pre-processing step provided a significant increase in the classifier’s performance. Finally, the proposed method performs better than other state-of-the-art networks considering real point clouds.
Financial markets are competitive environments influenced by several variables and sectors. Wrong decisions can compromise several areas and cause chain reactions that could disrupt various sectors of the economy. In recent years, intelligent models have been used as tools to aid decision-making in financial markets. Deep learning models stand out among them, as they can achieve good generalization with large datasets. The main goal of this paper is to introduce and evaluate deep learning for solving financial problems. We document the process and present the techniques employed to develop models using a dataset containing over 2 million financial data observations. We believe this paper could guide researchers working on similar problems by suggesting resources that can be used and steps that can be followed in similar scenarios, narrowing down the search for efficient financial machine learning models.
Considering the large number of people that suffer with epilepsy, there is great interest in developing a system capable of predicting the occurrence of the seizures. Warning patients about the proximity of a seizure can help them to avoid dangerous situations. Many efforts in the development of this system have been placed using, basically, machine learning techniques. In most of these works, regardless of the technique employed, it is assumed that the system is able to learn when the patient’s brain signals, obtained from electroencephalogram (EEG), change from interictal to pre-ictal state. Some published works do not explicitly clarify the method of choosing training and test data. This means that good results may be due to the use of an inadequate method. That is, the EEG segments may have been temporarily mixed up during training, which would help the network correctly predict the seizure. However, this would not work from a real point of view. The objective of this work is to investigate, through experiments, the effect of the way of choosing the training and test data on the accuracy of a seizure prediction system. Experimental results showed that the accuracy of the systems increases when training is performed with windows close to the seizure used for testing.
Machine learning techniques have shown success in classifying hand gestures. As the prevalence of prosthetic devices continues to rise, the adoption of non-invasive technologies, such as surface electromyography (sEMG), becomes paramount. This study systematically assesses the isolated influence of classification algorithms within hand gesture recognition (HGR) systems using sEMG data and dynamic time warping (DTW) based features. This approach effectively handles temporal variations in sEMG signals by leveraging DTW, ensuring input features are invariant to gesture speed. Six supervised learning classifiers were evaluated: the multilayer perceptron, support vector machine, logistic regression, linear discriminant analysis, k-nearest neighbors, and decision tree. Cross-validation was employed to fine-tune the segmentation hyperparameters, significantly improving results. To ensure reproducibility, the source code has been made available, the proposed system design has been detailed, and the evaluation protocols have been described. Our findings indicate that logistic regression outperformed other classifiers in this setup, achieving 95.2% accuracy in classifying six hand movements from ten healthy individuals, representing a 1.6% improvement over the best previously reported performance using the same publicly available dataset. Future research will assess the proposed HGR system’s generalization capability on larger datasets suitable for training more complex classifiers, including deep learning models.
The increasing demand for clean energy presents challenges in energy supply management, largely due to their intermittency. Photovoltaic power generation, in specific, is greatly affected by weather factors, which may render power grids susceptible to instability, quality and balance issues. In this context, photovoltaic power generation forecasting is crucial not only to enhance the management of diverse energy sources through generation planning, but also to ensure widespread adoption of photovoltaic energy. To address the predictability issue in generation, this study aims to investigate the combination of satellite data with meteorological data to predict the energy generation potential in photovoltaic panels within 30, 60, 120, and 180-minute horizons. For this purpose, images from the GOES-16 satellite are used in combination with data from a ground-based weather station, located at Florianópolis – Santa Catarina – Brazil. The data is fed to a convolutional neural network, where convolutions are employed to extract features from the satellite images, aiming to establish a relationship with solar irradiation. The output of the convolutional network serves as input for a multilayer perceptron network, which utilizes the data to predict the Global Horizontal Irradiance (GHI). Our results support that models incorporating satellite images provide forecasts approximately 41% better for the 30-minute horizon and 21% better for the 180-minute horizon, when compared to models without satellite images.
The inspection and maintenance of solar panels face significant challenges due to the dangers involved in checking for potential defects in photovoltaic panels. However, the advancement of studies to facilitate this verification is limited by the lack of image datasets for training algorithms that can perform this task. Based on this principle, the use of Convolutional Neural Networks (CNNs) for training and Data Augmentation for creating artificial data from real images is common. With that said, this work aims to explore configurations and models of Data Augmentation for the classification of defects in solar panels using CNNs. The proposed methodology consists of four experimental stages, where the application of Zoom, rotation, horizontal displacement, and vertical displacement transformations are evaluated, followed by filtering the best results and combining them to find the ideal classification model. The performance of the model proved to be most effective when using a 35-degree rotation for creating artificial images, thus achieving an 88% F1-Score and 87.64% accuracy.
Identifying Customer Induced Damage (CID) is a key part in warranty programs of electronics manufacturers. CID is defined as any damage in the unit performed by an unauthorized person including the customer in a Printed Circuit Board (PCB). In such cases, damaged units are not covered by warranty. The inspection of CIDs is usually performed by humans which may be costly and error prone. Modern computer vision techniques for object detection using deep neural networks can automatically and accurately detect CIDs on PCBs. The training of such networks requires a large labeled dataset of image examples of CIDs. Daily, hardware factories and repair centers generate hundreds of unlabeled images. Labeling them manually is laborious and time-consuming. Therefore, it is crucial to label the minimum amount of images such that the trained neural network can achieve comparable accuracy as if it were trained with the whole dataset. To this end, we propose an active learning approach that selects the most informative images for the object detector. For that, our approach is based on the uncertainty of the object detector, i.e., it selects new images based on class probability distribution given by the object detector. Also, we tackle some challenges that are intrinsic to this problem: i) it is a multiclass object detection problem since there are many types of defects; ii) there is a class accuracy imbalance; iii) there is a focus on recall, e.g. false positives are less harmful than false negatives, and iv) there are many images with no object which should not be selected for labeling. We evaluate this approach by using it to iteratively sample data, train and evaluate a model, and compare it with randomly sampled data. The results show that our method consistently outperforms random sampling by an average margin of 21.6%, proving to be a viable alternative for reducing the labeling cost and increasing detection accuracy in this domain.