
The growing use of machine learning in healthcare requires careful consideration of patient privacy, particularly when models are trained on small medical datasets where the risk of re-identification is heightened. This study analyzes how regression models operate under privacy-preserving constraints and assesses their susceptibility to privacy leakage. A stability-based membership-inference framework quantifies how much model outputs reveal about individual training samples. Noise-injection techniques are applied to reduce this risk, and their effects on privacy and predictive accuracy are evaluated. Experiments on a small real-world medical dataset show a trade-off between model utility and privacy, identifying conditions in which regression models remain reliable while limiting exposure to membership-inference attacks.
Scheduling scientific workflows in cloud computing environments, taking infrastructure provisioning into account, represents a complex optimisation problem, where minimising the total workflow completion time is a primary objective. This problem is computationally challenging, and metaheuristic methods such as genetic algorithms have been widely employed to obtain approximate solutions. In this study, an alternative approach based on a local search algorithm is proposed, employing four neighbourhood structures. These structures operate on the ordering of task assignments to virtual machines, enabling efficient exploration of the solution space through small, targeted modifications. A hybrid metaheuristic, known as a memetic algorithm, is used, combining the global search capabilities of genetic algorithms with the local refinement provided by the proposed local search. Experimental results demonstrate that the memetic algorithm produces competitive results compared with traditional genetic algorithms.
The analysis of retired electric vehicle battery packs is a critical enabler for sustainable circular economy strategies. However, real-world second-life battery data exhibit mixed chemistries, diverse topologies, and heterogeneous battery management system designs, which complicate reliable and scalable diagnostics. This paper introduces a physics-informed, ensemble-based semi-supervised machine learning framework to address these challenges. The proposed approach leverages multi-condition pulse testing to extract informative feature representations and integrates a simplified voltage-drop consistency constraint to enhance anomaly detection without reliance on manufacturer-specific models or proprietary parameters. Experimental results demonstrate strong generalization across battery chemistries and operating conditions while maintaining robust detection performance under a fixed false positive rate calibration.
This paper addresses a novel variant of the Vehicle Routing Problem that integrates two modern logistics problems: the use of electric vehicles and the optimisation of profit through zone-based pricing. The proposed problem, referred to as the Electric Vehicle Routing Problem with Zone-based Pricing, involves planning routes that maximise the profit, defined as the revenue obtained from serving a subset of customers that accept the service minus the travel energy costs. To tackle this complex problem, we propose two evolutionary algorithms as solution methods: a Genetic Algorithm, which searches directly in the space of potential routing solutions, and Genetic Programming, which evolves constructive heuristics that generate solutions. Experimental results demonstrate that Genetic Programming outperforms the Genetic Algorithm in all problem instances considered.
Cargo hijacking poses critical risks to global supply chains, yet most existing anomaly detection approaches are validated on single, homogeneous datasets and rely on static spatial boundaries that fail to capture realistic behavioral deviations. This study evaluates an unsupervised Long Short-Term Memory (LSTM) Autoencoder for real-time vehicle hijacking detection across heterogeneous logistics environments. The model was trained on a unified dataset of approximately 5.9 million GNSS observations integrating three distinct mobility profiles, and evaluated against road-constrained directional attacks generated via the Open Source Routing Machine (OSRM). A sensitivity-oriented decision threshold was adopted to prioritize threat detection over conservative filtering. Across 30 independent runs, the proposed framework achieved a mean Accuracy of 0.97 and a Recall of 0.99 on the unified dataset, with a clear separation between normal reconstruction error (MAE approximately 0.018) and attack error (MAE approximately 0.062). Inference latency averaged 0.69 s with a memory footprint of approximately 2.1 GB, without requiring GPU acceleration. These results confirm that training on diverse, heterogeneous data stabilizes the detection boundary and reduces false alarms without sacrificing sensitivity, providing a scalable and lightweight blueprint for securing logistics networks using only standard trajectory data.
Potato leaf blight remains one of the most destructive leaf diseases affecting potato crops in the Andean region, posing a significant threat to food security and the livelihoods of smallholder farmers. This work presents a computer vision framework for the automated detection of early blight (Alternaria solani) and late blight (Phytophthora infestans) in native Andean potato varieties using RGB imagery. Two datasets were employed: a localized dataset (2,766 images) and an extended-localized dataset incorporating additional distractor images that closely resemble those from localized dataset (3,666 images). Three convolutional neural network (CNNs) architectures, a custom CNN, EfficientNetB0, and MobileNetV3, were evaluated for classification performance and interpretability. MobileNetV3 achieved 100
Attention-Deficit/Hyperactivity Disorder (ADHD) remains difficult to assess objectively because diagnosis still depends largely on behavioral evaluation, while EEG-based computational approaches often face a trade-off between subject-independent robustness and physiological interpretability. To address this limitation, this work proposes an EEG-based framework that integrates Transformer-based temporal modeling with multi-scale Gaussian kernel connectivity regularized through α -Rényi mutual information, while incorporating Graph Spectral Analysis with XGBoost from Phase Locking Value-based connectivity to examine the class-discriminative structure of the learned representations. The approach was evaluated on a publicly available pediatric EEG dataset comprising 121 children under a 5-fold Stratified Group k-Fold cross-validation protocol designed to preserve subject independence. The experimental results show competitive performance, with 80.44 % accuracy, 80.37 % recall, and 81.35 % precision, surpassing EEGNet, Multi-Stream Transformer, IM-CBGT, and ANOVA-PCA SVM baselines; additionally, the analysis reveals that the learned connectivity patterns become progressively more discriminative from raw EEG to higher-level transformed representations. Overall, these findings support the potential of the proposed framework as a robust and interpretable alternative for objective EEG-based ADHD assessment.
The development of reliable computational tools for diagnosing and stratifying cognitive and neurodegenerative conditions remains a major challenge in translational neuroscience. Variability in clinical presentation, overlapping symptomatology, and heterogeneity in intermediate diagnostic categories complicate automated classification. This study proposes a methodology for constructing diagnostic tools based on spiking neural networks (SNNs) that uses relevant features calculated from virtual reality tasks to assess cognitive domains commonly affected in early dementia. These domains include episodic memory, executive function, and spatial navigation. Three diagnostic groups were considered: Normal Cognition (ED1), Subtle Cognitive Impairment (ED2), and Mild Cognitive Impairment (ED3). Feature selection was performed by ranking predictors according to the magnitude of their regression coefficients, and the most informative variables were used as inputs to the SNN classifier. To assess the influence of clinical labeling on model behavior, four stratification schemes (U0.5a, U0.5b, U0.5c, U1a) were defined and evaluated. Classification accuracy was computed on the training and independent test sets, and uncertainty was quantified using 95
Classical financial models struggle with the non-stationary noise of modern, news-driven markets, posing a severe risk to institutional investments and civil economic welfare during systemic shocks. This paper proposes Adapted DeepLOB, a hybrid CNN-LSTM architecture designed to predict directional trends and mitigate risk in highly volatile technology assets. The model fuses microstructural price data, macroeconomic indicators (S P 500, VIX), and NLP-quantified sentiment from financial news via FinBERT across a comprehensive historical period spanning from January 2010 to December 2023. Additonally, it integrates the Black-Scholes framework, using implied volatility to generate forward-looking risk metrics. Trained with a Focal Loss function to address market regime imbalance, the system establishes a robust defensive mechanism. Compared to a passive Buy Hold baseline, the model consistently increases risk-adjusted returns (Sharpe and Sortino ratios) while drastically reducing the Maximum Drawdown. The empirical results validate this multidimensional Artificial Intelligence approach not merely as a speculative tool, but as a resilient computational framework for capital preservation and systemic risk mitigation.
Lung cancer remains one of the leading causes of cancer-related mortality worldwide and represents a major public health concern. Although extensive research has identified numerous risk factors associated with this disease, further investigation is still required, particularly with respect to its underlying biological mechanisms. In prior work, a biological ontology was introduced to organize and harmonize heterogeneous biological data related to lung cancer and its subtypes. Nevertheless, lung cancer is inherently multifactorial, as biological processes interact with environmental and contextual factors that are commonly represented across disparate and unconnected data sources. To address this limitation, the present work extends the existing ontology to include both biological and environmental factors within a unified platform, which is embedded in a visual analytics tool. This approach enables the visualization and exploration of lung cancer–related risk factors and associated information across multiple heterogeneous data sources. By supporting cross-domain queries that integrate biological, environmental, and additional previously unconnected data, the proposed solution provides more functional and comprehensive access to knowledge relevant to lung cancer research.
With digital transactions, financial fraud has grown at a dramatic rate, with U.S. losses alone exceeding 12.5B per year and compromising the trust of global payment systems. Conventional ML methods fall short against complex relational patterns such as fraud rings and money laundering networks, which require graph-aware modeling. This survey offers a unified framework for Graph Neural Networks (GNNs) in financial fraud detection, addressing three research questions systematically. RQ1 demonstrates that GNNs outperform XGBoost with 12–25
This work presents a point-to-point trajectory planning framework for UAVs based on Q-Learning, integrated into a ROS2–Gazebo simulation environment. The problem is formulated as a discrete Markov Decision Process (MDP) over a two dimensional discretization of the space, with execution at constant altitude. Cylindrical obstacles are randomly generated in each experiment and represented using an inflated occupancy grid, which incorporates an explicit geometric safety margin during planning (2 m). Learning incorporates reward shaping based on potential to accelerate convergence while maintaining policy optimality. After training, the resulting trajectory is simplified, densified, and geometrically validated (by removing collinear points, interpolating intermediate waypoints, and sampling along segments on the inflated occupancy grid) before its execution in physical simulation. The results show stable convergence, high success rates, and spatial coherence between the learned value function and the geometry of the environment. The comparison between planned and executed trajectories confirms the transferability of the discrete model to a continuous dynamic system.
Affective touch plays a fundamental role in physical homeostasis and social well-being, differing from discriminative touch via its intrinsic connection to neurobiological reward and emotional regulation systems. This mechanism functions as a stress buffer by attenuating physiological reactivity and promoting states of psychological safety. Although it has traditionally been associated with the stimulation of hairy skin, recent research suggests that haptic interfaces can also modulate user valence and arousal. Within this framework, the proposed experiment studies tactile stimuli generated by a force-feedback haptic system to identify which specific haptic properties elicit positive emotional responses and modulate user affective state. Rather than assessing stable predispositions, the present study investigates immediate reactivity (state anxiety) to these force-rendered stimuli using a questionnaire based on the arousal dimension (calmness–activation). The resulting texture selection should be understood as a preliminary step that can subsequently inform the design and development of more complex haptic systems.
This work presents a rule-based approach for explaining the predictions of neural network classifiers, with a particular focus on Convolutional Neural Networks (CNNs). The FidexGlo algorithm was used to explain CNN decisions not at the pixel level but through patch-based conditions, where each antecedent expresses the average intensity of an image patch. Experiments were conducted on three standard benchmarks of grayscale images: MNIST, Fashion-MNIST, and EMNIST (letters). A ResNet-50 architecture was fine-tuned for each dataset. FidexGlo was then applied to the average values of image patches to extract rules closely aligned with the CNN’s behaviour. Qualitative visualisations demonstrate that the extracted rules identify meaningful discriminative regions—such as holes in digits, or shape boundaries. Overall, the study shows that our explainability method enables efficient, interpretable rule extraction from CNNs using patch-based explanations, offering a human-understandable view of deep model decisions. FidexGlo is available at https://github.com/Jean-Marc-B/dimlpfidex_Hepia .
This study addresses the challenge of 24-h ahead tropospheric ozone forecasting in the complex atmospheric basin of Seville, Spain. To capture the highly non-linear dynamics of photochemical pollutants, a hybrid CNN-LSTM-Attention framework is proposed, utilizing a 72-h sliding window in a multivariate approach. To explicitly prevent temporal data leakage and ensure robust evaluation, a purged walk-forward cross-validation scheme with fold-isolated preprocessing was implemented. Furthermore, novel optimization strategies, including Stochastic Weight Averaging (SWA) and Bayesian hyperparameter tuning, were applied during training to enhance structural robustness and convergence stability. Evaluated on an unseen 2023 testing set, deep sequence models demonstrated clear hierarchical superiority over traditional tree-based baselines. The CNN-LSTM-Attention SWA model achieved the highest performance ( R^2 = 0.699 , MSE = 368.30), significantly mitigating the temporal lag and systematic amplitude damping observed in simpler architectures. Conversely, while XGBoost reduced computational overhead to a few minutes, it systematically underestimated ozone peak concentrations. Ultimately, the synergy between convolutional feature extraction, attention mechanisms, and SWA stabilization yields superior generalization. This establishes a robust foundation for early warning systems in Mediterranean environments, successfully capturing the amplitude and phase of diurnal ozone spikes where conventional models fail.
This paper analyses various tools that are appropriate as teaching resources for data science or engineering at undergraduate or postgraduate levels. A comparison is performed between fsQCA, the online EDA (Exploratory Data Analysis) module within Nets4Learning and Minitab. We describe the functionalities of those frameworks and also the installation and running procedure. A visual tour is provided to showcase the main available tasks in these tools and to encourage other researchers or lecturers to introduce them in laboratory classes or as a complement for any theoretical lesson.
Reliable communication is one of the main challenges in disaster scenarios, where conventional infrastructure is often unavailable and mobile nodes exhibit highly dynamic behavior. This paper presents an artificial intelligence–based approach to enhance communication in wireless ad hoc networks under such conditions. A dataset was generated from scratch by integrating the ns-3 network simulator with BonnMotion to model realistic human mobility in disaster environments. From these simulations, two key features—channel utilization factor and queue packet size—were extracted and used to train a supervised learning model with the CatBoost algorithm. The model was validated with accuracy, precision, and F1-score, and then reintroduced into the simulation to support real-time decision-making. Experimental results show that the AI-enhanced strategy achieves substantial improvements in Packet Delivery Ratio (PDR), Throughput, and End-to-End Delay compared to a baseline without Quality of Service (QoS). These findings demonstrate the feasibility of integrating machine learning into communication layers to increase the resilience and efficiency of ad hoc networks for disaster response applications.
This study introduces a new measurement called the u_shape synchronization index (uSI). This metric uses voice recordings as a ’proxy’ for brain activity (phEG) to measure how well different brain regions, specifically those related to emotions and movement—work together during phonation. Building on a neuromechanical inversion pipeline, recorded phonation is transformed into probability-density phEG distributions. A u_shape synchronization index (uSI) is defined to capture the curvature of the μ -band relative to its immediate neighbors ( ϑ and high- α ), a signature of impaired motor–affective coupling in Autism Spectrum Disorder (ASD). This index was applied to longitudinal phEG samples collected from two ASD participants (one male, one female), each contributing 12 sustained-vowel recordings over six years, and compared these case series against a reference database of 12 male and 12 female healthy controls. Methodological validation combined one-sided bootstrap inference on uSI, for univariate contrasts, and multivariate deviation. Results show that the uSI reliably identifies pronounced μ -band troughs in multiple case recordings that exhibit subject-level heterogeneity and occasional discordance. It may be concluded that uSI is a promising, interpretable biomarker candidate that compresses complex spectral relationships into a single, testable index.
We present a standardized protocol and dataset design for multimodal acquisition in older adults with suspected cognitive impairment, aimed at supporting early identification research and data-intensive machine learning. The protocol combines four time-bounded discourse elicitation tasks, personal narrative, picture description, story narration, and procedural discourse, to capture complementary linguistic and cognitive demands. Speech is recorded with an external microphone and transcribed to provide aligned audio–text modalities, while an RGB camera facing the participant records facial behavior to enable analysis of nonverbal cues. In addition, we integrate an immersive VR/MR serious-game module inspired by our HoloDemTect framework, including daily-life activities such as a shopping-list task that requires selecting target items and placing them into a box. The VR/MR environment is fully instrumented to retain fine-grained telemetry (e.g., completion times, step-level latencies, interaction events, and error patterns), and gaze-related measures are recorded when supported by the hardware. To provide clinically meaningful reference outcomes for operational stratification and downstream classification, each session includes a brief self-assessment of cognition and instrumental functioning using Test Your Memory and the Lawton Brody IADL scale. Overall, the protocol yields a structurally complete and diverse dataset that enables joint analyses of speech, language, facial behavior, and immersive task performance, addressing common unimodality constraints in existing resources and facilitating multimodal deep learning baselines and future benchmarking.
This work proposes a graph-based methodology to prioritize the repair of telecontrol devices in medium-voltage electrical distribution networks. The grid is modeled as a graph-based digital twin built from infrastructure data. The priority of each telecontrol is defined as the installed power that becomes unsupported when the device fails, which is obtained by analyzing changes in network connectivity. The method was validated using real data from the distribution network of Granada (Spain) through simulation experiments. Results reveal a strong heterogeneity in the impact of telecontrol failures and show that repair priorities depend on the current network configuration. The proposed approach provides a scalable tool to support maintenance decision-making and improve grid resilience.