
Early detection of skin cancer (SC) is crucial for ensuring effective treatment and improved patient outcomes. Conventional convolutional neural networks (CNN) often face challenges such as overfitting due to limited training data, high computational requirements and insufficient focus on subtle lesion details. Additionally, the black-box nature of these models can impede clinical adoption by dermatologists. Therefore, incorporating the structure of ResNet50, we propose a new advanced model inspired by the reaction-diffusion process to enhance lesion detail recognition. The proposed reaction-diffusion-attention-based ResNet50 (RDA-ResNet50) employs RDA blocks to enable the model to focus on complex spatial and contextual details of the lesion regions. Besides, they helped to refine feature extraction and ameliorate inter-channel information flow. Each RDA block utilizes depthwise and pointwise convolutions to maintain a low computational footprint while efficiently capturing vital relationships. Furthermore, the diffusion process within the RDA blocks serves as a regularizer, smoothing feature maps to minimize noise and reduce the risk of overfitting. Moreover, gradient-based shapley additive explanations is used to visualize and understand the model predictions. The proposed system classifies dermoscopic images into nevus and melanoma categories with accuracy of 94.35
Accurate identification of HER2-positive breast cancer is essential for guiding targeted therapy, yet remains challenging in low-resource settings due to reliance on costly molecular diagnostics and limited access to specialized infrastructure. This study investigates whether a compact set of routinely collected clinical variables can support clinically meaningful HER2 status prediction using interpretable machine learning. A Rough Set Theory (RST)-based feature selection framework was applied to derive a minimal and clinically meaningful subset of features from the METABRIC dataset (n = 1,294). To ensure methodological rigor, target leakage was explicitly prevented, and model evaluation was conducted using stratified 5-fold cross-validation alongside an independent hold-out test set. Three classifiers—Decision Tree, Support Vector Machine (SVM, RBF kernel), and XGBoost—were trained under baseline conditions without resampling or cost-sensitive learning, enabling isolation of the effect of feature reduction. Performance was assessed using precision, recall, F1-score, and ROC AUC, with statistical significance evaluated via the Wilcoxon signed-rank test at the cross-validation level. RST reduced the feature set from 12 to 11 variables while preserving predictive performance relative to baseline models, with no statistically significant differences observed (p ≥ 0.05). Among the evaluated models, SVM achieved the best overall performance, with a ROC AUC of 0.72 and HER2-positive recall of 0.53 on the hold-out test set. All models demonstrated strong performance for the majority class, but reduced sensitivity for HER2-positive detection. The results indicate that the clinically curated feature set is highly informative and exhibits low redundancy, and that RST primarily serves to validate feature sufficiency while enabling modest model simplification. The findings also highlight the importance of model selection in imbalanced clinical classification tasks. Importantly, the use of routinely available clinical variables supports the development of interpretable and resource-efficient decision-support tools for HER2 classification, particularly in low-resource healthcare settings.
Enhanced Gazelle Optimization Algorithm (EGOA) is proposed as an improved Gazelle Optimization Algorithm (GOA)-based framework to address the convergence-precision trade-off associated with fixed step-size search strategies. The proposed enhancement introduces a self-adaptive step-size mechanism that promotes rapid exploration during the early search stage and finer exploitation near the optimum. EGOA was evaluated against large-step GOA (LGOA), small-step GOA (SGOA), Grey Wolf Optimizer (GWO), Particle Swarm Optimization (PSO) and the conventional Perturb and Observe (P O) algorithm using standard optimization benchmark functions and photovoltaic (PV) maximum power point tracking (MPPT) applications. In the optimization benchmarks, EGOA demonstrated a superior balance between convergence speed and solution accuracy compared with the fixed-step GOA variants and swarm-based baseline. In PV MPPT simulations, EGOA achieved exact MPPT under ideal irradiance and under all separated partial shading conditions (PSCs), where the average tracked power remained equal to the corresponding maximum power point (MPP) with near-zero standard deviation across repeated runs. By contrast, P O exhibited larger deviations and significantly higher variability under the more challenging shading cases, while GWO and PSO showed intermediate performance. Under sequential step-change test, EGOA consistently recovered the global MPP more rapidly than the benchmark algorithms, with reduced oscillatory behavior. These results demonstrate that adaptive step-size regulation strengthens GOA for PV MPPT applications by improving convergence speed, optimization accuracy and tracking robustness under both static and sequential step-change of PSCs.
The capacitated Chinese postman problem is a routing problem in which vehicles start and finish their routes at a depot and pass through all edges at least once. This study proposes a new mathematical model combining the rural postman problem and the capacitated Chinese postman problem. In the application section of the study, salting operations carried out during winter road icing in Palandöken, Aziziye, and Yakutiye, which are central districts of the Erzurum Metropolitan Municipality, were discussed. The study aims to find the shortest tour route using a genetic algorithm. As a result of the study, it was determined that the genetic algorithm performed well, and it aimed to contribute to the literature by introducing, for the first time in this field, the use of capacitated rural Chinese postman methods and a new mathematical model.
Heating, ventilation, and air conditioning (HVAC) systems are among the largest energy consumers in commercial buildings, and open-plan offices are particularly difficult to control because convective heat exchange between unpartitioned zones couples their thermal dynamics. Deep reinforcement learning (DRL) methods such as Deep Q-Networks (DQN) have shown promise for HVAC optimization, but their adoption is often limited by high computational cost, training instability, and demanding hardware requirements. This work investigates a lightweight and stable alternative: a coupling-aware, decentralized formulation of the on-policy SARSA (State–Action–Reward–State–Action) algorithm for real-time HVAC control in multi-zone open-plan offices. Each zone agent observes, in addition to its local conditions, an aggregate of neighboring-zone temperatures, allowing the tabular learner to compensate for inter-zone heat exchange. The agents were trained in a simulated six-zone office driven by TMY3 weather data and stochastic hybrid-work occupancy schedules, with a multi-objective reward balancing thermal comfort, energy use, and equipment switching. Relative to a clearly defined schedule-based baseline, the best-performing SARSA controller reduced annual HVAC energy consumption by 39.39
The study examines the use of electroencephalography (EEG) signals together with clinical data for diagnosing Alzheimer’s Disease (AD) and Frontotemporal Dementia (FTD). A machine learning framework was developed to combine EEG-derived features with demographic and cognitive information. The Fast Correlation-Based Filter method was used for feature selection, and several models, including XGBoost, Random Forest, AdaBoost, Neural Networks, and Gradient Boosting, were tested. Leave-One-Out Cross-Validation was used to evaluate model performance in the small-sample setting. The XGBoost model achieved the best performance with 89.8
Identifying critical nodes (CNs) in Wireless Sensor Networks (WSNs) is essential for ensuring connectivity, resilience, and overall operational efficiency. The failure or removal of such nodes can severely disrupt communication, leading to network fragmentation and performance degradation. However, the task of Critical Node Detection (CND) remains challenging due to the inherent complexity of real-world network topologies, the prohibitive computational cost of exhaustive search methods, and the limited scalability of traditional centrality-based approaches. To overcome these limitations, this paper introduces a hybrid optimization framework-Simulated Annealing-Improved Differential Evolution (SAIDE)-designed for efficient and scalable CND in WSNs. SAIDE leverages the global exploration capabilities of Differential Evolution (DE) and augments them with the local refinement and probabilistic acceptance features of Simulated Annealing (SA), resulting in a balanced and adaptive search strategy. The importance of the node is quantified using a composite influence metric that combines degree centrality, k-core decomposition, and inverse communication distance derived from the Friis transmission model. The primary optimization objective is to minimize the size of the Largest Connected Component (LCC) after node removal, thus maximizing network fragmentation. Extensive evaluations on synthetic network models-Random Geometric (RG), Erdos-Renyi (ER), and Barabasi-Albert (BA)-as well as real-world WSN datasets (ALE-WSN and LT-FS-ID) demonstrate that SAIDE consistently outperforms existing methods, including the TDE-degree baseline and advanced metaheuristics such as Multipopulation Differential Evolution (MPDE) and Memetic Algorithms (MA). SAIDE achieves more effective network fragmentation, faster convergence, and improved computational efficiency, establishing it as a robust and scalable solution for CND in large-scale WSNs.
The use of conventional heat transfer fluids has many limitations, and industrial applications require efficient heat transfer systems. Existing cooling systems suffer from low thermal conductivity and poor heat transfer, which result in reduced efficiency and high power consumption. This study proposes an optimised artificial neural network (ANN) and genetic algorithm (GA)-based artificial intelligence model to predict the thermal conductivity and viscosity of graphene nanoplatelets (GNP)–Cellulose nanocrystals (CNC)/Ethylene Glycol (EG)-water hybrid nanofluids, thereby improving heat transfer performance. Single GNP and hybrid GNP/CNC nanofluids are studied, with field-emission scanning electron microscopy and transmission electron microscopy used to characterise nanoparticles. A two-step preparation method was used for thermal conductivity and viscosity measurements across temperatures (30–80 °C) and volume concentrations (0.02–0.2 vol
Precise range estimation is considered one of the key elements in promoting electric car adoption, as it helps to avoid range anxiety. Traditional range estimation methods utilizing combinations of physics-based approaches and machine learning demonstrate some shortcomings, among which are unrealistic prediction values, poor generalizability, and insufficient physical interpretation. This paper presents an approach that utilizes the Physics-Informed Neural Networks (PINNs) method. The physics-informed neural network approach implies adding physical laws, such as energy conservation and equations of motion, which describe battery dynamics, to the process of training the neural network for predicting the driving range. Physics-Informed Neural Networks produce state-of-charge (SoC), power consumption, and range predictions while focusing on physical consistency by applying constraints to the model output of SoC, power, and rate. Experiments were carried out with real-life BMW i3 driving ranges and evaluated in a simulation traffic environment using SUMO. It was shown that the proposed PINN outperforms traditional neural networks in terms of accuracy (mean absolute errors): SoC MAE = 3.39
Type 1 Diabetes (T1D) requires continuous monitoring and accurate prediction of blood glucose (BG) for appropriate decision-making. In this study, a patient-specific BG prediction framework based on machine learning was evaluated using the OhioT1DM dataset under minimal preprocessing conditions. Several machine learning models were benchmarked and compared using standard regression metrics such as Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and coefficient of determination (R²), along with clinical assessment using Clarke Error Grid (CEG) analysis. Experiments were conducted using three literature-derived feature combinations: (1) BG alone, (2) BG + insulin-on-board (IOB) + meal (M), (3) BG + IOB + M + physical activity (PA). A sliding input window consisting of five historical time steps sampled at 5-minute intervals was used, while prediction targets were generated for 30- and 60-minute prediction horizons. A Tri-Ensemble model was constructed by combining the predictions of the three strongest benchmarked models using simple averaging as a robustness-oriented aggregation strategy. To improve interpretability and support future feature engineering, SHAP-based explainable AI (XAI) analysis was conducted as an exploratory feature relevance study over the complete available feature space. Patient-wise analysis was conducted to examine feature importance patterns and identify potentially informative variables for future blood glucose prediction studies. The proposed Tri-Ensemble demonstrated competitive predictive performance across multiple feature configurations and prediction horizons, with statistically significant improvements observed in several experimental settings.
We introduce MetaPKLot, a large-scale, harmonized dataset designed for vision-based parking lot management. It supports three key tasks: parking spot occupancy recognition, dwell time estimation, and parking spot extraction. MetaPKLot builds upon existing datasets, namely PKLot, CNRPark-EXT, and PLds, by adding over 1.3 million new annotations and revising approximately 900,000 existing ones, making it one of the largest publicly available human-annotated parking datasets to date. The new annotations include occupancy labels, segmentation masks, vehicle identifiers, and timestamps that enable dwell time computation. We also define standardized challenges and experimental protocols for each task, emphasizing not only predictive accuracy but also computational efficiency and generalization across unseen parking lots. Furthermore, we provide baseline methods that establish strong reference points for future research. By combining new annotations, standardized protocols, and baseline implementations, MetaPKLot serves as a comprehensive benchmark for developing and evaluating vision-based parking management systems under realistic, reproducible conditions.
The majority of Neural text classification models work by embedding the text in high dimensional vector spaces and measuring the semantic similarity within an implicit assumption of flat (Euclidean) geometry. However, real-world semantic structures, especially in fine-grained emotionally entangled datasets, may exhibit anisotropic correlations and non-uniform separability that are not fully captured by flat representations. We propose HyperSpectrum Geometry, a curvature metric-controlled hyperspherical manifold representation learning framework that combines a learnable global Riemannian metric tensor with normalized angular hyperspherical classification. Initially, our model maps sentence embeddings into a learned hyperspace, performs a Mahalanobis style metric deformation and finally, executes angular decision partitioning on a normalized deformed Riemannian hypersphere representation. We empirically evaluate the framework on structurally different datasets, including GoEmotions, AG News, TweetEval Emotion, DBpediaClassification, and Amazon Reviews Multi. The results show that fine-grained multi-label emotional entanglement produces substantially higher global metric deformation. In contrast, structurally separable topic and ontology datasets rely more strongly on angular partitioning with lower metric deformation. Therefore, the results reveal a deformation–separability pattern, that implies that, highly entangled multi-label classification requires stronger anisotropic metric deformation, whereas structured topic and ontology classification is mostly governed by angular separation. Over multiple random seeds, the model shows stable lightweight performance and at the same time, producing valuable geometric interpretations such as global metric deformation, metric condition number, and intra/inter class angular similarity. Our findings suggest that learnable metric deformation constitutes a useful and underexplored component of interpretable semantic representation learning.
Volatility forecasting is frequently framed as a modeling problem, yet attainable accuracy is fundamentally constrained by the information content of the data itself. This study evaluates multiple forecasting models, ranging from HAR and GARCH to Tree-based and Neural architectures, across 14 Global Equity Indices and Horizons form 1 day to 100 Trading days within a strictly chronological and capacity-controlled framework. Post diagnostics shows that realized volatility is organized by persistent regimes, strong cross-market synchronization, and discontinuous shocks. Under these conditions, forecasting performance is determined by parameter identifiability: low capacity model remain stable across horizons and deteriorate once parameter counts exceed available effective information. Attention mechanisms do not discover additional structure and instead coverage to weighting recent observations. The results indicate that for strongly dependent time series, nominal sample size is a misleading measure of learnability, and model capacity must be constrained relative to effective information rather than observation count.
The vehicle routing problem with time windows presents a complex challenge, involving the efficient scheduling of deliveries within predefined time windows while optimizing vehicle usage. In response, this paper introduces an advanced ant colony system approach that seamlessly integrates four heuristic techniques. The devised framework employs a ranking-based selection mechanism, utilizing the roulette procedure, to synergistically harness these heuristics, resulting in the generation of promising solutions. To counter stagnation and enhance solution diversity, we incorporate a no-improvement counter into the algorithm for solution selection. Additionally, we introduce a pioneering heuristic called 2-opt-special, inspired by the renowned 2-opt algorithm, to bolster the efficiency and effectiveness of the proposed approach. Computational results, derived from extensive comparisons with other ant colony-based algorithms and established solutions on well-known benchmarks comprising 100 and 400 customers, unequivocally affirm the effectiveness and competitiveness of our proposed algorithm in effectively addressing VRPTW challenges.
Disorders that affect the larynx and neurological control of speech—specifically Dysarthria, Laryngitis, Laryngozele, Vox senilis, Parkinson’s, and Spasmodic Dysphonia—can significantly disrupt vocal patterns and leave distinct acoustic signatures in a person’s voice. Early and accurate detection of these conditions is crucial for timely clinical intervention and effective patient care. In this work, we propose a deep learning–based framework for automatically classifying such vocal pathologies. Our approach leverages annotated speech recordings from the Saarbrücken Voice Database and the Italian Parkinson’s Voice and Speech dataset, from which mel-spectrogram features are extracted. After strengthening three advanced vision transformer architectures (DinoV3, EVA-02, and MaxViT), a comprehensive analysis of over a dozen advanced ensemble techniques was performed. The final system utilizes a Boosted Weighted Voting ensemble, enhanced with Temperature Scaling for model calibration. This sophisticated method learns the optimal voting weight for each model based on its performance on a validation set. The proposed ensemble framework achieves an accuracy of 86.84
We introduce ArabiDeepfake, a multi-domain Arabic deepfake-text dataset and benchmark designed for information security research. ArabiDeepfake spans five high-risk sources: news, government, social media, reviews, and e-commerce, with leak-free train/validation/test splits, balanced test sets, and deception-type labels (e.g., factual changes, omissions, satirical tone). Dialect/sector metadata is included where applicable. To mitigate contamination of the “real” class, we enforce authenticity controls (e.g., pre-ChatGPT timestamp filtering), alongside Arabic-only checks, de-duplication, and adjudicated spot reviews. Utilizing the MARBERTv2 encoder, we evaluate in-domain and cross-domain detection and compare a pooled single model against a five-model ensemble, reporting 95
Workplace safety monitoring is important for preventing accidents and providing secure industrial environments. Vision based Human Activity Recognition (HAR) has raised as an efficient model for identifying worker behaviors automatically from video data. However, traditional approaches undergo problems in capturing detailed spatial features and modeling long-range temporal dependencies in dynamic and safety critical scenarios. To address these limitations, this work presents a hybrid deep learning (DL) model that integrates Coordinate Attention (CA), Convolutional Neural Networks (CNN), and Transformer encoders for better HAR. In the proposed work, CNN is presented to extract hierarchical spatial features from video frames and the CA mechanism improves feature representation by embedding positional and channel-wise dependencies. Then, Transformer encoders are exploited for capturing temporal relations across frame sequences. The integrated model provides better spatio-temporal feature learning for accurate classification of workplace activities. The proposed model is evaluated on benchmark datasets that has routine office activities and industrial safety cases. Experimental results show that the proposed achieves superior performance and proved its effectiveness for workplace safety monitoring applications.
Groundwater prediction is vital for effective water resource management, that requires an understanding of the several factors influencing groundwater dynamics. While classic machine learning (ML) models have demonstrated strong predictive capabilities for groundwater levels by capturing complex nonlinear relationships, they face limitations such as extensive data requirements, overfitting, and lack of physical laws governing groundwater flow. In contrast, physics-informed neural networks (PINNs) integrate physical laws into neural architectures, offering enhanced prediction, and computational efficiency, especially in data-scarce environments. This study explores the applicability of PINNs to synthetic and real-world groundwater case studies, encompassing heterogeneous and homogeneous aquifers with varying boundary conditions and transient states. Using the groundwater flow partial differential equation as a constraint, PINNs are compared against classic ML techniques, including ensemble and deep learning models, across five diverse scenarios. Predictive performance and generalizability are evaluated using a calibrated numerical model based on MODFLOW 2005. The findings reveal that PINNs provide a compelling alternative to classic ML methods, particularly in handling complex heterogeneity of aquifer properties, which are crucial for advancing the capabilities of surrogate groundwater models.
Multimodal sentiment analysis (MSA) integrates textual, visual, and auditory information to better infer human emotions, yet it remains challenged by the imbalance in expressiveness across modalities and the loss of fine-grained cues during non-verbal encoding. These limitations hinder effective cross-modal interaction and restrict the performance of conventional fusion strategies. To address these challenges, we propose MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities. The core component, MagX, injects audio–visual cues directly into the XLNet backbone, enabling non-verbal information to contribute more meaningfully to contextualized textual representations. Built on this foundation, the Attention Multimodal Adaptation Gate (AMag) refines this interaction by applying multi-head attention to align modality-specific features with token-level representations, allowing the model to focus on sentiment-relevant non-verbal cues without overwhelming the textual backbone. In addition, the CrossCL module introduces a cross-modality contrastive learning objective that supervises the relationship between unimodal and fused representations across samples. This encourages consistent cross-modal alignment, mitigates imbalance between modalities, and improves the robustness of the learned feature space. Experimental evaluations on CMU-MOSI and CMU-MOSEI datasets show that MagXCL consistently outperforms previous methods, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.
In multi-class classification problems, one or more classes are often underrepresented, leading classifiers to favor the majority class(es). To address this issue, rebalancing strategies are necessary. Previous research on asymmetric label switching has shown that, when combined with principled neutral rebalancing methods, it consistently enhances classification performance. This study addresses multi-class imbalanced classification by introducing a framework that integrates binarization techniques with asymmetric label switching and neutral rebalancing mechanisms. A key aspect of this work is the derivation of Bayesian thresholds tailored for rebalanced dichotomies. We analyze the interaction between asymmetric label switching, rebalancing intensities, and various binarization strategies. Experimental evaluation across diverse datasets indicates that this integrated methodology provides consistent performance gains in multi-class imbalanced scenarios. These findings offer a basis for addressing imbalanced classification in domains such as medical diagnosis, fraud detection, and cybersecurity.