Due to the relatively small figure of research conducted on Arabic email spam about the number of Arabic-speaking individuals using the internet, comes challenges related to Natural Language Processing in morphologically rich languages such as Arabic. The target goal of this research is to create a classification framework for spam email detection in Arabic. The recent dataset obtained from Kaggle contains phishing and legitimate emails written in Arabic. Term frequency - Inverse document frequency (TF-IDF) and word embeddings are some of the text preprocessing methods used in this research. Support Vector Machine (SVM), Naïve Bayes, and Random Forest machine learning models are used, along with Long Short-Term Memory (LSTM) deep learning methods. The SVM performed best in the comparative analysis, peaking at an ROC-AUC of 98% and F1 score of 99%. These results will aid in shifting Arabic spam detection and email security in the Arabic digital world forward, showcasing the positive role machine learning can play in this process.
The rapid growth of the Internet of Things (IoT) has changed different industrial sectors, along with introducing significant cybersecurity challenges due to the increasing number of connected devices. Intrusion Detection Systems (IDS) play a crucial role in protecting IoT environments by identifying and mitigating cyber threats in real time. This paper provides a comprehensive survey of IDS techniques in IoT, highlighting the challenges, existing approaches, and future directions. The study explores Machine Learning (ML) and Deep Learning (DL) techniques, which enhance IDS performance by improving anomaly detection and threat classification. Additionally, we address key cybersecurity concerns, including integration with emerging technologies, scalability, resource constraints, and data privacy. By presenting a structured taxonomy and analyzing current IDS methodologies, this work aims to serve as a reference for researchers and practitioners seeking to develop more efficient, intelligent, and scalable IDS solutions for IoT security.
Predicting students’ performance is one of the essential educational data mining approaches aimed at observing learning outcomes. Predicting grade point average (GPA) helps to monitor academic performance and assists advisors in identifying students at risk of failure, major changes, or dropout. To enhance prediction performance, this study employs a long short-term memory (LSTM) model using a rich set of academic and demographic features. The dataset, drawn from 29,455 students at Saint Cloud State University (SCSU) over eight years (2016–2024), was carefully preprocessed by eliminating irrelevant and missing data, encoding categorical variables, and normalizing numerical features. Feature importance was determined using a permutation-based method to identify the most impactful variables on term GPA prediction. Furthermore, model hyperparameters, including the number of LSTM layers, units per layer, batch size, learning rate, and activation functions, were fine-tuned using experimental validation with the Adam optimizer and learning rate scheduling. Two experiments were conducted at both the college and department levels. The proposed model outperformed traditional machine learning models such as linear regression (LR), K-nearest neighbor (KNN), decision tree (DT), random forest (RF), and support vector regressor (SVR), and it surpasses two deep learning models, recurrent neural network (RNN) and convolutional neural network (CNN), achieving 9.54 mean absolute percentage error (MAPE), 0.0059 mean absolute error (MAE), 0.0001 root mean square error (RMSE), and an R² score of 99%.
The detection of spam reviews in multilingual environments remains a challenging task due to linguistic diversity, data imbalance, and semantic complexity. This paper proposes a novel hybrid model that integrates Twin Support Vector Machine (TwinSVM) with Harris Hawks Optimization (HHO) for simultaneous parameter optimization and feature selection. To enhance semantic understanding, sentiment-based features are incorporated alongside pre-trained word embedding models-BERT, FastText, and MUSE-across English, Arabic, and Spanish datasets. Our approach generates 24 high-quality datasets using embeddings with 100 and 400 dimensions, including a combined multilingual set. Experimental results demonstrate that our proposed HHO-TwinSVM model consistently outperforms conventional classifiers and metaheuristic-enhanced SVMs, achieving accuracy improvements of up to 9.44% and enhanced robustness in low-resource languages. This integrated framework represents a scalable and adaptable solution for multilingual spam detection. Four detailed experiments were conducted in this study, each designed to address and demonstrate a specific aspect of the proposed approach. Across all experiments, the method outperformed existing algorithms, achieving impressive accuracy rates of 92.9741 %, 89.0314%, 80.3580%, and 85.0859% on Arabic, English, Spanish, and multilingual datasets, respectively. Subsequently, sentiment analysis features were incorporated to further enhance detection performance, resulting in improvements of 1.0994%, 2.6674%, 9.4430%, and 8.7448%, respectively. A comprehensive analysis of the experimental results, including the influence of reviews and sentiment features, is also presented.
As the business world shifts to the web and tremendous amounts of data become available on multilingual mobile applications, new business and research challenges and opportunities have been explored. This research aims to intensify the usage of data analytics, machine learning, and sentiment analysis of textual data to classify customers’ reviews, feedback, and ratings of businesses in Jordan’s food and restaurant industry. The main methods used in this research were sentiment polarity (to address the challenges posed by businesses to automatically apply text analysis) and bio-metric techniques (to systematically identify users’ emotional states, so reviews can be thoroughly understood). The research was extended to deal with reviews in Arabic, dialectic Arabic, and English, with the main focus on the Arabic language, as the application examined (Talabat) is based in Jordan. Arabic and English reviews were collected from the application, and a new model was proposed to sentimentally analyze reviews. The proposed model has four main stages: data collection, data preparation, model building, and model evaluation. The main purpose of this research is to study the problem expressed above using a model of ordinal regression to overcome issues related to misclassification. Additionally, an automatic multi-language prediction approach for online restaurant reviews was proposed by combining the eXtreme gradient boosting (XGBoost) and particle swarm optimization (PSO) techniques for the ordinal regression of these reviews. The proposed PSO-XGB algorithm showed superior results when compared to support vector machine (SVM) and other optimization methods in terms of root mean square error (RMSE) for the English and Arabic datasets. Specifically, for the Arabic dataset, PSO-XGB achieved an RMSE value of 0.7722, whereas PSO-SVM achieved an RSME value of 0.9988.
Breast cancer recurrence prediction is crucial for patient care, necessitating advanced methodologies to enhance precision and facilitate therapeutic decision-making. This paper evaluates how oversampling techniques affect breast cancer recurrence prediction on the base classifiers. Our study is based on the Jordan Breast Cancer Dataset (JBRCA), which contains 21 features and 9723 instances. The JBRCA is a breast cancer dataset extracted from the King Hussein Cancer Center's (KHCC) registry database. By illustration, the study describes the imbalance problem facing most real datasets. This imbalance can lead to low model performance. Over-sampling techniques can help balance the data when there are fewer recurrence cases than non-recurrence. The proposed methodology includes data preparation and involves oversampling techniques for model development. We used base classification models such as logistic regression, decision tree, k-nearest neighbors, Gaussian Naive Bayes, and multilayer perceptron, combining them with oversampling techniques SMOTE, Random, ADASYN, SVMSMOTE, and Borderline-SMOTE to evaluate the compiled dataset. We examine the impact of the base classifiers along with oversampling techniques using accuracy, sensitivity, specificity, G-mean, and ROC-AUC to assess the model's efficacy.
The Internet of Things (IoT) has evolved as a significant area of impact and opportunity due to the billions of connected devices that have been distributed around the world. IoT devices, on the other hand, are susceptible to being compromised and hacked. When it comes to computing power and storage capacity, these Internet of Things devices are more vulnerable to cyberattacks than traditional endpoints such as smartphones, tablets, and laptops. This study introduces and assesses a machine learning-based cyberattack detection system. The suggested method employs a Support Vector Machine (SVM) classifier with the Harris Hawks Optimization (HHO) algorithm. The HHO technique improves SVM classifier hyperparameters, while the SVM performs malicious and normal stream classification based on the best-chosen model, and produces the optimal solution for feature weighting. The utility and capacity of the suggested approach to improve detection performance is demonstrated through proper scientific testing utilizing the LITNET-2020 benchmark dataset against six well-known classification algorithms and four metaheuristic-based classifiers using five reliable assessment measures.
The healthcare industry has been suffering from fraud in many facets for decades, resulting in millions of dollars lost to fictitious claims at the expense of other patients who cannot afford appropriate care. As such, accurately identifying fraudulent claims is one of the most important factors in a well-functioning healthcare system. However, over time, fraud has become harder to detect because of increasingly complex and sophisticated fraud scheme development, data unpreparedness, as well as data privacy concerns. Moreover, traditional methods are proving increasingly inadequate in addressing this issue. To solve this issue a novel evolutionary dynamic weighted search space approach (DW-WOA-SVM) is presented in the current study. The approach has different levels that work simultaneously, where the optimization algorithm is responsible for tuning the Support Vector Machine (SVM) parameters, applying the weighting procedure for the features, and using a dynamic search space to adjust the range values. Tuning the parameters benefits the performance of SVM, and the weighting technique makes it updated with importance and lets the algorithm focus on data structure in addition to optimization objectives. The dynamic search space enhances the search range during the process. Furthermore, large language models have been applied to generate the dataset to improve the quality of the data and address the lack of good dimensionality, helping to enhance the richness of the data. The experiments highlighted the superior performance of this proposed approach than other algorithms.
Cyber threats are an ongoing problem that is hard to prevent completely. This can occur for various reasons, but the main causes are the evolving techniques of hackers and the neglect of security measures when developing software or hardware. As a result, several countermeasures will need to be applied to mitigate these threats. Cyber-threat detection techniques can fulfill this role by utilizing different identification methods for various cyber threats. In this work, an intelligent cyber threat detection system employing a swarm-based machine learning approach is proposed. The approach involves using Harris Hawks Optimization (HHO) to enhance the Support Vector Machine (SVM) for improved threat detection through parameter tuning and feature weighting. Furthermore, various cyber-threat types have been considered, including Fake News, IoT Intrusion, Malicious URLs, Spam Emails, and Spam Websites. The proposed HHO-SVM has been compared to other approaches for detecting all these types collectively. The HHO-SVM outperforms all algorithms in most types (datasets). The proposed approach demonstrated the highest accuracy across seven datasets: FakeNews-1, FakeNews-2, FakeNews-3, IoT-ID, URL, SpamEmail-2, and SpamWebsites, achieving average accuracy of 68.251%, 68.729%, 79.049%, 95.254%, 100%, 96.681%, and 93.975%, respectively. Additionally, a thorough analysis of each cyber-threat type has been conducted to understand their characteristics and detection strategies.
The rapid expansion of medical data, characterized by its complex high-dimensional attributes, presents numerous promising opportunities and substantial challenges in healthcare analytics. Adopting effective feature selection techniques is essential to take advantage of the potential of such data. This research presents a modified algorithm called (mDA), which is the hybrid algorithm between the Evolutionary Population Dynamics and the Dragonfly Algorithm. This method combines Evolutionary Population Dynamics’s strength with the Dragonfly Algorithm’s flexible capabilities, offering a robust evolutionary machine learning approach specifically designed for medical data analysis. By integrating the dynamic population modeling of Evolutionary Population Dynamics with the adaptive search techniques of Dragonfly Algorithm, the proposed mDA significantly improves accuracy, reduces the number of features, and obtains the minimum average of the fitness scores. Comparative experiments conducted on seven diverse medical datasets against other established algorithms confirm the superior performance of the proposed mDA, establishing it as a valuable approach in examining complex medical data.
ABSTRACTThe Internet of Things has emerged as a significant and influential technology in modern times. IoT presents solutions to reduce the need for human intervention and emphasizes task automation. According to a Cisco report, there were over 14.7 billion IoT devices in 2023. However, as the number of devices and users utilizing this technology grows, so does the potential for security breaches and intrusions. For instance, insecure IoT devices, such as smart home appliances or industrial sensors, can be vulnerable to hacking attempts. Hackers might exploit these vulnerabilities to gain unauthorized access to sensitive data or even control the devices remotely. To address and prevent this issue, this work proposes integrating intrusion detection systems (IDSs) with an artificial neural network (ANN) and a salp swarm algorithm (SSA) to enhance intrusion detection in an IoT environment. The SSA functions as an optimization algorithm that selects optimal networks for the multilayer perceptron (MLP). The proposed approach has been evaluated using three novel benchmarks: Edge‐IIoTset, WUSTL‐IIOT‐2021, and IoTID20. Additionally, various experiments have been conducted to assess the effectiveness of the proposed approach. Additionally, a comparison is made between the proposed approach and several approaches from the literature, particularly SVM combined with various metaheuristic algorithms. Then, identify the most crucial features for each dataset to improve detection performance. The SSA‐MLP outperforms the other algorithms with 88.241%, 93.610%, and 97.698% for Edge‐IIoTset, IoTID20, and WUSTL, respectively.
The 2022 Qatar World Cup created massive global attention and generated widespread discussions on different social media platforms, including the X platform. The event was the subject of intense debate after Qatar was announced as the host. Opinions were divided, with supporters and critics weighing in based on political, ethical, cultural, and social considerations. Public sentiments evolved throughout three key phases-before, during, and after the event-shaped by numerous factors. This study aims to analyze these sentiments during these three stages based on a novel hybrid evolutionary approach. Three versions for each stage were produced by applying pre-trained word embeddings with 100 and 400 features and sentiment features combined with word embeddings. In total, nine different versions of datasets were employed to examine the proposed approach. Furthermore, five different metaheuristic algorithms were applied: the multi-verse optimizer (MVO), the genetic algorithm (GA), the particle swarm optimization (PSO), the salp swarm algorithm (SSA), and the whale optimization algorithm (WOA). The five metaheuristic algorithms were combined with the feature selection-support vector machine (FS-SVM) and weighting-support vector machine (WSVM) to examine the newly created dataset versions. The results reveal that people’s perspectives shifted from negative before the event to positive during and after the event. Moreover, a comparison of the proposed MVO-WSVM and MVO-SVM-Fs approaches with other metaheuristic algorithms showed the superior accuracy of the proposed approaches in sentiment prediction.
In the digital age, spam detection remains a critical challenge, particularly in educational environments where the nature and volume of communication differ significantly from other domains. The increasing frequency and sophistication of email-based hacking attacks pose significant security threats to academic institutions. This paper presents an approach to detecting spam in an educational context using machine learning algorithms. By leveraging a dataset collected from an academic institution, we highlight the unique characteristics of academic communication and their implications for spam detection. Our methodology involves a comprehensive feature selection technique to identify the most relevant attributes for effective spam filtering. To address the imbalance in the dataset, class balancing techniques are employed, ensuring the machine learning models are trained on a representative distribution of spam and non-spam messages. The impact of feature selection on the classification of spam emails is investigated using five machine learning algorithms. Evaluation of the classifiers determines the most effective approach for classification of spam emails in the educational context. The experimental results demonstrate that our approach achieves remarkable accuracy, with the Random Forest classifier reaching up to $\mathbf{9 9. 9 \%}$ accuracy. This high level of precision underscores the potential of tailored machine learning solutions in educational spam detection and sets a benchmark for future research in this area.
Online reviews are important information that customers seek when deciding to buy products or services. Also, organizations benefit from these reviews as essential feedback for their products or services. Such information required reliability, especially during the Covid-19 pandemic which showed a massive increase in online reviews due to quarantine and sitting at home. Not only the number of reviews was boosted but also the context and preferences during the pandemic. Therefore, spam reviewers reflect on these changes and improve their deception technique. Spam reviews usually consist of misleading, fake, or fraudulent reviews that tend to deceive customers for the purpose of making money or causing harm to other competitors. Hence, this work presents a Weighted Support Vector Machine (WSVM) and Harris Hawks Optimization (HHO) for spam review detection. The HHO works as an algorithm for optimizing hyperparameters and feature weighting. Three different language corpora have been used as datasets, namely English, Spanish, and Arabic in order to solve the multilingual problem in spam reviews. Moreover, pre-trained word embedding (BERT) has been applied alongside three-word representation methods (NGram-3, TFIDF, and One-hot encoding). Four experiments have been conducted, each focused on solving and demonstrating different aspects. In all experiments, the proposed approach showed excellent results compared with other state-of-the-art algorithms. In other words, the WSVM-HHO achieved an accuracy of 88.163%, 71.913%, 89.565%, and 84.270%, for English, Spanish, Arabic, and Multilingual datasets, respectively. Further, a deep analysis has been conducted to investigate the context of reviews before and after the COVID-19 situation. In addition, it has been generated to create a new dataset with statistical features and merge its previous textual features for improving detection performance.
Deep neural networks (DNNs) are currently being deployed as machine learning technology in a wide range of important real-world applications. DNNs consist of a huge number of parameters that require millions of floating-point operations (FLOPs) to be executed both in learning and prediction modes. A more effective method is to implement DNNs in a cloud computing system equipped with centralized servers and data storage sub-systems with high-speed and high-performance computing capabilities. This paper presents an up-to-date survey on current state-of-the-art deployed DNNs for cloud computing. Various DNN complexities associated with different architectures are presented and discussed alongside the necessities of using cloud computing. We also present an extensive overview of different cloud computing platforms for the deployment of DNNs and discuss them in detail. Moreover, DNN applications already deployed in cloud computing systems are reviewed to demonstrate the advantages of using cloud computing for DNNs. The paper emphasizes the challenges of deploying DNNs in cloud computing systems and provides guidance on enhancing current and new deployments.
Online media has an increasing presence on the restaurants’ activities through social media websites, coinciding with an increase in customers’ reviews of these restaurants. These reviews become the main source of information for both customers and decision-makers in this field. Any customer who is seeking such places will check their reviews first, which usually affect their final choice. In addition, customers’ experiences can be enhanced by utilizing other customers’ suggestions. Consequently, customers’ reviews can influence the success of restaurant business since it is considered the final judgment of the overall quality of any restaurant. Thus, decision-makers need to analyze their customers’ underlying sentiments in order to meet their expectations and improve the restaurants’ services, in terms of food quality, ambiance, price range, and customer service. The number of reviews available for various products and services has dramatically increased these days and so has the need for automated methods to collect and analyze these reviews. Sentiment Analysis (SA) is a field of machine learning that helps analyze and predict the sentiments underlying these reviews. Usually, SA for customers’ reviews face imbalanced datasets challenge, as the majority of these sentiments fall into supporters or resistors of the product or service. This work proposes a hybrid approach by combining the Support Vector Machine (SVM) algorithm with Particle Swarm Optimization (PSO) and different oversampling techniques to handle the imbalanced data problem. SVM is applied as a machine learning classification technique to predict the sentiments of reviews by optimizing the dataset, which contains different reviews of several restaurants in Jordan. Data were collected from Jeeran, a well-known social network for Arabic reviews. A PSO technique is used to optimize the weights of the features, as well as four different oversampling techniques, namely, the Synthetic Minority Oversampling Technique (SMOTE), SVM-SMOTE, Adaptive Synthetic Sampling (ADASYN) and borderline-SMOTE were examined to produce an optimized dataset and solve the imbalanced problem of the dataset. This study shows that the proposed PSO-SVM approach produces the best results compared to different classification techniques in terms of accuracy, F-measure, G-mean and Area Under the Curve (AUC), for different versions of the datasets.
This paper introduces and tests a novel machine learning approach to detect Android malware. The proposed approach is composed of Support Vector Machine (SVM) classifier and Harris Hawks Optimization (HHO) algorithm. More specifically, the role of HHO algorithm is to optimize SVM classifier hyperparameters while the SVM performs the classification of malware based on the best-chosen model, as well as producing the optimal solution for weighting the features. The effectiveness of the proposed approach and the ability to increase detection performance are demonstrated by scientific testing using CICMalAnal2017 sampled datasets. We test our method and its robustness on five sampled datasets and achieved the best results in most datasets and measures when compared with other approaches. We also illustrate the ability of the proposed approach to measure the significance of each feature. In addition, we provide deep analysis of possible relationships between weighted features and the type of malware attack. The results show that the proposed approach outperforms the other metaheuristic algorithms and state-of-art classifiers.
During the recent COVID-19 pandemic, people were forced to stay at home to protect their own and others’ lives. As a result, remote technology is being considered more in all aspects of life. One important example of this is online reviews, where the number of reviews increased promptly in the last two years according to Statista and Rize reports. People started to depend more on these reviews as a result of the mandatory physical distance employed in all countries. With no one speaking to about products and services feedback. Reading and posting online reviews becomes an important part of discussion and decision-making, especially for individuals and organizations. However, the growth of online reviews usage also provoked an increase in spam reviews. Spam reviews can be identified as fraud, malicious and fake reviews written for the purpose of profit or publicity. A number of spam detection methods have been proposed to solve this problem. As part of this study, we outline the concepts and detection methods of spam reviews, along with their implications in the environment of online reviews. The study addresses all the spam reviews detection studies for the years 2020 and 2021. In other words, we analyze and examine all works presented during the COVID-19 situation. Then, highlight the differences between the works before and after the pandemic in terms of reviews behavior and research findings. Furthermore, nine different detection approaches have been classified in order to investigate their specific advantages, limitations, and ways to improve their performance. Additionally, a literature analysis, discussion, and future directions were also presented.
Renewable energy sources are considered ubiquitous and drive the energy revolution. Energy producers suffer from inconsistent electricity generation. They often struggled with the unpredictability of the weather. Thus, making it challenging to balance supply and demand. Technologies like artificial intelligence (AI) and machine learning are effective ways to forecast, distribute, and manage renewable photovoltaic (PV) solar supplies. AI will make the energy forecasting system more connected, intelligent, reliable, and sustainable. AI can innovate how energy is used and help find solutions for decarbonizing energy systems. There are potential advantages to total energy forecasting. AI can support the growth and integration of PV solar energy. The article’s main objective is to use AI to forecast the output consumed power of the Yarmouk University PV solar system in Jordan. The total actual yield is 5548.96 MW h, and the performance ratio (PR) is 95.73%. Many techniques are used to predict the consumed solar power. The random forest model obtains the best results of root mean squared error and mean absolute error are 172.07 and 68.7, respectively. This accurate prediction allows for the maximum use of solar power and the minimal use of grid power. This work guides the operators to learn trends embedded in Yarmouk University’s historical data. These understood trends can be used to predict the consumption of solar power output. Thus, the control system and grid operators have advanced knowledge of the expected consumption of solar power at each hour of the day.
J. Merelo合作论文数Dept. of Computer Technology and Architecture;Universidad de Granada2