Intrusion detection systems (IDS) play a vital role in protecting computer networks from malicious activities. Dimensionality reduction techniques are commonly employed to enhance the effectiveness and accuracy of machine learning based IDS. In this study, we proposed an effective dimensionality reduction technique called feature importance-based autoencoder (FI-AE) for intrusion detection systems. Our proposed approach encompasses several key components. First, we introduce a novel feature importance method known as one-versus-all feature importance (OVA), which utilizes a random forest algorithm. Next, we train an autoencoder model using a weighted loss function that takes into account the feature importance values obtained through the OVA method. Finally, we utilized the trained autoencoder to reduce the number of features in the benchmark datasets, followed by the application of a random forest classifier to the reduced datasets. We tested our proposed model using three well-known datasets, namely NSL-KDD, UNSW-NB15, and CIC-IDS2017. The experiments revealed that the random forest classifier, combined with our proposed model, outperformed previous dimensionality reduction techniques in terms of accuracy and F1-score.
This article introduces a novel hybrid optimization method, the Young’s Double-Slit Experiment-Differential Evolution (YDSE-DE) technique, which is based on the integration of the YDSE approach with the robust capabilities of Differential Evolution algorithms. This integration enables the precise estimation of parameters in photovoltaic (PV) models, enhancing the modeling and simulation of solar energy systems. Employing a comparative assessment strategy, the authors demonstrate that the proposed YDSE-DE algorithm outperforms existing methods such as Ant Lion Optimizer (ALO) and Sooty Tern Optimization Algorithm (STOA) in terms of accuracy and computational efficiency. The effectiveness of this method is confirmed through rigorous testing, which shows improvements in the root mean square error (RMSE) metrics by significant margins across single-diode, double-diode, and three-diode PV models. RMSE values obtained using the YDSE-DE algorithm for the single-diode, double-diode, and three-diode models are documented as 0.0059117, 0.0015218, and 0.0018409, respectively. The present study underscores the considerable potential of the YDSE-DE algorithm in augmenting the precision and effectiveness of photovoltaic system simulation and design. The results develop and enhance the existing methodologies for PV parameter estimation and can be applied to optimize other complex systems requiring reliable and precise simulation. The novelty and scientific contribution of this work lie in the methodical synthesis of physical principles from optics with evolutionary computation, presenting a significant advancement in the field of renewable energy technology.
Coverless image steganography conceals information without modifying the carrier image, addressing vulnerabilities in traditional methods. However, existing approaches often require transmitting metadata, raising suspicion and security risks. To overcome these limitations, we propose Coverless Hybrid Image Data Encryption (C-HIDE), a robust steganographic method integrating Advanced Encryption Standard (AES) for data confidentiality and Elliptic Curve Cryptography (ECC) for secure key exchange. The system ensures secure transmission without altering cover images, making embedded data harder to detect. C-HIDE eliminates metadata transmission by enabling both sender and receiver to independently generate synchronized coverless image datasets (CIDs) using random seeds. Encrypted secret data is mapped to images whose hash sequences correspond to segments of the message, with Speeded-Up Robust Features (SURF) ensuring reliable image matching. At the receiver’s end, ECC-decrypted AES keys recover the original message while SURF retrieves relevant images. Experimental results demonstrate that C-HIDE achieves an embedding capacity of 574 bits per image, significantly surpassing DCT (256 bits) and DWT (128 bits) techniques. The system maintains 98.5% accuracy under attacks such as noise addition, cropping, and geometric transformations. Furthermore, it enhances security by eliminating metadata transmission, achieving a zero additional information ratio, unlike conventional methods requiring up to 25% extra data. By integrating encryption, minimizing detection, and removing metadata transmission, C-HIDE provides a secure, efficient, and scalable solution for covert communication in real-world applications.
Audio forensics plays a major role in the investigation and analysis of audio recordings for legal and security purposes. The advent of audio fake attacks using speech combined with scene-manipulated audio represents a sophisticated challenge in fake audio detection. Fake audio detection, a critical technology in modern digital security, addresses the growing threat of manipulated audio content across various applications, including media, legal evidence, and cybersecurity. This research proposes a novel transfer learning approach for fake audio detection. We utilized a benchmark dataset, SceneFake, that contains 12,668 audio signal files for both real and fake scenes. We propose a novel transfer learning method, which initially extracts mel-frequency cepstral coefficients (MFCC) and then class prediction probability value features. The newly generated transfer features set by the proposed MfC-RF (MFCC-Random Forest) are utilized for further experiments. Results expressed that using the MfC-RF features random forest method outperformed existing state-of-the-art methods with a high-performance measure accuracy of 0.98. We have tuned hyperparameters of applied machine learning approaches, and cross-validation is applied to validate performance results. In addition, the complexity of the computation is measured. The proposed research aims to enhance the accuracy measure, and efficiency of identifying manipulated audio content, thereby contributing to the integrity and reliability of digital communications.
The enormous potential for healthcare data management to deliver efficient and economical patient care has gained increasing attention in recent years. The Internet of Things (IoT), now reaching full maturity due to the rapid advancement of communication technologies, is expanding rapidly as large volumes of data are generated and transmitted. This growth introduces complex performance demands for controlling devices deployed globally. Most existing IoT healthcare infrastructures are highly centralized, leading to issues such as single points of failure and increased susceptibility to cyberattacks. To ensure secure and verified data interoperability, the blockchain paradigm introduces promising mechanisms for enhancing privacy, security, and decentralization. In this work, a blockchain-enabled edge-IoT framework is presented for securing smart healthcare management by minimizing network latency and computational complexity. The framework ensures data integrity and enables secure, immutable logging of patient information across distributed devices. It also supports real-time monitoring and control through IoT smart devices, reducing risks in critical healthcare scenarios. Formal security validation was conducted using the AVISPA tool to assess the authentication protocol. Simulation and performance analysis using iFogSim and IoTSim-Edge demonstrated a 94.6% reduction in execution time and up to 15% energy savings, validating the framework's scalability and effectiveness compared to conventional cloud-based models.
Integrating Artificial Intelligence into the Internet of Things is reworking industries by making them more clever, effective, and agile. AI techniques, like system mastering and deep gaining knowledge, similarly empower IoT by processing massive volumes of sensor records to derive higher insights for selection-making and optimization. However, this collaboration also brings up numerous demanding situations, including extended computational necessities, electricity inefficiencies, and privacy issues. The manuscript deals with the transformative capacity of AI in IoT systems, their blessings and disadvantages, and further improvement perspectives that might be orientated toward sustainability.
Estimating parameters in solar cell models is crucial for simulating and designing photovoltaic systems. The single-diode, double-diode, and three-diode models represent these systems. Parameter estimation can be viewed as an optimization problem to minimize the difference between measured and estimated data. This study presents PV parameter estimation using the enhanced Sinh Cosh Optimizer (I_SCHO), incorporating trigonometric operators from the Sine Cosine Algorithm (SCA). This integration improves the algorithm’s ability to navigate complex search spaces, avoid local optima, and expedite convergence. Assessment criteria include runtime, convergence behaviour, minimum RMSE, and system reliability measured by SD. Results show that I_SCHO consistently delivers superior accuracy and reliability compared to other methods. Experiments were conducted on five solar cells: RTC France, Photowatt-PWP201, Kyocera KC200GT, Ultra 85-P, and STM6-40/36 module. The study also includes a comparative analysis using state-of-the-art algorithms, demonstrating I_SCHO’s efficiency through RMSE, Power Voltage (P-V) and Current Voltage (I-V) curves.
IntroductionThe mountain gazelle (Gazella gazella) is a native species to the Middle East and has experienced a notable population decline due to human-induced habitat loss and fragmentation. In Saudi Arabia, the current status and distribution of this species remain poorly understood, necessitating data-driven conservation assessments.Methods, Results, and DiscussionThis study combined recent occurrence records with remote sensing and GIS-based environmental variables to model suitable habitats for the mountain gazelle using the MaxEnt algorithm. Key predictors included vegetation indices, land cover types, and elevation. The results identified core habitat areas in the western and southwestern regions, some of which fall outside current protected zones. These findings underscore the importance of expanding conservation areas and demonstrate how spatial modeling supports effective wildlife management in arid environments.
One of today`s most quickly developing technologies is the Internet of Things (IoT). There are now more threats and risks to its security than ever before. In order to tackle present and future IoT issues, machine learning is an effective technology that can be used to identify risks and threats in intelligent systems. In today’s world, credit card is the most popular payment mode for both online and offline. Consumers rely on online shopping and online bill payment, which cases of fraud associated with it are also increasing. With the developments in the communication channels, fraud is spreading all over the world resulting in huge financial losses. Fraud detection is the essential tool and probably the best way to stop fraud types. There is a technique of finding an optimal solution for a problem and implicitly generate the results using machine learning and genetic algorithm. The aim is to develop a model to detect fraudulent transactions and improves a credit card fraud detection solution with some machine learning algorithms such as GA, DT, LR, KNN, SVC, and ANN based on the RUST and SMOTE techniques. The experiments are conducted on the BCCFDD and DCCCD datasets to analyze the model using the dimension reduction transformers (T-SNE, PCA, and Truncated SVD). The performance of the classification model analyzed in terms of confusion matrix, the model ROC curve analysis, and accuracy. The evaluation finding is analyzed and compared. As proof of concept, a Credit Card Fraud Detection System (CCFDS) is developed to detect the credit card fraud based on the principles of the GA and showed the effectiveness of proposed approach. This algorithm is an optimization technique and evolutionary search based on the principles of genetic and natural selection, heuristic used to solve high complexity computational problems.
The article describes a new method for malware classification, based on a Machine Learning (ML) model architecture specifically designed for malware detection, enabling real-time and accurate malware identification. Using an innovative feature dimensionality reduction technique called the Interpolation-based Feature Dimensionality Reduction Technique (IFDRT), the authors have significantly reduced the feature space while retaining critical information necessary for malware classification. This technique optimizes the model’s performance and reduces computational requirements. The proposed method is demonstrated by applying it to the BODMAS malware dataset, which contains 57,293 malware samples and 77,142 benign samples, each with a 2381-feature vector. Through the IFDRT method, the dataset is transformed, reducing the number of features while maintaining essential data for accurate classification. The evaluation results show outstanding performance, with an F1 score of 0.984 and a high accuracy of 98.5% using only two reduced features. This demonstrates the method’s ability to classify malware samples accurately while minimizing processing time. The method allows for improving computational efficiency by reducing the feature space, which decreases the memory and time requirements for training and prediction. The new method’s effectiveness is confirmed by the calculations, which indicate significant improvements in malware classification accuracy and efficiency. The research results enhance existing malware detection techniques and can be applied in various cybersecurity applications, including real-time malware detection on resource-constrained devices. Novelty and scientific contribution lie in the development of the IFDRT method, which provides a robust and efficient solution for feature reduction in ML-based malware classification, paving the way for more effective and scalable cybersecurity measures.
Ongoing importance remains placed on parameter estimation in photovoltaic (PV) system design and simulation. Among the diode-based models utilized frequently are those consisting of a single diode, a double diode, and three diode. Minimizing the discrepancy between the calculated and measured values is often the primary aim when estimating parameters for these models. In recent years, the estimation of parameters in conventional PV models has been approached using various numerical, analytical, and hybrid techniques. However, these methods could be made more complex to produce credible results promptly and precisely. This article discusses the three fundamental PV models. A contemporary optimization algorithm, the Crayfish Optimisation algorithm (COA), is employed to extract the parameters for PV models. This includes the single-diode, double-diode, and three-diode models. An assessment uses to contrast the PV models. COA outperforms the subsequent competing algorithms, as demonstrated by the experimental results: Hunger Games Search, SOA, STOA, Synergistic Mimic Algorithm (SMA), TURBULAS Swarm Algorithm (TSA), and LAPOP (Lightning Attachment Procedure Optimisation) are all examples of optimization algorithms, along with HHO, HBO, LIPO, SOA, and STOA, respectively. By the negligible difference between measured and calculated data, this comparison illustrates that the parameters extracted by COA are optimal. As determined by the proposed COA algorithm, 0.00085477, 0.0010313, and 0.00092288 are the optimal RMSE values for SDM, DDM, and TDM.
In the field of data security, biometric security is a significant emerging concern. The multimodal biometrics system with enhanced accuracy and detection rate for smart environments is still a significant challenge. The fusion of an electrocardiogram (ECG) signal with a fingerprint is an effective multimodal recognition system. In this work, unimodal and multimodal biometric systems using Convolutional Neural Network (CNN) are conducted and compared with traditional methods using different levels of fusion of fingerprint and ECG signal. This study is concerned with the evaluation of the effectiveness of proposed parallel and sequential multimodal biometric systems with various feature extraction and classification methods. Additionally, the performance of unimodal biometrics of ECG and fingerprint utilizing deep learning and traditional classification technique is examined. The suggested biometric systems were evaluated utilizing ECG (MIT-BIH) and fingerprint (FVC2004) databases. Additional tests are conducted to examine the suggested models with:1) virtual dataset without augmentation (ODB) and 2) virtual dataset with augmentation (VDB). The findings show that the optimum performance of the parallel multimodal achieved 0.96 Area Under the ROC Curve (AUC) and sequential multimodal achieved 0.99 AUC, in comparison to unimodal biometrics which achieved 0.87 and 0.99 AUCs, for the fingerprint and ECG biometrics, respectively. The overall performance of the proposed multimodal biometrics outperformed unimodal biometrics using CNN. Moreover, the performance of the suggested CNN model for ECG signal and sequential multimodal system based on neural network outperformed other systems. Lastly, the performance of the proposed systems is compared with previously existing works.
Recently, low Earth orbit (LEO) satellites have emerged as key players in space information network (SIN) due to their ability to provide global coverage. However, they remain susceptible to threats such as denial of service (DoS), man-in-the-middle (MITM), and spoofing attacks. In this paper, we propose a cross-layer security framework (CLSF) to address these vulnerabilities. Our approach begins by employing a physically unclonable function (PUF) at the upper layer to establish mutual authentication between legitimate satellites and ground stations, while also securely exchanging frequency seeds for the next phase. Following this, dynamic seed frequency hopping (DSFH) is applied at the physical layer to counter DoS, MITM, and spoofing attacks. Additionally, the frequency transitions of malicious satellites are modeled using a Markov chain. Our results demonstrate that the proposed CLSF, which integrates PUF and DSFH, delivers strong security performance.
Accurate parameter estimation in photovoltaic (PV) system design and simulationis essential for optimizing performance. Traditional numerical, analytical,and hybrid methods often fail to deliver quick and precise results. This researchintroduces enhancements to three fundamental PV models (single-diode, doublediode,and three-diode) and employs the Improved Mountain Gazelle Optimizer(i_MGO) algorithm for parameter extraction. The innovative objective functionproposed aids in minimizing the discrepancy between calculated and measuredvalues. Rigorous experimental validation demonstrates that i_MGO outperformsexisting algorithms, achieving optimal parameter values with minimal rootmean square error (RMSE). The experimental findings illustrate that i_MGOperforms better than the following competing algorithms: Improved MountainGazelle Optimizer (i_MGO), Harris Hawks Optimization (HHO), LightningAttachment Procedure, Optimization Algorithm (LAPO), Sine Cosine Algorithm(SCA), Grey Wolf Optimizer (GWO), African Vultures Optimization Algorithm(AVOA), Hippopotamus Optimization Algorithm (HO), Electric Eel ForagingOptimization (EEFO), Synergistic Swarm Optimization Algorithm (SSOA),Coati Optimization Algorithm (COA), Gazelle Optimization Algorithm (GOA).This comparison demonstrates that the parameters extracted by i_MGO areoptimal, as the discrepancy between measured and calculated data is minimal.The optimal RMSE values for SDM, DDM, and TDM, as determined bythe proposed i_MGO algorithm, are 0.00081373, 0.00073908, and 0.00092975,correspondingly.
The quality of the air we breathe during the courses of our daily lives has a significant impact on our health and well-being as individuals.Unfortunately, personal air quality measurement remains challenging.In this study, we investigate the use of first-person photos for the prediction of air quality.The main idea is to harness the power of a generalized stacking approach and the importance of haze features extracted from first-person images to create an efficient new stacking model called AirStackNet for air pollution prediction.AirStackNet consists of two layers and four regression models, where the first layer generates meta-data from Light Gradient Boosting Machine (Light-GBM), Extreme Gradient Boosting Regression (XGBoost) and CatBoost Regression (CatBoost), whereas the second layer computes the final prediction from the meta-data of the first layer using Extra Tree Regression (ET).The performance of the proposed AirStackNet model is validated using public Personal Air Quality Dataset (PAQD).Our experiments are evaluated using Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Coefficient of Determination (R 2 ), Mean Squared Error (MSE), Root Mean Squared Logarithmic Error (RMSLE), and Mean Absolute Percentage Error (MAPE).Experimental Results indicate that the proposed AirStackNet model not only can effectively improve air pollution prediction performance by overcoming the Bias-Variance tradeoff, but also outperforms baseline and state of the art models.
Objective: Predicting the ability of a breast cancer patient to survive was a difficult research problem for many scholars. Since the early dates of the relevant research, significant progress has been recorded in many related areas. For example, with pioneering biomedical technologies, credits to low-cost computer hardware and software, high-quality data is gathered and stored automatically, and lastly, with better analytical methods, that massive data is processed efficiently and effectively. Therefore, the objective of this document is to submit a report on a research project in which we have benefited from the technological developments available to develop predictive models of breast cancer and whether it exists or not. Methods and materials: artificial neural network, support vector machine, decision trees, naïve bayes, and random forest algorithms are used along with the most common statistical method (logistic regression) to build prediction models using a large data set. We also used the Holdout method. To avoid the unbalanced nature of the classes, the parameters of the performance evaluation are predefined. Results: The results show that the Decision Tree (DT) is the top predictor with 89.1% accuracy on the holdout sample, surpassing all prediction accuracy reported in the literature; Artificial Neural Networks (ANN) came out to be the second with 88.9% accuracy; Naïve Bayes (NB) came out to be the third with 83.3% accuracy, Support Vector Machines (SVM) came out to be the fourth with 83.2% accuracy, and the Random Forest (RF) models came out to be the lowest of the five with 71.2% accuracy. Conclusion: A comparative study of multiple predictive models for breast cancer survival using a large set of data and 5-fold cross-validation gave us an insight into the relative ability to predict different data extraction methods. After analyzing the data, we have reached this conclusion: the model is able to help those who need it by predicting whether they have breast cancer or not. Furthermore, the proposed framework is valuable tool in cancer research and clinical practice.
Detecting abnormal sounds is crucial in various fields like security, environmental monitoring, and health surveillance. This proposal presents a method for identifying machine anomalous sounds using machine learning and unsupervised learning techniques. Our methodology includes data preprocessing, feature extraction, and training a deep neural network using a pre-trained AutoML (Automated Machine Learning) service to differentiate between normal and anomalous sounds. The system strives for real-time, high-accuracy detection, making it valuable across applications. This research addresses current limitations in sound anomaly detection, aiming to significantly impact the field. Our experiments demonstrate that our method surpasses state-of-the-art techniques, including pre-trained learning and self-supervised classification, in both overall anomaly detection performance and stability on the DCASE 2020 Challenge Task2 dataset.
Blockchain technology has gained widespread adoption in recent years due to its ability to enable secure and transparent record-keeping and data transfer. A critical aspect of blockchain technology is the use of consensus algorithms, which allow distributed nodes in the network to agree on the state of the blockchain. In this review paper, we examine various consensus algorithms that are used in blockchain systems, including proof-of-work, proof-of-stake, and hybrid approaches. We go over the trade-offs and factors to think about when choosing a consensus algorithm, such as energy efficiency, decentralization, and security. We also look at the strengths and weaknesses of each algorithm as well as their potential impact on the scalability and adoption of blockchain technology.
This paper promotes better life quality and lifestyle for patients. We attain this goal by creating a mobile application that analyses patient's medical records, such as diabetes, hypertension, and chronic kidney diseases. Then, we implement the system to diagnose patients with chronic conditions using machine learning techniques. Machine learning classifiers are used in this paper to decide whether a person has any chronic diseases. The investigated diseases are hypertension, diabetes, and chronic kidney disease. Four datasets were used to build the classifying models. Orange3 from Anaconda-Navigator, a data mining tool, was used to test machine learning algorithms. The study findings revealed the superiority of the tree algorithm with 100% accuracy for hypertension; it was the highest outcome for both males and females using Orange3. The highest precision, which is 100%, is observed by SVM, k-NN, decision trees, logistic regression, and CART for hypertension males’ data collection. In comparison, the highest precision is 100% in SVM, MLP, decision tree, random forest, logistic regression, and CART for the female dataset. We conclude that the two datasets for the same diseases share mostly the same algorithm accuracy. For kidneys, the Random Forest algorithm produced 100% accuracy, which is the highest value among other algorithms. For diabetes, neural networks have attested the best accuracy. It was 76.3%, yet the accuracy increased slightly as the kNN algorithm showed 83% accuracy.
Diabetes, in all of its types, costs countries of all income levels unacceptably enormous personal, societal, and economic expenses. To predict type-2 diabetes, this work aimed to develop an analytical predictive model based on machine learning techniques and a web-based personalized diabetes monitoring system. The history of a patient will be collected and ready for analysis purposes based on machine learning techniques by continuously monitoring the patient's vital data. A diabetes monitoring system is proposed by utilizing the patient's QR card that allows the patients and doctors to be connected to the Internet of Things. So, they can deliver real-time information (such as insulin records) about their health status and can visit different healthcare institutions. The proposed system can help doctors to make data-driven decisions and enhance patients' treatment. Several machine learning algorithms that are Decision Tree, Support Vector Classifier, Random Forest, Gradient Boosting, Multi-layer Perceptron's, Artificial Neural Network, k-Nearest Neighbors, Logistic Regression, and Naive Bayes are used. The proposed analytical model is evaluated based on two different datasets a synthetic dataset and PIMA Diabetes Dataset. The performance of the classification models was analyzed in terms of accuracy, recall, and precision based on the cross-validation strategy. The findings show that the ANN model has better prediction accuracy than other models. The evaluation findings are analyzed and compared with other existing models. The system has dashboard graphs displaying the number of patients in Saudi Arabia cities. It also contains visualized graphs that include more detailed classifications for patients' states (Normal, Pre-Diabetes, and Diabetes).