The dynamic connectivity and functionality of sensors has revolutionized remote monitoring applications thanks to the combination of IoT and wireless sensor networks (WSNs). Wearable wireless medical sensor nodes allow continuous monitoring by amassing physiological data, which is very useful in healthcare applications. These text data are then sent to doctors via IoT devices so they can make an accurate diagnosis as soon as possible. However, the transmission of medical text data is extremely vulnerable to security and privacy assaults due to the open nature of the underlying communication medium. Therefore, a certificate-less aggregation-based signature system has been proposed as a solution to the issue by using elliptic curve public key cryptography (ECC) which allows for a highly effective technique. The cost of computing has been reduced by 93% due to the incorporation of aggregation technology. The communication cost is 400 bits which is a significant reduction when compared with its counterparts. The results of the security analysis show that the scheme is robust against forging, tampering, and man-in-the-middle attacks. The primary innovation is that the time required for signature verification can be reduced by using point addition and aggregation. In addition, it does away with the reliance on a centralized medical server in order to do verification. By taking a distributed approach, it is able to fully preserve user privacy, proving its superiority.
The Internet of Things (IoT) is extensively used in modern-day life, such as in smart homes, intelligent transportation, etc. However, the present security measures cannot fully protect the IoT due to its vulnerability to malicious assaults. Intrusion detection can protect IoT devices from the most harmful attacks as a security tool. Nevertheless, the time and detection efficiencies of conventional intrusion detection methods need to be more accurate. The main contribution of this paper is to develop a simple as well as intelligent security framework for protecting IoT from cyber-attacks. For this purpose, a combination of Decisive Red Fox (DRF) Optimization and Descriptive Back Propagated Radial Basis Function (DBRF) classification are developed in the proposed work. The novelty of this work is, a recently developed DRF optimization methodology incorporated with the machine learning algorithm is utilized for maximizing the security level of IoT systems. First, the data preprocessing and normalization operations are performed to generate the balanced IoT dataset for improving the detection accuracy of classification. Then, the DRF optimization algorithm is applied to optimally tune the features required for accurate intrusion detection and classification. It also supports increasing the training speed and reducing the error rate of the classifier. Moreover, the DBRF classification model is deployed to categorize the normal and attacking data flows using optimized features. Here, the proposed DRF-DBRF security model's performance is validated and tested using five different and popular IoT benchmarking datasets. Finally, the results are compared with the previous anomaly detection approaches by using various evaluation parameters.
Cancer is a large group of diseases with one thing in common: they all happen when normal cells become cancerous cells that multiply and spread. Almost ten million people die from cancer every year, making it the second-highest cause of death worldwide. Although some highly skilled medical professionals have helped to lower this number, the complexity of identifying specific cancer types and determining appropriate treatments remains a significant challenge. Various cancer forms require different types of treatment; first, the expert must recognize the stage, behavior, growth of cells, and location, and then suggest treatment to the concerned patient. This study specifically addresses the challenge of recognizing the form of cancer among various cancer types and providing alternative treatment options. To resolve this issue, we extend the concept of a two-way decision-making approach to a three-way decision-making approach, which helps deal with uncertainties in selecting decision-makers. The conditional probability is a key factor in the three-way decision-making approach; we use the extended TODIM approach for this to overcome the limitations of the traditional TODIM approach. For this, we develop a series of Hamacher aggregation operators using double-hesitant hierarchy linguistic term numbers (DHHLTNs), which include two types of hesitation: the first hesitant hierarchy linguistic term (FHHLT) and the second hesitant hierarchy linguistic term (SHHLT), making them better at expressing ambiguity and softness. We also develop score and accuracy functions to rank DHHLTNs and define a distance equation to determine the distance between two DHHLTNs. The decision-making process becomes complex due to unknown weight vectors. The entropy measure approach is introduced to locate unknown weight vectors. This paper contributes to the existing literature by demonstrating the effectiveness of the proposed technique in identifying critical types of cancer, showing its feasibility and efficacy compared to other MCDM techniques. The proposed technique shows highly accurate and reliable results in identifying and prioritizing critical forms of cancer, indicating that it is more precise and provides improved expert information.
This article examines the operational functionality of intelligent transport systems to enhance smart cities by reducing traffic congestion. Given the increasing populations of smart cities, there is a growing demand for public transit systems to address the issue of traffic congestion. Therefore, the suggested system is developed using a few parametric design models, which combine point-to-point protocol and mode control optimization. The multi-objective parametric design for a smart transportation system is conducted using min-max functions to minimize the waiting time period for end users. Furthermore, customers are given the option to utilize a line following mechanism that offers suitable connectivity, along with independent identification and revitalize functions. The predicted model effectively eliminates the delay produced by transportation devices when positioning units are involved, ensuring that individual messages are delivered without any interruptions. In order to evaluate the results of the proposed system model, four different scenarios were examined. A comparison analysis revealed that the suggested method achieves a suitable directional flow for 96% of smart transport units. Additionally, it reduces delays and waiting periods by 2% and 6% respectively, while increasing energy consumption by 29%.
A gastrointestinal disease is a group of cancers which mainly affects the digestive system, along with the stomach, small intestine, oesophagus, rectum, and colon. Accurate classification and earlier diagnosis of this cancer are crucial for better patient outcomes. Deep learning (DL) algorithm, especially convolutional neural network (CNN), is trained to categorize endoscopic images of gastrointestinal tissue as either benign or malignant. Gastrointestinal cancer (GC) classification with DL is the process of using artificial intelligence (AI), especially the DL algorithm, to categorize endoscopic images of gastric tissue as benign or malignant. It could help clinicians to identify the earliest symptoms of cancer and make treatment decisions, resulting in improved patient outcomes. The study designs a new gastrointestinal disease Detection and Classification using Hybrid Rice Optimization with Deep Learning (GDDC-HRODL) model. The presented GDDC-HRODL model intends to classify the medical images for GC. To achieve this, the GDDC-HRODL technique initially preprocesses the input data to improve image quality. In addition, the presented GDDC-HRODL algorithm employs the HybridNet model to produce feature vectors and the hyperparameter tuning process takes place using the HRO algorithm. For GC classification purposes, the GDDC-HRODL technique uses an attention-based long short-term memory (ALSTM) model and its hyperparameters can be selected by the ant lion optimization (ALO) algorithm. The design of hyperparameter tuning processes helps to accomplish enhanced GC classification performance. The experimental analysis of the GDDC-HRODL algorithm on the medical dataset demonstrates its betterment in the GC classification process.
Red palm weevil (RPW) is a pest that can cause severe damage to plantations and affects palm trees. Classical approaches to detection depend on visual analysis, which is inaccurate and time-consuming. Hence, deep learning techniques emerge as a potential solution used for automating the process of detection and presenting efficient and precise results. The initial detection of the RPW remains a difficult task for good production as the identification will protect palm trees infected from the RPW. So advanced technologies like artificial intelligence (AI) and computer vision (CV) can be used in preventing the spread of the RPW on palm trees. Various scholars still working on identifying a precise method for the classification, identification, and localization of the RPW pest. This article develops an automated Red Palm Weevil Detection using Gorilla Troops Optimizer with Deep Learning (RPWD-GTODL) method. The goal of the presented RPWD-GTODL approach lies in the accurate detection and localization of the RPW effectually. To accomplish this, the presented RPWD-GTODL technique initially uses the Gabor filtering (GF) technique to pre-process the images. For RPW detection, the RPWD-GTODL technique uses a Mask RCNN object detector with MobileNetv2 as a backbone network. Moreover, the detection performance of the RPWD-GTODL technique can be boosted by the design of the GTO algorithm for the hyperparameter selection of the MobileNetv2 model. The performance validation of the RPWD-GTODL technique is tested using the RPW dataset and the results demonstrate the enhanced performance on RPW detection process with maximum accuracy of 99.27%.
For irrigation in agriculture, water is a natural resource. Recycling water use is vital for the sustainable development of ecological environment and for resource conservation. Different substances that are thought to be pollutants and contribute to the deterioration of water quality are present in the wastewater from daily life and industrial activity. This research propose novel method in agricultural water management using feature extraction as well as classification based on DL methods. Inputs are collected as agriculture field water management as well as processed for noise removal, normalization and smoothening. Processed input data features are extracted utilizing kernel convolutional component analysis network. The extracted features has been classified using Quadratic reinforcement NN. Experimental analysis are carried out in terms of accuracy, precision, recall, positive predictive value, RMSE and mAP. Proposed technique attained accuracy of 92%, precision of 86%, recall of 65%, positive predictive value of 71%, RMSE of 55%, MAP of 51%.
It is significantly more challenging to extend the visibility factor to a higher depth during the development phase of a communication system for subterranean places. Even if there are numerous optical fiber systems that provide the right energy sources for intended panels, the visibility parameter is not optimized past a certain point. Therefore, the suggested method looks at the properties of a fiber optic communication system that is integrated with a certain energy source while having external panels. A regulating state is established in addition to characteristic analysis by minimizing the reflection index, and the integration of the general adversarial network (GAN) optimizes both central and layer formations in exterior panels. Thus, the suggested technique uses the external noise factor to provide relevant data to the control center via fiber optic shackles. As a result, the normalized error is smaller, boosting the suggested method's effectiveness in all subsurface areas. The created mathematical model is divided into five different situations, and the results are simulated using MATLAB to test the effectiveness of the anticipated strategy. Additionally, comparisons are done for each of the five scenarios, and it is found that the proposed fiber-optic method for energy sources is far more effective than current methodologies.
In all developing countries, the application of biomedical signals has been growing, and there is a potential interest to apply it to healthcare management systems. However, with the existing infrastructure, the system will not provide high-end support for the transfer of signals by using a communication medium, as biomedical signals need to be classified at appropriate stages. Therefore, this article addresses the issues of physical infrastructure, using Hadoop-based systems where a four-layer model is created. The four-layer model is integrated with Fuzzy Interface System Algorithm (FISA) with low robustness, and data transfers in these layers are carried out with reference health data that are collected at various treatment centers. The performance of this new flanged system model aims to minimize the loss functionalities that are present in biomedical signals, and an activation function is introduced at the middle stages. The effectiveness of the proposed model is simulated by using MATLAB, using a biomedical signal processing toolbox, where the performance of FISA proves to be better in terms of signal strength, distance, and cost. As a comparative outcome, the proposed method overlooks the conventional methods for an average percentage of 78% in real-time conditions.
Today, many people under the age of 10 are being examined for brain-related issues, including tumours, without displaying any symptoms. It is not unusual for children to develop brain-related concerns such as tumours and central nervous system disorders, which may affect 15% of the population. Medical experts believe that the irregular eating habits (junk food) and the consumption of pesticide-tainted fruits and vegetables are to blame. The human body is naturally resistant to harmful gears, but only up to a point. If it exceeds the limit, a cell manipulation process is automatically initiated that can remove dangerous inactive tissues from the cell membrane and later grows into tumour blockage in the human body. Thus, the adoption of an advanced computer-based diagnostic system is highly recommended in order to generate visually enhanced images for anomaly identification and infectious tissue segmentation. In most cases, an MR image is chosen since it is easier to distinguish between affected and nonaffected tissue. Conventional convolution neural network (CCNN) mapping and feature extraction are difficult because of the vast volume of data. In addition, it takes a lengthy time for the MRI scanning process to obtain diverse positions for anomaly identification. Aside from the discomfort, the patient may experience motion abnormalities. Recurrent neural network (RNN) classifies tumour regions into several isolated portions much faster and more accurately, so that it can be prevented. To remove motion artefacts from dynamic multicontrast MR images, a novel long short-term memory- (LSTM-) based RNN framework is introduced in this research. With this method, the MR image’s visual quality is improved over CCNN while simultaneously mapping a larger volume and extracting more quiet characteristics than CCNN can. DC-CNN, SMSR-CNN, FMSI-CNN, and DRCA-CNN results are compared. For both low and high signal-to-noise ratios, the suggested LSTM-based RNN framework has gained reasonable feature intelligibility (SNRs). In comparison to previous approaches, it requires less computing and has higher accuracy when it comes to detecting infected portions.
The demographic and population modeling methods have been under investigation trends since the 1980s. Extrapolation, prediction, and theoretical computational analysis of exogenous variables are approaches to the forecasting of population processes. Such methods can be exploited to predict individual birth preferences or experts’ views at the population level. Predicting demographic changes have been problematic while its precision usually depends on the case or pattern; numerous methods have been explored; however, so far there is no clear guidelines where the proper approach ought to be. Like certain fields of industry and policy, planning is focused on projections for the future composition of the population, the potential creation of population sizes and institutions which are significant. In order to recognize potential social security issues as one determinant of overall macroeconomic growth, countries that have reduced mortality and low fertility, the case with some of the Asian nations, desperately require accurate demographic estimates. This introduction provides a stochastic cohort model that uses stochastic fertility, migration, and mortality modeling approaches to forecast the population by gender and literacy. This work focus on the population and literacy ratio of India as this nation holds the second largest population in the world. Our approach is based on artificial neural network algorithm that can forecast the population literacy ratio and gender differences based on living states populations using social networks data. We concentrated primarily on quantifying future planning challenges as previous research appeared to neglect potential risks. Our model is then used to forecast/predict gender-wise population for each major state/city. The findings offer clear perspectives on the projected gender demographic composition, and our model holds high precision results.
In recent times, internet of things (IoT) applications on the cloud might not be the effective solution for every IoT scenario, particularly for time sensitive applications. A significant alternative to use is edge comput-ing that resolves the problem of requiring high bandwidth by end devices. Edge computing is considered a method of forwarding the processing and communication resources in the cloud towards the edge. One of the consid-erations of the edge computing environment is resource management that involves resource scheduling, load balancing, task scheduling, and quality of service (QoS) to accomplish improved performance. With this motivation, this paper presents new soft computing based metaheuristic algorithms for resource scheduling (RS) in the edge computing environment. The SCBMA-RS model involves the hybridization of the Group Teaching Optimization Algorithm (GTOA) with rat swarm optimizer (RSO) algorithm for optimal resource allocation. The goal of the SCBMA-RS model is to identify and allocate resources to every incoming user request in such a way, that the client???s necessities are satisfied with the minimum number of possible resources and optimal energy consumption. The problem is formulated based on the availability of VMs, task characteristics, and queue dynamics. The integration of GTOA and RSO algorithms assist to improve the allocation of resources among VMs in the data center. For experimental validation, a comprehensive set of simulations were performed using the CloudSim tool. The experimental results showcased the superior performance of the SCBMA-RS model interms of different measures.
The majority of people in the modern biosphere struggle with depression as a result of the coronavirus pandemic’s impact, which has adversely impacted mental health without warning. Even though the majority of individuals are still protected, it is crucial to check for post-corona virus symptoms if someone is feeling a little lethargic. In order to identify the post-coronavirus symptoms and attacks that are present in the human body, the recommended approach is included. When a harmful virus spreads inside a human body, the post-diagnosis symptoms are considerably more dangerous, and if they are not recognised at an early stage, the risks will be increased. Additionally, if the post-symptoms are severe and go untreated, it might harm one’s mental health. In order to prevent someone from succumbing to depression, the technology of audio prediction is employed to recognise all the symptoms and potentially dangerous signs. Different choral characters are used to combine machine-learning algorithms to determine each person’s mental state. Design considerations are made for a separate device that detects audio attribute outputs in order to evaluate the effectiveness of the suggested technique; compared to the previous method, the performance metric is substantially better by roughly 67%.
Mobile edge computing (MEC) is a paradigm novel computing that promises the dramatic effect of reduction in latency and consumption of energy by computation offloading intensive; these tasks to the edge clouds in proximity close to the smart mobile users. In this research, reduce the offloading and latency between the edge computing and multiusers under the environment IoT application in 5G using bald eagle search optimization algorithm. The deep learning approach may consume high computational complexity and more time. In an edge computing system, devices can offload their computation-intensive tasks to the edge servers to save energy and shorten their latency. The bald eagle algorithm (BES) is the advanced optimization algorithm that resembles the strategy of eagle hunting. The strategies are select, search, and swooping stages. Previously, the BES algorithm is used to consume the energy and distance; to improve the better energy and reduce the offloading latency in this research and some delays occur when devices increase causes demand for cloud data, it can be improved by offering ROS (resource) estimation. To enhance the BES algorithm that introduces the ROS estimation stage to select the better ROSs, an edge system, which offloads the most appropriate IoT subtasks to edge servers then the expected time of execution, got minimized. Based on multiuser offloading, we proposed a bald eagle search optimization algorithm that can effectively reduce the end-end time to get fast and near-optimal IoT devices. The latency is reduced from the cloud to the local; this can be overcome by using edge computing, and deep learning expects faster and better results from the network. This can be proposed by BES algorithm technique that is better than other conventional methods that are compared on results to minimize the offloading latency. Then, the simulation is done to show the efficiency and stability by reducing the offloading latency.
The word radiomics, like all domains of type omics, assumes the existence of a large amount of data. Using artificial intelligence, in particular, different machine learning techniques, is a necessary step for better data exploitation. Classically, researchers in this field of radiomics have used conventional machine learning techniques (random forest, for example). More recently, deep learning, a subdomain of machine learning, has emerged. Its applications are increasing, and the results obtained so far have demonstrated their remarkable effectiveness. Several previous studies have explored the potential applications of radiomics in colorectal cancer. These potential applications can be grouped into several categories like evaluation of the reproducibility of texture data, prediction of response to treatment, prediction of the occurrence of metastases, and prediction of survival. Few studies, however, have explored the potential of radiomics in predicting recurrence-free survival. In this study, we evaluated and compared six conventional learning models and a deep learning model, based on MRI textural analysis of patients with locally advanced rectal tumours, correlated with the risk of recidivism; in traditional learning, we compared 2D image analysis models vs. 3D image analysis models, models based on a textural analysis of the tumour versus models taking into account the peritumoural environment in addition to the tumour itself. In deep learning, we built a 16-layer convolutional neural network model, driven by a 2D MRI image database comprising both the native images and the bounding box corresponding to each image.
In Machine Learning , if one class has a significantly larger number of instances (majority) than the other (minority), this condition is defined as class imbalance. With regard to datasets, class imbalance can bias the predictive capabilities of Machine Learning algorithms towards the majority (negative) class, and in situations where false negatives incur a greater penalty than false positives, this imbalance may lead to adverse consequences. Our paper incorporates two case studies, each utilizing a unique approach of three learners (gradient-boosted trees, logistic regression, random forest) and three performance metrics ( Area Under the Receiver Operating Characteristic Curve , Area Under the Precision-Recall Curve , Geometric Mean ) to investigate class rarity in big data. Class rarity, a notably extreme degree of class imbalance, was effected in our experiments by randomly removing minority (positive) instances to artificially generate eight subsets of gradually decreasing positive class instances. All model evaluations were performed through Cross-Validation. In the first case study, which uses a Medicare Part B dataset, performance scores for the learners generally improve with the Area Under the Receiver Operating Characteristic Curve metric as the rarity level decreases, while corresponding scores with the Area Under the Precision-Recall Curve and Geometric Mean metrics show no improvement. In the second case study, which uses a dataset built from Distributed Denial of Service attack attack data (POSTSlowloris Combined), the Area Under the Receiver Operating Characteristic Curve metric produces very high-performance scores for the learners, with all subsets of positive class instances. For the second study, scores for the learners generally improve with the Area Under the Precision-Recall Curve and Geometric Mean metrics as the rarity level decreases. Overall, with regard to both case studies, the Gradient-Boosted Trees (GBT) learner performs the best.
This introduction presents an overview of the key concepts discussed in the subsequent chapters of this book. The book focuses on sampling to reduce the impact of class imbalance on machine learning models. It demonstrates that classification performance across several imbalanced big datasets across different application domains can be significantly improved using Random Undersampling without substantially altering the composition of the original data. The book provides an overview of related works. It describes the Machine Learning (ML) classification algorithms and libraries, to include the evaluation strategy with validation techniques and performance metrics. The book introduces the datasets and how they were processed, model training, and performance evaluation. It involves a real-world Medicare fraud problem, with severe class imbalance. To ease the process of using ML, engineers build the algorithms within software modules or packages, making sure that they work reliably, quickly, and at-scale.
Severe class imbalance between the majority and minority classes in large datasets can prejudice Machine Learning classifiers toward the majority class. Our work uniquely consolidates two case studies, each utilizing three learners implemented within an Apache Spark framework, six sampling methods, and five sampling distribution ratios to analyze the effect of severe class imbalance on big data analytics. We use three performance metrics to evaluate this study: Area Under the Receiver Operating Characteristic Curve, Area Under the Precision-Recall Curve, and Geometric Mean. In the first case study, models were trained on one dataset (POST) and tested on another (SlowlorisBig). In the second case study, the training and testing dataset roles were switched. Our comparison of performance metrics shows that Area Under the Precision-Recall Curve and Geometric Mean are sensitive to changes in the sampling distribution ratio, whereas Area Under the Receiver Operating Characteristic Curve is relatively unaffected. In addition, we demonstrate that when comparing sampling methods, borderline-SMOTE2 outperforms the other methods in the first case study, and Random Undersampling is the top performer in the second case study.
High class imbalance between majority and minority classes in datasets can skew the performance of Machine Learning algorithms and bias predictions in favor of the majority (negative) class. This bias, for cases where the minority (positive) class is of greater interest and the occurrence of false negatives is costlier than false positives, may result in adverse consequences. Our paper presents two case studies, each utilizing a unique, combined approach of Random Undersampling and Feature Selection to investigate the effect of class imbalance on big data analytics. Random Undersampling is used to generate six class distributions ranging from balanced to moderately imbalanced, and Feature Importance is used as our Feature Selection method. Classification performance was reported for the Random Forest, Gradient-Boosted Trees, and Logistic Regression learners, as implemented within the Apache Spark framework. The first case study utilized a training dataset and a test dataset from the ECBDL'14 bioinformatics competition. The training and test datasets contain about 32 million instances and 2.9 million instances, respectively. For the first case study, Gradient-Boosted Trees obtained the best results, with either a features-set of 60 or the full set, and a negative-to-positive ratio of either 45:55 or 40:60. The second case study, unlike the first, included training data from one source (POST dataset) and test data from a separate source (Slowloris dataset), where POST and Slowloris are two types of Denial of Service attacks. The POST dataset contains about 1.7 million instances, while the Slowloris dataset contains about 0.2 million instances. For the second case study, Logistic Regression obtained the best results, with a features-set of 5 and any of the following negative-to-positive ratios: 40:60, 45:55, 50:50, 65:35, and 75:25. We conclude that combining Feature Selection with Random Undersampling improves the classification performance of learners with imbalanced big data from different application domains.
This paper aims to address a key research issue regarding the ECBDL’14 bioinformatics big data competition. The ECBDL’14 dataset was the big data target in the competition, and it consisted of 631 attributes and about 32 million instances, of which about 98% belonged to the negative class. The ECBDL’14 competition dataset has recently been used in the literature to assess the effect of class imbalance on big data analytics. The contribution of our paper is two-fold. First, a survey of several literature works that utilized the EDBDL’14 dataset, either fully or partially, is presented. Second, compared to the Random Oversampling approach used by the winning algorithm for the competition, we utilize a Random Undersampling approach in conjunction with a Feature Selection approach. Through Random Undersampling, different class distributions were generated, ranging from slightly imbalanced to balanced. Prior to sampling, we perform Feature Selection by computing Feature Importance with the Random Forest learner within the Apache Spark framework. Subsequently, classification performance is computed for Random Forest, Logistic Regression, and Gradient-Boosted Trees in the same big data analytics framework. The key results of our study indicate that our proposed solution had a higher prediction performance (minimally) compared to that of the highest value of the winning algorithm. However, it is important to note that Random Undersampling, compared to Random Oversampling, imposes a lower computational burden and results in a faster training time, which is beneficial to data analytics. We conclude that our solution clearly outperforms the ECBLD’14 winning algorithm.