The proliferation of Internet of Things (IoT) devices has introduced unprecedented attack surfaces, necessitating highly efficient intrusion detection systems (IDS). This paper proposes Q-LDPSO, a Quantum-Enhanced Leaders-Driven Particle Swarm Optimizer for IoT IDS feature selection. Two principal novel contributions are made to the base LDPSO algorithm: (i) Quantum Population Initialisation (QPI), where particles are encoded as qubits in Hadamard superposition states and collapsed to binary vectors with provably superior spatial diversity and a Quantum Bidirectional Search Strategy (QBSS) for leader particles using quantum rotation gate updates in both convergent and divergent directions. The proposed system Q-LDPSO drives feature selection for an RF + XGBoost ensemble classifier evaluated on CIC-IoT2023, N-BaIoT, and BoT-IoT datasets, achieving 99.63% accuracy, 99.71% detection rate, 0.15% false alarm rate, and 99.68% F1-score outperforming all 9 compared algorithms, with superior convergence, diversity, and feature compactness. Ablation studies are conducted for confirm synergistic contributions of both novel components.
Malware remains one of the most significant threats to modern computer and digital systems. Owing to its continuously evolving nature, preventing and detecting malware has become increasingly complex. Traditional rule-based detection techniques are often ineffective against novel or obfuscated malware variants. Consequently, recent research has focused on leveraging machine learning and deep learning methods to identify evolving malicious patterns that can detect previously unseen threats. Since 2015, there has been a substantial growth in studies applying artificial intelligence approaches for malware detection and classification. This paper presents a contemporary review of state-of-the-art machine learning and deep learning algorithms used in malware analysis. To ensure methodological rigor and transparency, the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) framework was adopted. The review critically evaluates the strengths and limitations of existing models, analyses their performance in detecting and classifying diverse malware types, and discusses emerging challenges such as dataset imbalance, adversarial attacks, and model generalizability. Finally, the paper outlines key research directions aimed at improving the robustness, scalability, and interpretability of AI-driven malware detection systems.
IntroductionIn the evolving landscape of healthcare and medicine, the merging of extensive medical datasets with the powerful capabilities of machine learning (ML) models presents a significant opportunity for transforming diagnostics, treatments, and patient care.MethodsThis research paper delves into the realm of data-driven healthcare, placing a special focus on identifying the most effective ML models for diabetes prediction and uncovering the critical features that aid in this prediction. The prediction performance is analyzed using a variety of ML models, such as Random Forest (RF), XG Boost (XGB), Linear Regression (LR), Gradient Boosting (GB), and Support VectorMachine (SVM), across numerousmedical datasets. The study of feature importance is conducted using methods including Filter-based, Wrapper-based techniques, and Explainable Artificial Intelligence (Explainable AI). By utilizing Explainable AI techniques, specifically Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP), the decision-making process of the models is ensured to be transparent, thereby bolstering trust in AI-driven decisions.ResultsFeatures identified by RF in Wrapper-based techniques and the Chi-square in Filter-based techniques have been shown to enhance prediction performance. A notable precision and recall values, reaching up to 0.9 is achieved in predicting diabetes.DiscussionBoth approaches are found to assign considerable importance to features like age, family history of diabetes, polyuria, polydipsia, and high blood pressure, which are strongly associated with diabetes. In this age of data-driven healthcare, the research presented here aspires to substantially improve healthcare outcomes.
This paper delves into the dynamic intersection of machine learning (ML) and healthcare, envisioning a paradigm shift in diagnostic accuracy, personalized treatment, and streamlined administration. It meticulously explores various ML algorithms, spanning deep learning, decision trees, and clustering techniques, pivotal in domains like early cancer detection, diabetes detection, heart disease detection, autism spectrum disorder detection, and Parkinson’s disease detection. Rigorous model evaluation, employing accuracy, precision, F1-score, specificity, and mean squared error metrics, ensures algorithm dependability. However, data privacy challenges, amplified by intricate regulations, persist. Ethical considerations add complicated dimensions, including algorithmic bias and cultivating patient trust. Addressing these necessitates robust education for healthcare professionals and alignment with legal frameworks. Despite challenges, the paper advocates for a conscientious integration of ML, emphasizing its transformative potential in healthcare and urging judicious technology amalgamation to propel advancements in patient care and clinical outcomes.
The texture is identifiable in optical and easy ways. Texture classification is an imperative region in texture analysis, where it gives descriptors for classifying the images. The categorization of normal and abnormal matter by magnetic resonance (MR), computed tomography (CT), and texture images has made noteworthy evolution in modern years. Recently, different novel robust classification techniques have been introduced to classify the different kinds of images for prediction. However, the accuracy of classification was not improved with lesser time. To address these issues, the edge-preserved Tversky indexive Hellinger and deep perceptive Czekanowski classifier (ETIH-DPCC) technique is introduced to segment and classify the images with more accuracy. The ETIH-DPCC technique includes diverse processes namely preprocessing, segmentation, feature extraction, as well as classification. At first, different types of images, such as magnetic resonance imaging, CT, and texture, are used as input. With the acquired input, edge-preserving normalized adaptive bilateral filtering is employed to carry the image preprocessing. In this stage, the noisy pixels are removed and edges are preserved. Then, the Tversky-indexed quantile regression is applied to segment the images into diverse texture regions. After that, the feature extraction is performed on the segmented region using Hellinger kernel feature extraction, where a more informative feature for image prediction is extracted. During this process, the irrelevant features are avoided to decrease the dimensionality and feature extraction time. These extracted features are finally classified into positive and negative classes for disease prediction using DPCC. DPCC comprises multiple layers to deeply analyze the association between training and testing features. With this, the prediction accuracy is improved. Experimental outcomes show that the ETIH-DPCC technique efficiently enhances prediction accuracy and less time compared to conventional methods.
Droughts typically develop gradually, and early prediction is crucial for the government to formulate effective mitigation plans. Our approach does not involve predicting specific drought index values. Instead, we forecast whether a particular year will experience drought. Insufficient investigation has been carried out regarding variations in additional climatic indicators like shortwave radiation, wind speed, sea level, and pollution in the context of droughts in the state of Tamil Nadu, India. In the study period taken from 1995 to 2020, only three years (2002, 2009, and 2017) experienced drought occurrences, resulting in an imbalanced dataset. To enhance the classification performance of this imbalanced dataset, a weighted dataset is constructed using a feature weighting approach known as the Single Objective Scorer (SOS) based Multi-objective PSO(MPSO) in conjunction with the Gradient Boosting Classifier. The proposed model facilitates objective-based multi-population formation and neighborhood learning. Precision and recall are crucial metrics, particularly in measuring imbalanced dataset classification performance. The application of multi-objective optimization techniques helps to strike a suitable balance between precision and recall. In addition to the Standardized Precipitation Index (SPI) and Standardized Precipitation Evapotranspiration Index (SPEI), 14 climatic indicators based on land, atmosphere, and sea are utilized. By employing the weighted dataset created with SOS-based MPSO, a significant improvement in recall value of 0.81 is achieved. Based on the weights assigned to the features, it is identified that the Mean Sea Level of the Arabian Sea and CO 2 are significant indicators for predicting meteorological drought. The Explainable AI techniques SHAP and LIME are employed for interpreting the drought prediction model, providing insights into its workings.
In today’s world, there has been a significant increase in the use of devices, gadgets, and mobile applications in our daily activities. Although this has had a significant impact on the lives of the general public, people who are Partially Visually Impaired SPVI, which includes a much broader range of vision loss that includes mild to severe impairments, and Completely Visually Impaired (CVI), who have no light perception, still face significant obstacles when trying to access and use these technologies. This review article aims to provide an overview of the NUI, Multi-sensory Interfaces and UX Design (NMUD) of apps and devices specifically tailored CVI and PVI individuals. The article begins by emphasizing the importance of accessible technology for the visually impaired and the need for a human-centered design approach. It presents a taxonomy of essential design components that were considered during the development of applications and gadgets for individuals with visual impairments. Furthermore, the article sheds light on the existing challenges that need to be addressed to improve the design of apps and devices for CVI and PVI individuals. These challenges include usability, affordability, and accessibility issues. Some common problems include battery life, lack of user control, system latency, and limited functionality. Lastly, the article discusses future research directions for the design of accessible apps and devices for visually impaired individuals. It emphasizes the need for more user-centered design approaches, adherence to guidelines such as the Web Content Accessibility Guidelines, the application of e-accessibility principles, the development of more accessible and affordable technologies, and the integration of these technologies into the wider assistive technology ecosystem.
Huntington’s Disease (HD) is a devastating neurodegenerative disorder characterized by progressive motor dysfunction, cognitive impairment, and psychiatric symptoms. The early and accurate diagnosis of HD is crucial for effective intervention and patient care. This comprehensive review provides a comprehensive overview of the utilization of Artificial Intelligence (AI) powered algorithms in the diagnosis of HD. This review systematically analyses the existing literature to identify key trends, methodologies, and challenges in this emerging field. It also highlights the potential of ML and DL approaches in automating HD diagnosis through the analysis of clinical, genetic, and neuroimaging data. This review also discusses the limitations and ethical considerations associated with these models and suggests future research directions aimed at improving the early detection and management of Huntington’s disease. It also serves as a valuable resource for researchers, clinicians, and healthcare professionals interested in the intersection of machine learning and neurodegenerative disease diagnosis.
It is hard to exaggerate how crucial user reviews are in figuring out an organization's e-commerce income. Before purchasing any goods or service, online customers rely on product and service reviews. As a result, firms must consider how reliable Internet reviews are because they might directly affect their reputation and bottom line. Because of this, some businesses pay spammers to post false reviews. These false reviews profit from consumer buying choices. As a result, during the past twelve years, there has been substantial study into techniques for spotting false reviews. However, there is still a need for a survey that can analyze and summarize the diverse techniques. In this paper, we are going to put the SVM and Naive Bayesian machine learning model system into place that can spot false reviews with an accuracy of 0.801 and 0.687.
Powder bed fusion (PBF) applies to various metallic materials used in the metal printing process of building a wide range of complex parts compared to other AM technologies. PBF process has several variants such as DMLS (direct metal laser sintering), EBM (electron beam melting), SHS (selective heat sintering), SLM (selective laser melting), and SLS (selective laser sintering). For PBF to reach its maximum potential, machine learning (ML) algorithms are used with suitable materials to achieve goals cost-effectively. Various applications of neural networks, including ANNs, CNNs, RNNs, and other popular techniques such as KNN, SVM, and GP were reviewed, and future challenges were discussed. Some special-purpose algorithms were listed as follows: GAN, SeDANN, SCNN, K-means, PCA, etc. This review presents the evolution, current status, challenges, and prospects of these technologies in terms of material, features, process parameters, applications, advantages, disadvantages, etc., to explain their significance and provide an in-depth understanding of the same.
Since the invention of computer resources, intrusions into the computing environment have been a fairly widespread form of hostile conduct. Over the past three decades, a variety of security measures have been put in place, but as technology has advanced, so too have security risks. Preventing virus infections and other dangers to the computing infrastructures is crucial since computers are used by everyone in the globe, whether directly or indirectly. One of the most common types of harmful programs that attaches to others and typically runs before the host programs is the virus. The majority of current detection techniques relies on a single class of scanning algorithm. However, they are unable to detect all types viruses. Thus, a combined scanning method in single antivirus program is proposed which is more effective in scanning wide range of viruses. Signature-based virus detection is accomplished by comparing observed events to pre-defined signatures. An identifiable threat has a pattern known as a signature (Jayakumar et al. in Fusion of heterogeneous intrusion detection systems for network attack detection. Sci World J [1]). Based on file modification, heuristic methods are used to discover unknown computer viruses. The virus program is proposed to highlight all the crucial features of the antivirus application.
Intrauterine fetal demise in women during pregnancy is a major contributing factor in prenatal mortality and is a major global issue in developing and underdeveloped countries. When an unborn fetus passes away in the womb during the 20th week of pregnancy or later, early detection of the fetus can help reduce the chances of intrauterine fetal demise. Machine learning models such as Decision Trees, Random Forest, SVM Classifier, KNN, Gaussian Naïve Bayes, Adaboost, Gradient Boosting, Voting Classifier, and Neural Networks are trained to determine whether the fetal health is Normal, Suspect, or Pathological. This work uses 22 features related to fetal heart rate obtained from the Cardiotocogram (CTG) clinical procedure for 2126 patients. Our paper focuses on applying various cross-validation techniques, namely, K-Fold, Hold-Out, Leave-One-Out, Leave-P-Out, Monte Carlo, Stratified K-fold, and Repeated K-fold, on the above ML algorithms to enhance them and determine the best performing algorithm. We conducted exploratory data analysis to obtain detailed inferences on the features. Gradient Boosting and Voting Classifier achieved 99% accuracy after applying cross-validation techniques. The dataset used has the dimension of 2126 × 22, and the label is multiclass classified as Normal, Suspect, and Pathological condition. Apart from incorporating cross-validation strategies on several machine learning algorithms, the research paper focuses on Blackbox evaluation, which is an Interpretable Machine Learning Technique used to understand the underlying working mechanism of each model and the means by which it picks features to train and predict values.
In the industrial machining process, there have been major advances in near-net-shaped forming, which leads machining to be considered a significant modern phenomenon. Machining turns a huge number of metals into chips every year. This study aimed to determine the wear and mechanical properties of various cutting inserts. Polycrystalline diamond (PCD) and Ceramic Inserts were selected as coated inserts. It was discovered that tool wear at the cutting edge impacts various factors, including the amount of cutting forces created during machining; the surface finish of the workpiece is also compromised, resulting in reduced tool life. Owing to the frequent replacement of cutting tools, the decreased wear rate of cutting tools exponentially raises the costs that companies/machine shops would incur. After the second iteration, this insert began to develop crater wear, which resulted in a poor surface finish and high heat generation. However, the surface finish of this instrument was discovered to be the best during the first iteration. From the outcome, the PCD coated tool with feed speeds and low depth of cuts performed the efficient machining process. The surface finish is also accurate for PCD coated tool. The bat and whale algorithms’ optimization involved to find the best technical parameters to achieve the lowest possible error value based on rake and face wear. The bat and whale algorithms were used to determine the optimized rake and face wear values. The bat algorithm outperforms the whale algorithm in terms of wear value predictions.
Automatic image caption generation is an intricate task of describing an image in natural language by gaining insights present in an image. Featuring facial expressions in the conventional image captioning system brings out new prospects to generate pertinent descriptions, revealing the emotional aspects of the image. The proposed work encapsulates the facial emotional features to produce more expressive captions similar to human-annotated ones with the help of Cross Stage Partial Dense Network (CSPDenseNet) and Self-attentive Bidirectional Long Short-Term Memory (BiLSTM) network. The encoding unit captures the facial expressions and dense image features using a Facial Expression Recognition (FER) model and CSPDense neural network, respectively. Further, the word embedding vectors of the ground truth image captions are created and learned using the Word2Vec embedding technique. Then, the extracted image feature vectors and word vectors are fused to form an encoding vector representing the rich image content. The decoding unit employs a self-attention mechanism encompassed with BiLSTM to create more descriptive and relevant captions in natural language. The Flickr11k dataset, a subset of the Flickr30k dataset is used to train, test, and evaluate the present model based on five benchmark image captioning metrics. They are BiLingual Evaluation Understudy (BLEU), Metric for Evaluation of Translation with Explicit Ordering (METEOR), Recall-Oriented Understudy for Gisting Evaluation (ROGUE), Consensus-based Image Description Evaluation (CIDEr), and Semantic Propositional Image Caption Evaluation (SPICE). The experimental analysis indicates that the proposed model enhances the quality of captions with 0.6012(BLEU-1), 0.3992(BLEU-2), 0.2703(BLEU-3), 0.1921(BLEU-4), 0.1932(METEOR), 0.2617(CIDEr), 0.4793(ROUGE-L), and 0.1260(SPICE) scores, respectively, using additive emotional characteristics and behavioral components of the objects present in the image.
Micro-electric discharge machining (Micro-EDM) is deployed for machining hard-to-machine materials, such as various grades of titanium alloys, heat-treated alloy steels, composites, tungsten carbides, and so forth. Mild steel is known for its easy machinability. However, conventional machining of mild steel can often lead to the built-up edge formation on the tool. There is a minimal focus on machining ductile materials using nonconventional machining processes. This is due to the rapid work hardening in cold forming conditions. In the present study, the aluminium alloy 6061 and mild steel AISI 304 were taken as a work piece. Input pulse on factors considered as three levels and orthogonal array utilized to optimize the EDM parameters. Numerical results confirm the influence of input parameters in the response. The highest MRR is obtained at Ton = 40 μs and Toff = 4 μs, and the least MRR is acquired at Ton = 20 μs and Toff = 3 μs. The fruit fly algorithm and the cockroach swarm algorithm were used to predict the optimal minimized MRR value. The experimental results show that the cockroach swarm algorithm was performing better than the fruit fly algorithm in the MRR minimization process.
The novel coronavirus is a family of animal transferred viruses that can cause illness in humans. This virus took over the world in 2019 and WHO deemed it as an epidemic naming it as COVID-19. A lot of research has gone in the prediction of the trends and classification of cases using various machine learning and deep learning techniques. With the outbreak of this pandemic, efficient detection of the disease at a faster rate has become very crucial. This study proposes a convolutional neural network (CNN)-based deep learning approach for classification of COVID-19 positive cases from normal cases using X-Ray radiology scans of the patients. The model consists of a large custom dataset of images extracted from an open source dataset and is then trained using our proposed model. Different optimizer algorithms are also compared in order to check which one of them gives the most accuracy. The model is further tested using the categorical accuracy metrics and then a graphical analysis of the results is provided. A comparative study is also conducted with an already existing support vector machines (SVM) model. The images were trained according to three classifications: normal, COVID infected, and viral pneumonia infected patients. The main objective of our research is to help further the research in early diagnosis of COVID-19 using modern deep learning techniques.
Drought is the least understood natural disaster due to the complex relationship of multiple contributory factors. Its beginning and end are hard to gauge, and they can last for months or even for years. India has faced many droughts in the last few decades. Predicting future droughts is vital for framing drought management plans to sustain natural resources. The data-driven modelling for forecasting the metrological time series prediction is becoming more powerful and flexible with computational intelligence techniques. Machine learning (ML) techniques have demonstrated success in the drought prediction process and are becoming popular to predict the weather, especially the minimum temperature using backpropagation algorithms. The favourite ML techniques for weather forecasting include singular vector machines (SVM), support vector regression, random forest, decision tree, logistic regression, Naive Bayes, linear regression, gradient boosting tree, k-nearest neighbours (KNN), the adaptive neuro-fuzzy inference system, the feed-forward neural networks, Markovian chain, Bayesian network, hidden Markov models, and autoregressive moving averages, evolutionary algorithms, deep learning and many more. This paper presents a recent review of the literature using ML in drought prediction, the drought indices, dataset, and performance metrics.
This paper aims to evaluate the performance of multiple non-linear regression techniques, such as support-vector regression (SVR), k-nearest neighbor (KNN), Random Forest Regressor, Gradient Boosting, and XGBOOST for COVID-19 reproduction rate prediction and to study the impact of feature selection algorithms and hyperparameter tuning on prediction. Sixteen features (for example, Total_cases_per_million and Total_deaths_per_million) related to significant factors, such as testing, death, positivity rate, active cases, stringency index, and population density are considered for the COVID-19 reproduction rate prediction. These 16 features are ranked using Random Forest, Gradient Boosting, and XGBOOST feature selection algorithms. Seven features are selected from the 16 features according to the ranks assigned by most of the above mentioned feature-selection algorithms. Predictions by historical statistical models are based solely on the predicted feature and the assumption that future instances resemble past occurrences. However, techniques, such as Random Forest, XGBOOST, Gradient Boosting, KNN, and SVR considered the influence of other significant features for predicting the result. The performance of reproduction rate prediction is measured by mean absolute error (MAE), mean squared error (MSE), root mean squared error (RMSE), R-Squared, relative absolute error (RAE), and root relative squared error (RRSE) metrics. The performances of algorithms with and without feature selection are similar, but a remarkable difference is seen with hyperparameter tuning. The results suggest that the reproduction rate is highly dependent on many features, and the prediction should not be based solely upon past values. In the case without hyperparameter tuning, the minimum value of RAE is 0.117315935 with feature selection and 0.0968989 without feature selection, respectively. The KNN attains a low MAE value of 0.0008 and performs well without feature selection and with hyperparameter tuning. The results show that predictions performed using all features and hyperparameter tuning is more accurate than predictions performed using selected features.
This paper focuses to solve two important problems faced in this digital era. The first problem is handling large amount of digital data. Digital sectors are finding difficulty in handling repetitive works and standardization of process. Our University is a eco-friendly, digital university each and every work of our students are maintained digitally. So for a course instructor to open, view, correct and finding plagiarism in the assignment or homework submitted by students is a tough task. So in this paper this tough task is simplified by applying Robotic Process Automation (RPA). The process of downloading and reading the students’ digital document is automated. The second problem is the plagiarism. There has been a massive threat in the academic and technological integrity of innovation in the name of paraphrasing. In a survey conducted by Donald McCabe of Rutgers University over 70 percent of students admit to paraphrasing and copying of others work. They have also fabricated bibliography. This problem can be addressed using the state-of-the-art document similarity deep learning neural networks. In this paper, we propose the design and implementation of a Siamese Neural Network (SNN), Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) to identify paraphrasing, using only features of similarity between words where the dataset is evolutionary and the model keeps learning with incremental sources of the document’s plagiarism. This paper will also provide the Tensorflow based implementation of such a model. The proposed system will predict the similarity between documents and predict the plagiarism percentage. Keywords: RPA, SNN, CNN, LSTM, Tensorflow
Machine learning is a part of artificial intelligence in which the learning was done using the data available in the environment. Machine learning algorithms are mainly used in game development to change from presripted games to adaptive play games. The main theme or plot of the game, game levels, maps in route, and racing games are considered as content. Context refers to the game screenplay, sound effects, and visual effects. In any type of game, maintaining the fun mode of the player is very important. Predictable moves by non-players in the game and same type of visual effects will reduce the player's interest in the game. The machine learning algorithms works in automatic content generation and nonpayer character behaviours in gameplay. In pathfinding games, puzzle games, strategy games adding intelligence to enemy and opponents makes the game more interesting. The enjoyment and fun differs from game to game. For example, in horror games, fun is experienced when safe point is reached.