This paper introduces a novel hybrid framework for early cervical cancer risk prediction that integrates domain-informed feature engineering, evolutionary optimization, and transformer-based modeling. We propose four key innovations: (1) the derivation of clinically meaningful features based on epidemiological relationships, such as Sexual Activity Duration and STD Diagnosis Rate; (2) architecture-aware feature selection using Particle Swarm Optimization (PSO) tailored for transformer networks; (3) an imbalance-aware training strategy combining Synthetic Minority Oversampling Technique (SMOTE) and focal loss to address extreme class skew; and (4) clinically actionable interpretability via Shapley Additive Explanations (SHAP), Local Interpretable Model-agnostic Explanations (LIME), and attention weights, ensuring transparent and trustworthy decision support. These components are embedded within a TabTransformer architecture, which utilizes self-attention mechanisms to analyze tabular health records and capture the complex interdependencies among risk factors. A comparative evaluation demonstrates the superior results of our approach, achieving 95.3% +/- 0.9% accuracy and a 94.8% +/- 1.1% F1 score on the UCI Cervical Cancer Risk dataset while maintaining clinical relevance and transparency. Furthermore, its interpretability, validated through Shapley Additive Explanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME), confirmed alignment with established risk factors, reinforcing its potential clinical relevance.
Feature selection is a crucial step in the data preprocessing stage of machine learning. It involves selecting a subset of relevant features for use in model construction. Feature selection helps in improving model performance by reducing overfitting, enhancing generalization, and decreasing computational cost. Techniques for feature selection can be broadly classified into filter methods, wrapper methods, and embedded methods. This paper presents a feature selection method based on Particle Swarm Optimization (PSO). The proposed algorithm makes use of a guided particle scheme whereby three filter-based methods are incorporated. The proposed algorithm addresses the issue of premature convergence to global optima compared to other PSO feature-based methods. In addition, the algorithm is tested on very-high-dimensional genome data that include up to 44,909 features. Results of an experimental comparison with other state-of-the-art feature selection algorithms show that the proposed algorithm produces overall better results.
Air pollution remains a critical issue in many countries, particularly in developing regions like India. In November 2019, New Delhi experienced an alarming Air Quality Index (AQI) level of 900, far exceeding the ‘severe’ threshold. Accurate air pollution forecasting is essential for informed decision-making due to its direct impact on public health. Effective forecasting depends on selecting suitable methods and evaluation measures to maximize prediction accuracy. This study investigates the prediction of air quality levels in India, emphasizing techniques to enhance forecasting precision and identify areas for further model improvement. Our research evaluates various Neural Networks, deep learning, and machine learning algorithms for predicting AQI. The findings reveal that among the techniques explored, the Long Short-Term Memory (LSTM) model outperforms others, demonstrating very good accuracy in capturing complex patterns in AQI data. This contribution serves as a valuable benchmark for experts aiming to refine forecasting methods and provides a reference for emerging researchers in the field.
Cancer classification using high-dimensional genomic data presents significant challenges in feature selection, particularly when dealing with datasets containing tens of thousands of features. This study presents a new application of the Simultaneous Perturbation Stochastic Approximation (SPSA) method for feature selection on large-scale cancer datasets, representing the first investigation of the SPSA-based feature selection technique applied to cancer datasets of this magnitude. Our research extends beyond traditional SPSA applications, which have historically been limited to smaller datasets, by evaluating its effectiveness on datasets containing 35,924 to 44,894 features. Building upon established feature-ranking methodologies, we introduce a comprehensive evaluation framework that examines the impact of varying proportions of top-ranked features (5%, 10%, and 15%) on classification performance. This systematic approach enables the identification of optimal feature subsets most relevant to cancer detection across different selection thresholds. The key contributions of this work include the following: (1) the first application of SPSA-based feature selection to large-scale cancer datasets exceeding 35,000 features, (2) an evaluation methodology examining multiple feature proportion thresholds to optimize classification performance, (3) comprehensive experimental validation through comparison with ten state-of-the-art feature selection and classification methods, and (4) statistical significance testing to quantify the improvements achieved by the SPSA approach over benchmark methods. Our experimental evaluation demonstrates the effectiveness of the feature selection and ranking-based SPSA method in handling high-dimensional cancer data, providing insights into optimal feature selection strategies for genomic classification tasks.
Quantitative Structure-Activity/Property Relationship (QSAR/QSPR) is a machine learning approach to predict chemical and physical properties of pure compounds; however, it has limited application in multi-component compounds. The complex and layered nature of multi-component materials presents challenges in computing molecular representation, thus limiting the application of QSAR and QSPR. In this study, a new method has been proposed to derive numerical representation based on a combinatorial approach. It calculates all the possible interactions between different components in reaction using the Cartesian product over sets of descriptors of constituents, considering each multi-component material as a mixture system. A Python package was developed to calculate mixture descriptors based on this arithmetic equation, which can be used in machine learning-based QSAR and QSPR models.
The Quantitative Structure-Activity Relationship (QSAR) approach for predicting the biological activity and physicochemical properties of mixtures is gaining prominence, driven by the growing demand for highly engineered materials designed for specific functions. Developing mixture descriptors that effectively capture the intricacies of multi-component materials presents a significant challenge due to their structural complexity. We implemented a series of existing and new mixing rules to drive the mixture descriptors and develop mixture-based-QSAR (mxb-QSAR) models. We evaluated 12 additive mixture descriptors, and a novel non-additive combinatorial descriptor derived from the Cartesian product. These descriptors were used to model the fouling release (FR) property of 18 silicone oil-infused PDMS coating polymers by characterizing the removal of Ulva. linza. Various linear and nonlinear mxb-QSAR models were obtained using these 13 mixture descriptors. The best model, derived from the newly proposed Cartesian-based combinatorial mixture descriptors, employed a decision tree in combination with a two-stage feature importance feature selection. This model achieved a coefficient of determination R2 of 0.987 for both training and test sets, along with a cross-validation Q2 LOO of 0.791. The success of the nonlinear model and combinatorial descriptors underscores the significance of complex relationships among variables, as well as the synergistic effects of the components on fouling release properties.
In data science and machine learning, efficient and scalable algorithms are paramount for handling large datasets and complex tasks. Classification algorithms, in particular, play a crucial role in a wide range of applications, from image recognition and natural language processing to fraud detection and medical diagnosis. Traditional classification methods, while effective, often struggle with scalability and efficiency when applied to massive datasets. This challenge has driven the development of innovative approaches that leverage modern computational frameworks and parallel processing capabilities. This paper presents the Bison Algorithm, applied to classification problems. The algorithm, inspired by the social behavior of bison, aims to enhance the accuracy of classification tasks. The Bison Algorithm is implemented using PySpark, leveraging the distributed computing power to handle large datasets efficiently. This study compares the performance of the Bison Algorithm on several dataset sizes using speedup and scaleup as the performance measure.
Multi-component materials/compounds and polymeric/composite systems pose structural complexity that challenges the conventional methods of molecular representation in cheminformatics, which have limited applicability in such cases. Therefore, we have introduced an innovative structural representation technique tailored for complex materials. We implemented different mixing rules based on linear and nonlinear relationships’ additive effect of different components in composites treating each multi-component material as a mixture system. We developed and improved mixture descriptors based on 12 different mixture functions grouped into three main categories: property-based descriptors, concentration-weighted descriptors, and deviation-combination descriptors. A python package was developed for this purpose, allowing users to compute 12 different mixture-descriptors to use as input for the generation of mixture-based Quantitative Structure-Activity/Property Relationship (mxb-QSAR/QSPR) machine learning models for predicting a range of chemical and physical properties across various complex systems.
Chest X-ray imaging plays a vital role in the treatment of respiratory diseases such as pneumonia. Recent technological innovations have significantly improved the efficiency of the image analysis process especially Artificial and convolutional neural networks. However, there is a need to improve the accuracy. Thus, it is important to develop an automated, early diagnosis system that can deliver quick decisions and significantly lower diagnosis error. Recent advancements in emerging Artificial Intelligence approaches, particularly Deep Learning algorithms, have made the chest X-ray screening a viable option for early Pneumonia detection. Therefore, this paper focuses on using Teaching Learning Based Optimization (TLBO) with Convolutional Neural Network (CNN) to improve the accuracy of detecting pneumonia in Chest X-ray images. The research study was conducted on the chest X-ray images of the pneumonia data set and is compared with previous work. Results confirm that TLBO with CNN is a good choice for detecting pneumonia at the accuracy of 98.88% and is an improvement over benchmark studies.
Developments in technology facilitate the use of machine learning methods in medical fields. In cancer research, the combination of machine learning tools and gene expression data has proven its ability to detect cancer patients. However, processing such high-dimensional and complex data is still a challenge. This paper analyzed the impact different dimensionality reduction techniques have on machine learning models used for cancer prediction. Dimensionality reduction techniques such as principal component analysis (PCA), PCA with a kernel, and autoencoder were utilized to reduce the dimensionality of the RNA sequencing data. Two machine learning classifiers, namely neural network and support vector machine, were trained and tested using the original, dimensionally reduced, and cancer-relevant data. Various metrics, such as accuracy, precision, recall, F-Measure, receiver operating characteristic curve, and area under the curve, were used to assess the performance of classifiers. The results showed that dimensionality reduction positively affects the performance of the classifiers. Additionally, autoencoder performed better than PCA and PCA with a kernal. These findings indicate the potential of dimensionality reduction in improving the analytical results of machine learning classification models on high-dimensional data.
Geochemical maps are of great value in mineral exploration.Integrated geochemical anomaly maps provide comprehensive information about mapping assemblages of element concentrations to possible types of mineralization/ore,but vary depending on expert's knowledge and experience.This paper aims to test the capability of deep neural networks to delineate integrated anomaly based on a case study of the Zhaojikou Pb-Zn deposit,Southeast China.Three hundred fifty two samples were collected,and each sample consisted of 26 variables covering elemental composition,geological,and tectonic information.At first,generative adversarial networks were adopted for data augmentation.Then,DNN was trained on sets of synthetic and real data to identify an integrated anomaly.Finally,the results of DNN analyses were visualized in probability maps and compared with traditional anomaly maps to check its performance.Results showed that the average accuracy of the validation set was 94.76%.The probability maps showed that newly-identified integrated anomalous areas had a probability of above 75%in the northeast zones.It also showed that DNN models that used big data not only successfully recognized the anomalous areas identified on traditional geochemical element maps,but also discovered new anomalous areas,not picked up by the elemental anomaly maps previously.
Internal corrosion is a major concern in ensuring the safety of transmission and gathering pipelines in Structural Health Monitoring (SHM). It usually requires numerous sensors deployed inside the piping system to comprehensively cover the locations with high corrosion rates. This study presents a hybrid modeling strategy using Computational Fluid Dynamics (CFD) and Genetic Algorithm (GA) to improve the sensor placement scheme for corrosion detection and monitoring. The essence of the proposed strategy harnesses the well-validated physical modeling capability of the CFD to simulate the oil-water two-phase flow and the stochastic searching ability of the GA to explore better solutions on a global level. The CFD-based corrosion rate prediction was validated through experimental results and further used to form the initial population for GA optimization. Importantly, fitness was defined by considering both sensing effectiveness and cost of sensor coverage. The hybrid modeling strategy was implemented through case studies, where three typical pipe fittings were used to demonstrate the applicability of the sensor layout design for corrosion detection in pipelines. The GA optimization results show high accuracy for sensor placement inside the pipelines. The best fitness of the U-shaped, upward-inclined, and downward-inclined pipes were 0.9415, 0.9064, and 0.9183, respectively. Upon this, the hybrid modeling strategy can provide a promising tool for the pipeline industry to design the practical placement.
Copy-move forgery is one of the most used manipulations for tampering with digital images. The authenticity of the image becomes more crucial when the images are used in important processes. keypoints-based algorithms have been reported to be very effective in revealing copy-move evidence due to their robustness against various attacks. However, these approaches sometimes fail to make good prediction because of different factors such small number of keypoints detected, or wrongly detected keypoints. Matching the correct keypoints and filtering the wrong keypoints are other difficult tasks. One reason behind these issues is the parameters used to configure the key point detection algorithm. In this paper, another CMF (copy-move forgery) detection algorithm is proposed, by applying particle swarm optimization to find the best parameters for the algorithm for all different phases. Furthermore, filtering is achieved through two stages to remove most of the wrong keypoints detected. Additionally, triangulation is used as another technique applied to the algorithm in order to increase the detection area. Experimental results shows that the algorithm has good performance.
Abstract Epilepsy is a chronic neurological disorder that is caused by unprovoked recurrent seizures. The most commonly used tool for the diagnosis of epilepsy is the electroencephalogram (EEG) whereby the electrical activity of the brain is measured. In order to prevent potential risks, the patients have to be monitored as to detect an epileptic episode early on and to provide prevention measures. Many different research studies have used a combination of time and frequency features for the automatic recognition of epileptic seizures. In this paper, two fusion methods are compared. The first is based on an ensemble method and the second uses the Choquet fuzzy integral method. In particular, three different machine learning approaches namely RNN, ML and DNN are used as inputs for the ensemble method and the Choquet fuzzy integral fusion method. Evaluation measures such as confusion matrix, AUC and accuracy are compared as well as MSE and RMSE are provided. The results show that the Choquet fuzzy integral fusion method outperforms the ensemble method as well as other state-of-the-art classification methods.
Ant Colony Optimization is one of the most used methods applied to complex problems. Image processing is a difficult task in particular when complex images are involved. This paper uses a fuzzy euclidean metric as a distance measure between pixels from two images. This is the first evaluation of Ant Colony Optimization image edge detection in the context of fuzzy index and fuzzy euclidean metrics. The Canny edge detection is considered as the ground truth when evaluating similarities between the considered fuzzy images including medical ones. Experiments were run and successful comparisons were conducted using existing data sets as well as well-known non-fuzzy similarity metrics such as Jaccard’s index, Dice’s coefficient and the Pratt’s Figure of Merit were applied.
With the help of an electroencephalogram (EEG) the electrical activity of the brain is measured, and this can help identify chronic neurological disorders such as epilepsy. Epileptic episodes are detected by monitoring patients in order to provide preventive measures. Current research studies are using a combination of time and frequency features to recognize epileptic seizures automatically. In order to automatically detect epileptic seizures, different machine learning approaches have been used. Gradient boosting decision tree (GBDT) is a machine learning technique that is known for its efficiency, accuracy, and interpretability. In terms of performance of GBDT, many machine learning tasks such as multi-class classification, learning to rank, etc. have reported competitive performance. In this paper, epileptic seizure recognition data is investigated and split into a binary and multi-class data set for which the GBDT method is applied. In addition, the SHAP (Shapley Additive Explanations) method is used as an explanation tool to interpret the machine learning models that are produced via training for both the binary and the multi-class data set.
An Intrusion Detection System (IDS) is a system that protects against network attacks. This protection is achieved by monitoring the activity within a network of connected computers in order to analyze and predict the activity for intrusions. In the event that an attack would happen, the system would respond accordingly. In the past, different machine learning techniques have been proposed, which can be broken into clustering algorithms and classification algorithms. In this chapter, the CICIDS2017 data set is investigated, which contains benign and the most up-to-date common attacks resembling true real-world data. A machine learning approach is chosen whereby a comparison between a deep neural network approach and an ensemble method called super learner is performed. Furthermore, other algorithms such as gradient boosting machine, distributed random forest, and the XGBoost from the AutoML library are also compared.
While early detection of diseases helps in managing and improving patient outcomes, most detection methods employed today are largely manual, costly, and time-consuming. Accordingly, computer-aided diagnosis is emerging as an innovative solution to improving the accuracy of detection by eliminating human errors and lowering the cost of diagnosis. One of the diseases that can benefit immensely from computer-aided diagnosis is pneumonia, which is an acute pulmonary infection accounting for thousands of hospitalizations and deaths globally. Current pneumonia detection approaches entail manually examining radiology images such as X-rays. Because of subjective variability, the outcomes of the examination are not always accurate. As a result, researchers have started to develop models based on machine learning to aid in detecting pneumonia based on chest X-ray images. Most of the models developed are based on deep learning, especially convolutional neural networks. However, these models require vast data sets for training and their accuracy values can be improved. For that reason, this paper developed a detection model based on Reinforcement Learning (RL) with convolutional neural network (CNN). The chest X-ray images of pneumonia is a data set that is used for the experiments. The obtained results confirm that applying the RL model is a good choice for detecting pneumonia. The efficacy of these model's performance was evaluated by measuring the precision, recall, F1-score, accuracy, and confusion matrix.
Domain Name System (DNS) is the Internet's system for converting alphabetic names into numeric IP addresses. It is one of the early and vulnerable network protocols, which has several security loopholes that have been exploited repeatedly over the years. The clustering task for the automatic recognition of these attacks uses machine learning approaches based on semi-supervised learning. A family of bio-inspired algorithms, well known as Swarm Intelligence (SI) methods, have recently emerged to meet the requirements for the clustering task and have been successfully applied to various real-world clustering problems. In this paper, Particle Swarm Optimization (PSO), Artificial Bee Colony (ABC), and Kmeans, which is one of the most popular cluster algorithms, have been applied. Furthermore, hybrid algorithms consisting of Kmeans and PSO, and Kmeans and ABC have been proposed for the clustering process. The Canadian Institute for Cybersecurity (CIC) data set has been used for this investigation. In addition, different measures of clustering performance have been used to compare the different algorithms.
William Naylor合作论文数The University of Bath9
Camelia-Mihaela Pintea合作论文数Tech Univ CJ3