
Microarray technology provides an enormous opportunity to measure large-scale gene expressions simultaneously. However sparse and high-dimensional feature space posed significant challenges during data analysis, mainly in learning network structure. Hence, feature selection (FS) has become an essential phase in microarray data analysis to obtain significant genes that could enhance the performance of subsequent process. This study aims to propose a hybrid FS methods based on Binary Bat Algorithm (BBA) and Dynamic Bayesian Network (DBN). The proposed method is tested on cancer gene expression dataset that is publicly available. The fold change analysis is conducted to measure a gene expression level between two diverse conditions prior to subsequent process of FS. Experimental results show that BBA has achieved better results compared to other baseline methods when trained with DBN low-order conditional independence and DBN full-order conditional independence with accuracy of 89.1
Clustering is a standard paradigm for associating some similar objects to a cluster while associating others to their respective clusters in an unsupervised manner. Different techniques are proposed in the domain of clustering depending upon the underlying methodology slightly vary from one another. Some of these are regarded as Hierarchical Clustering, K-means Clustering, Spectral Clustering, Affinity Propagation and Density based spatial Clustering. Their strengths and weaknesses are different for different datasets. In this research, we have used three benchmark datasets to analyze the merits and demerits for each of the clustering techniques. After applying these algorithms to the datasets named letter, glass and wine, we presented the critical review on why a specific clustering algorithm lacks with respect to execution time and clustering quality measured by Silhouette’s coefficient and why an algorithm performs better than others on the same dataset.
Pakistan is one of the major coal producing country and also have the most serious coal mine accidents around the world. The proposed study performs experiments on a dataset contain reviews based on accident reasons occurred during mining work to reduce the accident rate. The aim of this study achieved by categorizing the reasons into different classes on the behalf of human behavior, roof dropping, and smoke inhalation. Then perform preprocessing on the reviews to clean the data. After preprocessing, the bag-of-words and TF-IDF are used singly and in combination to preserve meaningful information in extracted feature form. Finally, the Random Forest, Naive Bayes classifier, SVM, Decision Tree, Logistic Regression and proposed ensemble LSD (LR+SVM+DT) models are used to classify the accident reasons to analyze the most occurring once. The performance of the proposed approach is evaluated using average accuracy, precision, recall, and f1 score. The experiments reveal that a combined feature improved the performance of machine learning models. Ensemble LSD outperforms the other models and achieves 94
The biggest challenge in Electronic Commerce (E-Commerce) is understanding the demand of the market and customer perception regarding the products of online stores. Online user evaluations are faster than a direct poll for obtaining customers' perceptions. To predict the quality of service provided, sentiment analysis methodology is used, which identifies and classifies product reviews into positive and negative sentiments through different classifying algorithmic models. This study has used Random Forest, Logistic Regression, and Gaussian Naive Bayes to perform sentiment analysis to analyze customers' likes and dislikes. The “Clothing Products Reviews” dataset from ‘Kaggle’ is used. Due to the significant number of positive sentiments, the results showed the dependability dimensions that cause the biasness in results. Synthetic Minority Oversampling Technique (SMOTE) is applied to balance the positive and negative sentiments. Random Forest, Naïve Bayes, and KNN classifiers are combined into one Stacking Classifier to improve the results and get 97
The COVID-19 outbreak is moving individuals globally. Monitoring social media and internet news is now vital to grasping this phenomenon and its impact. This study’s goal is to offer a method for capturing important concepts and themes addressed in the mainstream media and social networks and then apply it to the COVID-19 outbreak. This study compares articles and news, then visualizes the evolution and influence of the COVID-19 epidemic articles. The COVID-19 articles dataset from Kaggle is utilized for experiments. Various articles use the dataset to examine COVID-19 reviews. The studies use datasets and models such as Decision Tree (DT), Logistic Regression (LR), and Extra Tree Classifier (ETC), with F1 score, precision, and recall being evaluated. To improve accuracy, Term Frequency-Inverse Document Frequency (TF-IDF) is applied to the extracted keywords. The assembled model enhances this research, and the Logistic Regression (LR) model provides the highest accuracy at 90%.
Facial Expression Recognition (FER) has applications in various areas, including human behavior, human-machine interaction, mental disorder detection, human psychology, mood and response behavior, better interpersonal relationship, effective communication, video surveillance, face detection tracking system, and real-time emotion detection system, etc. Therefore, processing and analysis of human faces is an active research area for human expression recognition. In this research work, through the parsing of face, a framework for human facial expression recognition has been proposed, in which seven facial expressions have been recognized, i.e., happy, sad, neutral, anger, disguise, fear, and surprise. A light-weighed CNN based model has been designed and tested on publically available datasets. To verify the robustness of the proposed model, several experiments have been performed. The proposed light weighted CNN based model achieved 95
In this study, a method is proposed in which accurate prediction at early stage is important of cardiac patient for efficiently treating. The novelty of this research is using feature extraction with the help of some techniques of signals processing. Ten different kinds of sensors are used are metal oxide semiconductors that are used for sensing the different gasses that are emanating from the human body. Moreover, ECG, SPO2 and oxygen sensors are used for further processing. Various experiments are performed that identify 5, 10, 15 and 20 subjects every subject is identified and scanned as 1000 different features. The signals that are received are analogue and with the help of Arduino they convert to digital signals. An architecture is trained on the dataset that is developed. Sensitivity, f-measures, accuracy and specificity are the standards that are used for the evaluation of the model that is proposed as identification of human odour. The accuracy for this model is more than 85%.
Ordinary linear regression models have been widely implemented to measure the causal relationship between exogenous factors and purchase decision. While inconsistent samples are not yet investigated by previous studies through these models. In this paper, we are interested to discuss a purchase decision model by ordinary regression using data reduction strategy of rough sets in handling these samples type. The primary data and information were collected from 265 random customers of the textile stores in Padang City, Indonesia. The results showed regression model has better performance after data reduction in selecting purchase’s factors. In this case, the customer’s psychology is the most significant factor related to decision making in purchasing of textile material whether before and after data reduction. The proposed rough-regression model is very appropriate to support categorical data analysis with many uncertainty factors, especially in cross-sectional survey research types.
There are common factors that characterize individuals who commit crime. Statistically, relationship between factors could be measured and formed using regression models. While modeling crime rates was widely approached using the ordinary regression model. However, this model is not much capable for categorical response variables such as crime type. In this paper, we are interested to classify crime type using some independent variables using multinomial logistic regression and K -Nearest Neighbor ( K -NN) models. While both are powerful models for classification purposes. The secondary crime data was collected from 2019 (pre-pandemic) and 2020 (during a pandemic) in the police office of Payakumbuh Region, West Sumatra, Indonesia. The results indicate that the crime type was influenced by employment status (unemployed persons) and time occurring (daytime) for both years periods. In the testing data phase, the average of accuracy levels of multinomial logistic regression and K -NN are 66.81 K -NN model is better approach to be used for the prediction and classification of crime type if compared with multinomial logistic regression. Both models could be considered for supporting the police divisions on decision making and prevention strategy.
Android is one of the popular open-source mobile operation system. The popularity of Android attracts the malware developer to attack the Android operating system to get credential information for bad purposes using malware. There is numerous method available to detect the presence of Android botnet using static and dynamic analysis. Static analysis unable to monitor the botnet behavior during runtime. Therefore, we use dynamic analysis by proposing android botnet detection technique using network analysis which will be able to capture behavior of android application during runtime. The benign application dataset is taken from APKPure and malware application dataset from Koodous and Github. We extracted five features that were captured during network traffic analysis using Wireshark and SSL Packet Capture. The network traffic analysis approach shows promising result with accuracy value 76.5
Air pollution is mixed particles and gases which could harm towards our health, especially towards our respiratory system. The polluted particles can exist in both outside and inside surroundings. Furthermore, chronic diseases and cancer are also associated with air pollution where it could oxidative stress and inflamed human cells. Thus, this study serves to assist public community to be aware of the air quality in their surroundings to reduce risks of having respiratory problems. Data visualization through AirAwareMalaysia dashboard is produced by using data analytics and prediction technique which could help the users to see and identify the low air quality area. Interactive features are implemented where users able to click buttons to view the air quality. The Cross-Industry Standard Process for Data Mining (CRISP-DM) model is used where data analytics and visualization techniques are applied. Finding shows that the implementation of data analytics, data visualization, and data prediction analysis are key features in assisting users to develop surrounding awareness towards air quality. Finally, the study could achieve the following objectives: 1) To identify data requirements to visualize the air pollution index in Malaysia based on air pollutants parameters for AirAwareMalaysia 2) To design a user interface for web visualization dashboard for AirAwareMalaysia and 3) To develop visualization dashboard using the datasets for AirAwareMalaysia.
This paper assessed the implementation of trigonometric basis function as input enhancement for Functional Link Neural Network (FLNN) trained with Ant Lion Optimizer (ALO) learning algorithm. The previous work of FLNN trained with ALO model used the tensor model to introduce nonlinearities in its inputs features enhancements. One of the major concerns of using the tensor model is that when FLNN has large number of input features, the network may contain many higher-order terms which could lead to combinatorial explosion in the number of weights as the order of the network becomes excessively high. To avoid this, trigonometric basis function in implemented in the model. The result on classification performance made by FLNN with trigonometric basis architecture and FLNN with tensor model architecture both trained with ALO were carried out. From the result achieved, the implementation of the FLNN with trigonometric basis trained with ALO performs the classification task quite well and yields better accuracy on the unseen data.
The nature inspired algorithms have motivated the practitioners to solve complex real-world problems. These algorithms are more capable to approach the optimal solution faster than conventional methods. The proposed algorithm uses the exploration capability of the Improved Flower Pollination algorithm with dynamic switch probability and swap operator (IFPDSO) and exploitation capability of Pattern Search (PS) to approach the optimal solution efficiently. The hybridization of IFPDSO and Pattern Search (IFPDSO-PS) has been validated on various benchmark functions and compared with other hybrid algorithms to evaluate its better performance.
Clustering is an effective technique for identifying patterns and structures in labeled and unlabeled datasets in the medical sector. Density-based clustering is a sophisticated machine learning technique for identifying distinctive patterns in large datasets. However, this approach has certain drawbacks like inability to determine local densities and overlapping clusters or clusters with blur boundaries. This paper embeds fuzzy logic with density-based clustering, for improved clustering separability. In order to validate the usability of the proposed approach, we use five real-world datasets belonging to medical domain.
The Artificial Neural Network Autoregressive model (ANN-AR) is a recently adopted approach for forecasting Naira to the USD exchange rate. The Bayesian Regularized Neural Network (BRNN) is an alternative to ANN-AR based on the probabilistic interpretation of network weights. It is useful for solving overfitting problems inherent in ANN when large historical data are not available. In this paper, we developed a BRNN for modelling the monthly time series data of Naira to USD for four years. Performance analysis was observed using actual exchange rate values for the Year 2017 against the model predicted outcomes based on four evaluation metrics. Results from the analysis established the appropriateness of a BRNN for modelling short-term exchange rates in Nigeria.
The vagueness and hesitancy from imprecise information in multi-criteria decision-making problems can be solved using an intuitionistic fuzzy environment. This paper presents a new TOPSIS method for ranking alternatives that is combined with intuitionistic fuzzy sets (IFS). Previous research suggests that the intuitionistic fuzzy weighted averaging (IFWA) operator was employed to combine judgments of all experts and also subjective evaluation to weight the criteria. In this paper, the role of IFWA operator is preserved while the criteria weights are obtained using the entropy measure of IFS. The distance between positive and negative ideal solutions is calculated using Euclidean distance. The alternatives’ ranking order is concluded, based on the obtained values of relative closeness coefficient. A numerical example and comparable results demonstrate the proposed approach's potential in decision-making problems.
XYZ company is a telecommunications company that offers a variety of services to Indonesians, including high-speed internet connection via fiber optic lines. The customer data for 1.000 subscribers will be utilized to conduct an analysis of the disruption generated by the communication network or the main device and other supporting devices connected to the internet network in this study. The classification of data in certain classes will be known using the Naive Bayes Classifier Algorithm, and the results of the classification will be used as a solution to calculate the interference that frequently occurs, namely code 1035 interference caused by customer data not having internet service, voice service, or IPTV service. The code 1054 interference was caused by a mismatch of the customer’s active device in the system during the transition or migration from Copper Cable to Fiber Optic (GPON00-GPON05). The probability of interference code 1035 is TP Rate = 0.988, FP Rate = 0.033, Precision Recall = 0.988, F-Measure = 0.988, MCC = 0.971, ROC Area = 1.000 and PRC Area = 1.000. And accuracy by class code 1054 is TP Rate = 0.984, FP Rate = 0.007, Precision Recall = 0.985, F-Measure = 0.982, MCC = 0.979, ROC Area = 1.000 and PRC Area = 1.000.
The speech emotion recognition is a challenging and an exigent task in the field of data science. Existing studies have only focused on one-dimensional Convolutional Neural Network (CNN) architecture for speech emotion recognition. This one-dimensional architecture’s speech recognition accuracy is low when dealt with RAVDESS, TESS and URDU datasets using non-optimal parameters. To overcome this problem, this research work proposed an efficient two-dimensional CNN architecture with an optimized combination of parameters to achieve better accuracy. The proposed method is compared with Support Vector Machine (SVM) and one-dimensional CNN using RAVDESS, TESS and URDU datasets based on accuracy. Based on the conducted experiments, it can be seen that, the proposed method has outperformed with an accuracy of 76.08
With the rapid progress in space technology development, a lot of digital images are transmitted back to ground telescope receiver. Thus, image security plays a significant role during the transmission process. Existing algorithms in image security were developed based on Latin squares, DNA sequence etc. However, the suitability of the techniques is questionable and high cost was needed. In this paper, Combined Spatial and Frequency Domains (CSFD) algorithm is proposed. It discusses the RGB color image security algorithm based on the images combined spatial and frequency domain. Fourier transform is applied to an RGB color telescope image before it is processed in spatial domain. Then, the R, G, and B components are extracted out before they are encrypted by two levels of encryption algorithm. The outputs were evaluated using NPCR, UACI, PSNR, and information entropy. The results showed that this algorithm has better performance compared to other recently developed algorithms.
The World Health organization estimates that heart disease is liable for the death of 12 million people per year throughout the world. Nearly half of all deaths are caused by cardiovascular disease each year. The earlier cardiovascular disease may be detected and treated, the lower the risk of significant outcomes for individuals who are already at risk. Developing a better heart disease prediction model can be extremely beneficial in predicting cardiovascular disease risk more accurately. Authors developed Gated Recurrent Unit (GRU) and a random forest (RF) based heart disease prediction model in this study to weed out risk factors for the heart disease. After screening out the main attributes, the proposed approach GRU-RF is tested against traditional Deep Neural Network, and K-nearest neighbour algorithms. The prediction accuracy of proposed approach is 87