
The recent surge in the abundance of fake news appearing on social media and news websites poses a potential threat to high-quality journalism. Misinformation hurts people, society, science, and democracy. This reason has led many researchers to develop techniques to identify fake news. In this paper, we discuss a stance prediction technique using the Deep Learning approach, which can be used as a factor to determine the authenticity of news articles. The Fake News Stance Prediction is the process of automatically classifying the stance of a news article towards a target into one of the following classes: Agree, Disagree, Discuss, Unrelated. The stance prediction task's input is the news articles containing a pair: a headline as the target and a body as a claim. This paper proposes a deep learning architecture using Bi-directional Long Short Term Memory and Autoencoder for stance prediction. We illustrate, through empirical studies, that the method is reasonably accurate at predicting stance, achieving a classification accuracy as high as 94%. The proposed stance detection method would be useful for assessing the credibility of news articles.
With the advent of computer vision, there has been significant growth in the research and development of facial recognition based automated attendance systems. Although current systems have been successful in alleviating human interaction and manual efforts, there still exist several challenges such as severe misclassifications, undetectable face angles, and different lighting conditions which result in a drastic drop in the accuracy. The system introduced in this paper has achieved an overall accuracy of 93.33%. A concept termed the “two-tier authentication” method has been developed to improve the overall accuracy of the system and to integrate a mechanism of time allowance for students. This method facilitates granting attendance to students based on the number of recognized faces as well as the probability of each prediction allowing for a more robust method of marking attendance to students. The novelty of this approach is to introduce an accurate statistical sequence for the execution of a proxy-free automated attendance system that employs state-of-the-art algorithms. Composed of 3 distinct parts, every sub-system performs a specific task namely, face detection, generation of face embeddings (FaceNet), and face classification. A comparative study was carried out to select the most appropriate detection and classification algorithms where Faster R-CNN and Support Vector Classifier outperformed their corresponding competitors respectively.
Cancer develops when cells in any part of the body start to grow out of control. It can spread to other parts of the body. Melanoma is a type of skin cancer that is developed when melanocytes i.e. cells which produce melanin (the pigment which is responsible for the perceived color of skin) begin to grow out of control. Melanoma is dangerous as it has a high tendency to spread to other parts of the body, if not detected early and left untreated. In this paper, we use deep learning techniques to build a classification system to categorise a skin lesion into malignant and benign. This system relies on a dataset which consists of skin lesion images from various sites on the body. We augment the dataset using appropriate transformations and evaluate the classification system using various metrics. The different models used in this implementation are compared based on the metrics to find the superior performing model. ResNet-50 as per the results of sensitivity, specificity and accuracy has the best results among the other three with values 99.7%, 55.67%, 93.96% respectively.
Most software employees are facing challenges on integrating programs written in different languages for implementing the techniques for their software development. This can be achieved by automating the conversion of programs using natural language programming techniques. This research presents a novel `Semantic Rule-based Automatic Code conversion System (SRACS)' that uses semantic layering, keyword identification, and a semantic rule-based constructor. The code snippets for `Hello World', `For Loop', `While Loop', `If else', `Factorial' and `Travelling Salesman Program' are converted from Java to Python and vice versa, and the accuracies are presented. An average accuracy of 71.57% is achieved for the conversion of the code snippets from Java to Python, and a 77.07% is achieved for Python to Java. The accuracy is based on the `accuracy in the conversion of the variables', `accuracy in the conversion of the attributes' and on the `proper indentation of the code in the target code'.
In Software Development Life-cycle, Verification and Validation plays a very important role, especially in the case of Safety-Critical Industries like Aerospace. Display dashboard consists of multiple static and dynamic objects having affine transformation, graphics overlap, shadows and less inter symbol discriminative features compared to natural images. Manual Software graphics verification is an error-prone and time-consuming activity. In this paper, we propose a novel software graphics verification pipeline to verify graphics symbols and alphanumeric objects as per Software requirements. To the best of our knowledge, our proposed approach is the first study on deep learning-based graphics symbol detection from complex synthetic background which requires high model accuracy. We experiment using Single-shot Multibox Detector (SSD) and You Only Look Once (YOLO v2) to detect different Graphical symbols from display simulator real-time captured video frames. These detected objects are further classified based on their nature. Objects containing alphanumeric digits can be recognized using Optical Character Recognition and dynamic symbols are detected using object detection to infer other properties. Finally, all the extracted properties can be compared with test expectations to verify their correctness. The result shows superior accuracy of the SSD algorithm over other state-of-the-art object detection algorithms for detecting real-time graphics symbols.
Cricket has the second-largest fan-base after football. Interest in any game is a factor of quality of the game which in turn depends on the quality of players. It is therefore important to have good players and that they are paid well. Sports industry largely relies on the advertising sector for sponsorship and financing of games. Advertisement companies spend a fortune to acquire the best slots during a game to catch the maximum viewership. This implies that advertising companies have a lot of interest in the duration of a match. Indian Premier League (IPL) has a huge fan-base and is one of the major events where companies spend a large amount of money to advertise their products. Due to this, a short game, which ends prior than expected, results in loss of opportunity in terms of time-slots lost and hence revenue and fan interest. The prediction of duration of a game will be beneficial for both sport and advertisement industry. In this paper, we use machine learning algorithms to predict the duration of a match in terms of the number of balls expected to be delivered in the match. The work introduces four different approaches, using historical data, to predict the number of balls in a match.
In the typical use case of browsing Linked Data in DBpedia, the user would find an average of 180 facts attached to each entity. These facts are ordered alphabetically based on predicates, but a logical ordering of these facts is a better option. In this article, we present a Nexus based predicate ranking of Linked Data facts named NPRank. The key idea of NPRank is, the importance of a predicate is directly proportional to its familiarity among its group called Nexus. NPRank is a language and endpoint independent model allowing seamless integration and querying of data from multiple endpoints. Nexus score generated to rank predicates also assists in fragmentation of large data and bring in more hidden data from the SPARQL endpoints. Our experiments, conducted with the ranking of the Linked Data facts, corresponding to most visited pages of Wikipedia; from 275 active SPARQL endpoints, achieves better performance than the state-of-the-art methods.
This paper introduces a QvLA-RACH access scheme to reduce resource wastage by enhancing slot utilization in cellular M2M communications. The proposed QvLA-RACH scheme employs a Q-value update technique to reduce idle slots by controlling the probability of collision. the scheme also uses a cooperative Q-Learning strategy to update the Q-value. The performance of the proposed scheme is evaluated compared to the existing scheme using extensive simulation. The results show that the proposed QvLA-RACH scheme achieves better performance in terms of throughput and access delay by 49.7% and 18%, respectively.
This paper deals with the design of a model based on Adaptive Neural Fuzzy Inference System (ANFIS) for the prediction of early age (3 days) compressive strength of concrete. The model is generated by a dataset having 8 parameters. these are converted into 7 inputs viz. Cement, Flyash, BFS, water, Superplasticizer, Coarse aggregates, and fine aggregates and 1 output i.e. 3 days compressive strength. The model was trained and tested using hybrid method of learning. The results produced Training and checking errors as 0.153MPa and 1.212MPa respectively making ANFIS very much appropriate for this purpose.
Stein's Unbiased Risk Estimate (SURE) is considered as an indirect method for predicting Mean Squared Error (MSE) in the absence of ground-truth, as its computation requires only noisy observation and denoised image. SURE is usually used as an objective function for optimizing the operational parameters of denoising algorithms, adequate for real-time images. Hence, a close analysis of the performance of SURE on standard test images is worthy of investigation. Pearson's Correlation (r) of SURE with Mean Absolute Error (MAE) between denoised images and ground-truth is analyzed in this paper, on Shepp-Logan Phantom and simulated Magnetic Resonance (MR) images, at different noise levels. Denoised images which differ in terms of MAE against ground-truth are produced by varying the standard deviation of a Gaussian smoothing kernel $(\mathbf{0.01\leq\sigma\leq 0.04})$ of fixed dimension, $\mathbf{9\times 9}$. Values of correlation between SURE and MAE on Shepp-Logan and simulated MR images are $\mathbf{r}=-\mathbf{0.99\pm 0.02}$ and $\mathbf{r}= \mathbf{0.48\pm 0.36}$, respectively. Concordance of SURE with MAE is observed to be poor on simulated MR images, especially at higher noise levels. SURE is suitable for optimizing the parameters of denoising kernels only when the underlying function used to compute the kernel is fully differentiable by the noisy observation.
Accent is a distinctive way of pronouncing a language, especially one associated with a particular country, area or social class. While dialects are usually spoken by groups united by geography or class and are a variety of same language differing in vocabulary and grammar as well as pronunciation, the accent is how the same language is spoken differently by people of different ethnicities. Accent plays a vital role when it comes to speech recognition by voice assistance systems. The modern-day voice assistant systems tend to misinterpret the words spoken by a person with a strong accent influenced by his native language. Machine learning algorithms, especially Support Vector Machine (SVM) and Random Forest when applied on a proper training set can play a vital role in the classification of accent. We propose in this paper the learning framework for speech recognition of Indian accent by analysing the features of Indian accented English and classifying based on sounds that are typical to Indian accented English.
Smart energy management is a major area of interest to meet the rising energy demand for which several countries are deploying smart meters. Presently, there is a need to better visualize the high-volume of data captured by smart meters to provide a means to effectively gather various analytical insights which can help in better understanding the energy usage patterns. This article presents a cascade application of two competitive learning algorithms - Self-organizing Map (SOM) and K-means clustering, to discover knowledge from smart meter data. A SOM is applied to construct a 2-D topologically preserving map which is useful in understanding and visualizing the consumer load profiles. Then K-means is applied on the codebook vectors of the SOM to determine the clusters containing consumers with similar energy consumption patterns. The identified consumer clusters enable the utility firms in preparing segment-specific tariffs to efficiently shape the future energy usage patterns.
Surface remeshing intends to yield high quality, high regularity and low complexity meshes that are geometrically faithful to original models and free from the low-quality elements. Unfortunately, attaining balance between quality, regularity, complexity and approximation error becomes tedious during remeshing. In the work presented here, authors propose a surface remeshing techniques based on quadric based simplification and high-quality approximation that tries to attain balance of remeshing goals. Given a triangular mesh and user assigned approximation error, the output mesh (remesh) achieves a higher minimal interior angle and low mesh complexity with implicit feature preservation. The proposed approach recapitulates in two-steps. First, the mesh complexity is optimized to lessen the extent of vertices, and then followed by enhancement of the quality of the elements using local operators. This approach can be efficiently incorporated in preprocessing for numerous applications. Investigations have demonstrated that the proposed approach is efficient and robust. The results of proposed approach attain high quality triangles along with preservation of the features of the original geometry.
With the rapid increase in the dependency on technology and the internet in our personal and professional life, the computer networks have become very congested, and the frequency of presence of an intrusion in a network has also increased. An active IDS (Intrusion Detection System) protects the network from intrusions and provide security to the system. Machine learning techniques are the most efficient technologies to develop IDS as substantial network data can be easily trained and tested using ML models. Any general machine learning models work in three phases: Data pre-processing, Feature selection and training, and testing the developed models. The major contribution of this paper is the extraction of sub-optimal feature set from NSL-KDD data set having 41 features and then implementing different ML models to find the best suitable model using this set of features. It is observed that the ML models SVC and MLPClassifier performed better as compared to CNN in terms of complexity, accuracy and training time when trained and tested using the selected optimal feature set. CNN is an excellent deep learning algorithm that gives good results for image data perform better in comparison to simple text data machine learning models like SVC and MLP Classifier. MLP Classifier gave a higher accuracy of 98.19%.
Glaucoma detection is a significant problem to be solved in medical field. Few research works have been designed to detect glaucoma in its early stage. But, performance of glaucoma disease detection using existing techniques was not effectual. Moreover, time complexity of conventional glaucoma disease detection was more. In order to overcome such limitations, Damped Least-Squares Recurrent Deep Neural Classification (DLRDNC) Technique is proposed. The DLRNL Technique designs DLS-Recurrent Deep Neural Classifier in order to increase the prediction performance of glaucoma disease at an early stage with minimal time. The DLRNL Technique conducts simulation process using metrics such as disease detection accuracy, disease detection time and false positive rate with respect to different number of fundus image. The simulation results depict that the DLRNL Technique is able to increase the accuracy and also reduces the amount of time required for glaucoma disease detection when compared to the state-of-the-art works.
Feature selection is one of the most important preprocessing steps in Machine Learning. This can be broadly divided into search based methods and ranking based methods. The ranking based methods are very popular because they need much lesser computational power. There can be many different ways to rank the features. One of the ways to measure effectiveness of a feature is by evaluating its ability to separate the classes involved. These interclass Separability based measures can be directly used as a feature ranking tool for binary classification problems. Bhattacharya Distance which is the most popular among them has been used majorly in a recursive setup to select good quality feature subsets. Jeffries-Matusita (JM) distance improves Bhattacharya distance by normalizing it between 0 and 2. In this paper, we have ranked the features based on JM distance. The results are comparable with mutual information, Relief and Chi Squared based measures as per experiments conducted over 24 public datasets but in much lesser time. JM distance also provide some intuition about the dataset prior to any feature selection or machine learning algorithm. A comparison has been done on classification accuracy and JM scores of these datasets, which can provide a good intuition on how good a dataset is for classification and point out the need of or lack of further feature collection.
Gait recognition is an expanding stream in biometrics, intended to recognize individuals through the investigation of their walking pattern. This pattern is obtained from a distance, without the active participation of the people. One of the difficulties of the appearance-based gait approach is to enhance the performance of frontal gait recognition, as it carries less spatial and temporal data when compared with other view variations. As a result, to increase the performance of the frontal gait recognition, this paper presents a method which uses two-step procedure; the Hierarchical centroid Shape descriptor (HCSD) and the similarity measurement. The proposed method was assessed on the broadly used CASIA A, CASIA B, and CMU MoBo gait databases. The experimental outcomes showed that the proposed method gave promising results and outperforms certain state-of-the-art methods in terms of recognition performance.
Research on Gene Regulatory Networks (GRN) is primarily guided by differential co-expression analysis. Through this approach, we can observe a biologically significant difference in gene regulatory control under varied metabolic conditions. Although quite successful, the traditional analysis gets restricted either from the point of view of a single gene-specific topology where only pairwise co-expression plays a major role or in the type of gene coexpression. In the latter context, state of the art examples mostly considers linear correlative regulatory patterns in a network. In the current work, we overcome these limitations using an established gene regulatory framework which accounts for the complete co-expression structure. An overall improvement on the regulatory functions could be observed incorporating nonlinear gene to gene regulation. In this regard, the nonlinear differential orchestrated action of genes in different conditions present in certain unreported gene regulatory pathways of p-53 dataset have been found to be significant, where linear measures of differential regulations reflected meager importance.
In this work, we propose a hybrid binary classifier which combines a decision tree with a support vector machine. The proposed hybrid model has the advantages of improved accuracy and easy interpretability. The model will be useful for feature selection cum classification tasks in real-world supervised learning problems. Numerical evidence is also provided using 25 standard data sets from various fields to assess the performance of the model. Performance of the proposed hybrid binary classifier is quite better when compared to individual classifiers.
We live in an era where technology and automation is impacting every domain possible, increasing efficiency, productivity, scalability and at the same time cutting down investments and efforts. IoT (Internet of Things) has given us the power of connecting everyday objects to internet and hence they can be controlled and monitored from anywhere in the world. However, traditional warehouses still remains an exception. Humans toiling their way through the day, controlling and monitoring all the environmental conditions of the warehouses, is the general Picture of warehouses around the globe today and that comes with a cost of human errors resulting in wastes of resources and decreasing efficiency. Therefore, we require an automated smart warehouse which not only provides multi-parameter monitoring and control but is able automatically vary the parameters according to the environment. Existing solutions fail to address this problem which are either impractical or offer partial solution. In this paper we present a smart warehousing system which through various sensors connected to a micro-controller performs not only multi-parameter monitoring by sending the data to the cloud but also provides automated control by processing the data in the server and sending results back to the micro-controller. Moreover, the system also provides with a website and an android application which present various graphs and give complete control and automation through a tap of a button.