
Optimization of searching the best possible action depending on various states like state of environment, system goal etc. has been a major area of study in computer systems. In any search algorithm, searching best possible solution from the pool of every possibility known can lead to the construction of the whole state search space popularly called as minimax algorithm. This may lead to a impractical time complexities which may not be suitable for real time searching operations. One of the practical solution for the reduction in computational time is Alpha Beta pruning. Instead of searching for the whole state space, we prune the unnecessary branches, which helps reduce the time by significant amount. This paper focuses on the various possible implementations of the Alpha Beta pruning algorithms and gives an insight of what algorithm can be used for parallelism. Various studies have been conducted on how to make Alpha Beta pruning faster. Parallelizing Alpha Beta pruning for the GPUs specific architectures like mesh(CUDA) etc. or shared memory model(OpenMP) helps in the reduction of the computational time. This paper studies the comparison between sequential and different parallel forms of Alpha Beta pruning and their respective efficiency for the chess game as an application.
The threat landscape is exponentially increasing which become more worsen when a greater number of emerging devices are connected to internet such as Internet of Things (IoTs), Embedded Systems, Cyber Physical devices etc. To control the damage of cyber threats, there is a need to monitor the cyber criminals continuously to understand the tools and technique used by the attackers in order to develop cyber defense mechanism to protect cyber Eco-systems. In this research, a multi-Honeypot platform as a tool is presented for cyber threat intel generation to implement the multiple classes of Honeypots such as Windows, IoTs, Embedded etc. Honeypot is widely used by the security researchers, security companies to understand the tools and tactics about the attackers but these are quite complex to deploy and maintain especially due to diverse set of IT systems and intensive resource requirements to deploy High Interaction Honeypots. This complexity is reduced in this research by implementation of Para-Virtualization based approach to enable multiple classes of Honeypot sensors of different platforms on a light weight hardware. It is addressed that time window to collect the data and to conclude it as a cyber threat intel with support of evidences should be probabilistically determined. After applying the analysis such as behavior analysis and deep learning methods to determine about the unknown threat patterns, the attack data sets are correlated into different cyber threat events and converted into an actionable cyber threat intelligence to disseminate the information in an automated manner. In the end, threat intelligence is generated and experiments are documented. The deep learning-based analysis inspired by neural networks is integrated in the design to determine the unknown classified threat events.
Migration of customer to other companies is “Churn”. Customers who abandon services of any company or leave the organization are known as churners. Now a day churn prediction is one of the biggest challenge for organizations/Companies. Reputation/Ranking, finance and growth plans of company is directly affected by churn. This makes research on churn more valuable. In this paper python platform is used to implement two of the classification models of supervised machine learning. As the data used is a labeled data thus this is under the category of supervised domain of machine learning. Confusion matrix is observed to state the conclusion which says that KNN is a better approach to predict customer churn over Logistic Regression. KNN is 2.0% more accurate than Logistic Regression to predict customer churn. Here “accuracy” parameter says that Logistic algorithm is least effective to predict the customer churn as compare to K-Nearest Neighbors.
In Computer vision, object recognition is a very important component and also very challenging. Intention of This paper is to exploit a high confidence object detection framework that boosts up the classification performance with less computational burden and cost efficient. Features are extracted from images by using Histograms of Oriented Gradients (HOG) technique and then for generating principle components as well as reducing dimensions Principal Component Analysis (PCA) has been applied on the extracted features. For classification of objects Support Vector Machine (SVM), Random Forest, Input mapped classifier, M5P classifier and Gaussian process classifier have been employed. A comparative study on performance of those approaches have been conducted. Moreover, for better clarification of the dataset, statistical and automated analysis have been considered. Overall findings demonstrates that, Principle Component Analysis (PCA)based Support Vector Machine (PSO) outperforms other approaches by depicting accuracy of 94.02% and highest F-Score measurement.
Rooftop solar energy potential has traditionally been estimated by surveying the number of large buildings in a given area. In this work, we propose a fast and low-cost method to estimate the rooftop photovoltaic solar energy generated in a particular area by utilizing satellite imagery - even though it may be of low resolution. We employ a deep learning based approach to carry out image segmentation on low resolution satellite images of Bangalore, India. Three different model architectures (U-Net, SegNet, FCN) were trained on manually hand-labelled data and tested for the task of semantic segmentation of satellite images. U-Net was found to yield the best results. By using a custom modified U-Net to segment the images, we arrive at the building rooftop area available for solar panels. To calculate the annual solar energy generated in gigawatt-hour (GWh), the incident solar insolation values from the U.S. National Renewable Energy Laboratory (NREL) based on observed weather patterns in Bangalore, and standard values from datasheets of photovoltaic manufacturers are used.
Presently, the prediction of share is a challenging issue for the research community as share market is a chaotic place. The reason behind it, there are several factors such as government policies, international market, weather, performance of company. In this article, a model has been developed using long short term memory (LSTM) to predict the share price of DLF group. Moreover, for the experimental purpose the data of DLF group has been taken from yahoo financial services in the time duration of 2008 to 2018 and the recurrent neural network (RNN) model has been trained using data ranging from 2008 to 2017. This RNN based model has been tested on the data of year 2018. For the performance comparison purpose, other linear regression algorithms i.e. k-nn regression, lasso regression, XGboost etc has been executed and the proposed algorithm outperforms with 2.6% root mean square error.
Executable files such as .exe, .bat, .msi etc. are used to install the software in Windows-based machines. However, downloading these files from untrusted sources may have a chance of having maliciousness. Moreover, these executables are intelligently modified by the anomalous user to bypass antivirus definitions. In this paper, we propose a method to detect malicious executables by analyzing Portable Executable (PE) files extracted from executable files. We trained a supervised binary classifier using features extracted from the PE files of normal and malicious executables. We experimented our method on a large publicly available dataset and reported more than 95% of classification accuracy.
Categorization of software bugs is an important task in software repository mining. Most of the information about the software bugs are in textual form, and it is difficult to categorize these bugs into a particular category as the some of the terms present in the software bugs can be common to multiple categories. Fuzzy similarity technique can be utilized to identify the belongingness of these bugs into different categories. In this paper, a binary software bug categorization technique using fuzzy similarity measure is proposed to classify the bugs as bugs or non-bugs. The fuzzy similarity of a software bug is computed and based on a user-defined threshold value the bug can either be assigned to bug or non-bug category. Experiments are performed on available software bug data sets and performance of proposed fuzzy similarity based classifier is evaluated using the parameters accuracy, F-measure, precision, and recall. The proposed algorithm is also compared with the existing standard machine learning algorithms.
Recent research results show that ontology can be used to improve the accuracy of document clustering. Previous studies mainly focused on the preprocessing part of text document using ontology. In this paper, we propose a hybrid approach, concentrating on both the preprocessing task as well as the clustering algorithm. This is with an objective of reducing the number of features and execution time, eliminate synonymous problems and enhance the accuracy of clustering. Cosine similarity is used as similarity measure. The preprocessing part uses a WordNet Ontology based feature extraction method. In clustering, the initial centroids are found by applying the Red Black Tree based sorting method. The data points are allocated to the suitable clusters using a novel approach, by maintaining the path of similarity between data points and nearest cluster centroids. Experimental results on some of the existing clustering algorithms with cosine similarity are compared with our novel clustering technique. Results show that the proposed hybrid approach executes better on the Newsgroup dataset with considerable improvements in dimensionality reduction, running time, and accuracy.
The agricultural production is affected by the climate changes i.e. humidity, rain, extremes of temperature etc. Additionally, abiotic stresses are causative element to the etiology of disease as well as pest on crops. The production of the crops can be improved by diagnosis as well as detecting the accurate disease on time or in early stage. Moreover, it is very difficult for accurately detecting and treatment based on the technique which used in disease and insect pests diagnosis. Few researchers have made efforts on predicting disease as well as pest crops using machine learning algorithms. Therefore, this paper presents disease identification in corn crops by analyzing the leaves in the very early stage. We have used PlantVillage dataset for experiments and analysis. The validity of the results has been cheeked on various performance metrics such as precision, accuracy, recall, storage space, running time of the model and AUC-RoC. The obtained results shows the proposed technique outperform in comparison with the traditional machine learning algorithms. Developed model is able to achieve the accuracy of 94%.
Nowadays, face detection is common. Its used in many areas. With the help of face detection benchmark data-sets, many signs of progress have been made. Face detection methods used nowadays is not matching the real-world requirements. With new advancing technologies and services, we need to upgrade our existing systems. By using a data-set called as WIDER FACE which is very large in size than already existing data-sets, we can improve the performance. WIDER FACE data-set has many faces in it which may be challenging as it includes faces under different conditions. Moreover, we can see that in face detection task, WIDER FACE data-set is best for training and testing the accuracy of the model. But existing face detection algorithms and models are not up-to the mark. They have major limitations. So we created a WIDER FACE detection system which will help us overcome all those issues.
This paper represents handwritten digit recognition on a very well-known dataset which is MNIST dataset using the Linear Binary Pattern (LBP) and Scale-Invariant Feature Transform (SIFT) feature extraction methods. From the dataset, features have been extracted using this extracting methods. After this, to reduce the number of features or reducing the dimension we have used Principal Component Analysis (PCA) for better performance for our proposed classifier. Then it has been trained by various classifiers. Then the accuracies and errors of those classifiers have been compared and demonstrated. Also some statistical analysis has been done for better understanding of the dataset. From those comparison, it has been shown that our proposed model (SIFT+PCA+CNN) has the better accuracy and less error than other classifiers. The results are competitively compared to previous works and they provide a baseline for evaluation of future work.
We present an experimental study of implementing Latent Dirichlet Allocation (LDA) and Comparative Visual Analytics to trace socio-political issues highlighted within large corpora of political speech transcripts. In this experiment, over 500 speech transcripts are scraped by building scrapers to analyze this big-data of transcripts and derive insights from it. Based on LDA topic modelling algorithm, latent "topics", referred as issues in this paper, were discovered from the speech transcripts and visualized using 'pyLDAvis', which is an interactive visualization tool used upon LDA Model results. Along with LDA, graphical visualizations were generated such as Lexical Dispersion Plots and 'Topic Bar Plots using Matplotlib library of Python. Within comparative analytics, visual graphs were generated for speeches by two different candidates and juxtaposed to compare and interpret their discourse. Linguists have performed Political Discourse Analysis (PDA) using manual approaches but analyzing such a large volume of speeches is practically time consuming and extremely complex. Our experiment which focuses on identifying socio-political issues within speech transcripts using NLP based text analytics proves to be a beneficial technique for understanding Political Discourse Analysis (PDA).
The age of digital information uses images in fields like military and medical applications, but the security of those transacted images is still a question mark. To overcome this challenge, an efficient encryption system has to be developed which should accomplish confidentiality, integrity, security and it should also prevent the access of images by unauthorized users. Such an image encryption system has to be developed which provides enhanced security for images. This system uses two important techniques which is based on chaos theory namely: Confusion and Diffusion. Confusion part uses block scrambling and modified zigzag transformation while the Diffusion part uses 3D logistic map and key generation followed by additive cipher. This system also protects from statistical and differential attacks. The experimental results such as Entropy, Histogram analysis, Mean Absolute Error, Number of Pixels Change Rate (NPCR), Unified Average Changing Intensity (UACI) proves that the security of the images has been preserved at a higher level and also prevents the unauthorized access to the sensitive information.
Millimeter wave communication in combination with MIMO system and NOMA multiple access technology are emerging technologies for 5G and next generation wireless communication. For the mmWave (Millimeter wave) based communication, fundamental challenge of massive MIMO communication is capacity improvement while minimizing cost and energy components. There are various solutions presented so far but the challenge of energy and spectrum efficiency yet to optimize along with capacity improvement and scalability. The novel concept, beamspace MIMO-NOMA which is integration of two technologies, Non orthogonal multiple access and beamspace MIMO. In Beamspace MIMO communication, only few number of RF chains are preferred to serve users, so energy consumption is decreased while with the help of NOMA more than one user can be served in each beam. This system can significantly restrict energy consumption. But the problem of such methods is the scalability. Therefore designing scalable spectrum and energy efficiency solution for mmWave communications is main research problem. In this paper, we proposed modified beamspace MIMO-NOMA in which user clustering is used. We used precoding technique to mitigate interferences. Also, Modified iterative power allocation is used for the intensification of achievable sum rate. Results obtained through simulation indicate that proposed method is more efficient than the beamspace MIMO-NOMA.
Entity Resolution (ER) is a prerequisite to several Web applications including enhancing semantic searches and information extraction from the Web, strengthening the Web of Data by interlinking entity descriptions from autonomous sources, and supporting reasoning using related ontologies. While designing an ER system, it is assumed that each entity profile consists of an exclusively identified set of attribute-value pairs, each entity profile matches to a solitary real-world object, and two similar profiles are identified, while they co-occur in at least one block. ER is an inherently quadratic problem (i.e., O (n2)), given that every entity must draw a comparison with others. Moreover, existing ER techniques relinquishes to scale for large entity collections, Web data. The most well-known solution for addressing large-scale ER in the literature is blocking, which is an approximate solution where similar entities are grouped into blocks and comparisons are limited to within blocks. The process of entity resolution and the types of entity resolution in relational and Web data are discussed in this paper. Further, the paper reviews the literature on the approaches introduced by former researchers on the entity resolution system. The data integration, block building, and block processing phases, and the challenges involved for designing an efficient ER system are discussed. This paper concludes with the measures required to evaluate entity resolution approaches.
The evolution of cluster computers based on multicores, many cores and GPGPU accelerators is encouraging application developers to write hybrid parallel programs. Hybrid parallel programming is quite complex as it requires use of multiple programming paradigms such as MPI, OpenMP, CUDA/OpenCL to exploit the varied computational power available in a system. The paper brings out the challenges faced by application developers desiring to use heterogeneous HPC clusters. It describes a unified development environment which eases the complete development lifecycle of hybrid parallel programs on a HPC cluster. The software is capable of providing access to multiple clusters of different architectures, owing to the modularity of design and web based approach. The paper also serves as a good resource for researchers interested to gain an insight into hybrid parallel programming.
Recent Development in Hardware and Software Technology for the communication email is preferred. But due to the unbidden emails, it affects communication. There is a need for detection and classification of spam email. In this present research email spam detection and classification, models are built. We have used different Machine learning classifiers like Naive Bayes, SVM, KNN, Bagging and Boosting (Adaboost), and Ensemble Classifiers with a voting mechanism. Evaluation and testing of classifiers is performed on email spam dataset from UCI Machine learning repository and Kaggle website. Different accuracy measures like Accuracy Score, F measure, Recall, Precision, Support and ROC are used. The preliminary result shows that Ensemble Classifier with a voting mechanism is the best to be used. It gives the minimum false positive rate and high accuracy.
Software defined Network is a network defined by software, which is one of the important feature which makes the legacy old networks to be flexible for dynamic configuration and so can cater to today's dynamic application requirement. It is a programmable network but it is prone to different type of attacks due to its centralized architecture. The author provided a solution to detect and prevent Distributed Denial of service attack in the paper. Mininet [5] which is a popular emulator for Software defined Network is used. We followed the approach in which collection of the traffic statistics from the various switches is done. After collection we calculated the packet rate and bandwidth which shoots up to high values when attack take place. The abrupt increase detects the attack which is then prevented by changing the forwarding logic of the host nodes to drop the packets instead of forwarding. After this, no more packets will be forwarded and then we also delete the forwarding rule in the flow table. Hence, we are finding out the change in packet rate and bandwidth to detect the attack and to prevent the attack we modify the forwarding logic of the switch flow table to drop the packets coming from malicious host instead of forwarding it.
In Healthcare 4.0, Remote patient monitoring (RPM) becomes a more powerful and flexible patient observation through wearable sensors at any time and anywhere. The most focused application area of RPM which allows doctors to get real-time information of their patient remotely with the help of wireless communication system. Thus, RPM reduces the time and cost of the patient. It also provides the quality care to the patient. To enhance the security and privacy of the patient data, in this paper, we have presented a Permissioned blockchain-based healthcare architecture. We have also discussed the challenges and their solutions. We have described the applications of blockchain. We also have given the usage of Machine learning with blockchain technology which can impact the healthcare industry.