Minimum spanning tree (MST) have been employed in practice for various exploratory data analyses, e.g., to discover clusters of arbitrary shapes and sizes from diversified datasets. However, the computational complexity of these algorithms becomes a bottleneck when they are applied on very large datasets. The main overhead associated with these algorithms is the proximity search in the construction of a similarity graph which incurs O(N2) time on a set of N data points. To conquer this issue, several graph sparsification techniques have been proposed which take O(N3/2) time. This paper proposes a O(N4/3) time local neighborhood similarity graph construction technique using two levels of partitioning and merging, which is asymptotically O(N1/6) factor improvement over the existing methods. To the best of our knowledge, this is the asymptotically fastest known algorithm using two levels of partitioning and merging. Experimental analysis shows that the proposed sparse graph construction technique reduces 99.46% edges of the complete graph by preserving the relevant neighborhood information. Also, MST constructed from the proposed graph captures the local neighborhood of data points efficiently which is shown in terms of edge error and weight error rates. Finally, we demonstrate that the proposed approximate MST based clustering technique outperforms the best-known existing algorithms in terms of clustering accuracy and execution time on diversified datasets.
Abstract A large amount of music data is now available on the Internet thanks to the advent of the World Wide Web. In addition to seeking for expected music for clients, it becomes necessary to build a recommendation service. The majority of existing music recommendation systems relies on collaborative or content-based engines. However, a user's music selection is not only based on prior preferences or musical content. However, it is also reliant on the user's mood. This paper presents an mood-based music (MoodSIC) recommendation framework that automatically learns a user's mood and suggest a list of songs pertaining to the mood. MoodSIC first detects the listener’s mood using various parameters such as skin temperature, facial texture, voice input, facial expression and then render a recommendation of the songs. The MoodSIC system is delivered as a web-application which uses MongoDB as a back-end for storing the songs. The proposed recommendation system provides a user-friendly interface to detect user mood using webcam, generate recommended playlist and autoplay music to the liking from generated playlist.
This study presents a deep learning model to serve as an image caption generator that generates descriptions or captions of the images in proper natural language sentences, which will then be read aloud by the text to speech translator. With the growing demand for tools like this in various fields such as assisting the visually impaired, self-driving vehicles, and virtual assistants. Hence, the development of such systems has become increasingly important. The proposed system utilizes a combination of Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) with attention models, specifically by using the Inception V3 model and a variant of RNN called Gated Recurrent Units (GRU).
Lithium-ion batteries, among many energy storage systems, offer high energy density, low voltage dips, long lifespan, and wide working temperatures. They have been widely adopted in a variety of applications, including as electric vehicles, aerospace, energy management systems, etc. Accurate prediction of remaining useful lifetime (RUL) and health status of lithium-ion batteries have received lot of attention in the recent years. Machine learning approaches have recently gained popularity as a means of empirically learning and predicting battery behaviour. However, the complex and nonlinear behaviour of lithium-ion batteries pose challenges for traditional machine learning approaches. This paper investigates the application of two non-linear machine learning models, namely artificial neural network (ANN) and 1-D Convolution Neural Network (1-D CNN), for predicting the RUL. NASA prognostics battery dataset is utilized for the present study. Experimental results indicate that the 1-D CNN achieves better prediction accuracy as compared to ANN and other traditional machine learning.
Falls in elderly people are common. Detecting near-fall situations can prevent fall-related injuries. Wearable technology has made a significant impact in this direction and fall-detection from wearable sensors has become an important research problem in ambient assisted living. Although a number of machine learning algorithms exist for wearable fall-detection, most of them are based on supervised learning. These algorithms require a huge amount of training data and generating such data is very time-consuming process. This paper employs deep embedded clustering, an unsupervised learning approach, for wearable fall-detection. For experimental purpose, Kaggle fall-detection dataset is considered. Results indicate that deep embedded clustering achieves higher accuracy in attaining fall-detection.
Social media such as Twitter is these days one of the fundamental news hotspots for a great many individuals all over the planet because of their minimal expense, simple access, and quick spread. In this paper a real-time disaster mining from twitter disaster event mining using apache spark for natural disasters is proposed by fusing text, images and geo-spatial data. The way of transmitting messages or news reports is not good for developing societies where even the emergency information even can't be transmitted to locals who may get affected due to the disaster. Locating disasters from different locations of the world and mapping location collected from the Twitter sources from different areas at risk of natural disaster and gather them to Geo-parsed real-time tweet data streams. In addition, the Geo-statistical analysis to generate real-time disaster mapping is plotted in a base map. Then report is generated for two case studies tweet disaster mapping to find post-event impact assessment is carried out and detection accuracy is computed. The dataset applied Streaming K-means algorithm and CNN is validated to be a very effective technique to extract event clusters of the identical spatial entities in diversified contexts such as text, photos, geo-tags are observed with an accuracy of 91% which is reasonably acceptable.
Telecom companies record customer's actions, which generates a large amount of data that can lead to crucial insights on customer behaviour and demands. Most telecom companies use customer segmentation to increase customer satisfaction, which entails dividing targeted customers into different groups based on demographics or usage perspectives such as gender, age group, buying behaviour, usage pattern, special interests, and other characteristics that represent the customer. With more number of attributes and great sparsity of telecom data, identifying targeted customers become difficult and various machine learning algorithms have been proposed for the same. Deep learning has gained huge popularity in various business analytics and operations. However, use of deep learning for customer segmentation is very limited. This paper aims to segment Telecom customer data using deep embedded clustering algorithm. For experimental purpose, Kaggle's telco customer churn dataset is considered. Results of our study indicate that deep embedded clustering algorithm is able to attain better segmentation results as compared to traditional clustering algorithms such as K-means and Hierarchical clustering approaches.
In the era of fourth industrial revolution or Industry 4.0, Computational Intelligence, Data Science and smart Information and Communication technology (ICT) are gaining huge importance in our society.Governments around the world started relying heavily on information and communication technologies to build smart and hyper-connected infrastructures that allow cities to provide better services to people and reduce energy consumption, as the global urban population is growing substantially.These intelligent technologies allow cities to introduce new ways of monitoring the atmosphere, buildings, street lighting, traffic, crowds, crime, etc.In building a smarter and more intelligent city around the world, Computational intelligence and ICT are the underlying technologies; but if not wisely applied, ICT can be a major environmental problem.At present, about 2 percent of global greenhouse gas (GHG) emissions are accounted for by the global ICT industry.This footprint is expected to increase dramatically to about 14 percent by 2040, according to a recent report.Intelligent use of advances in ICT will assist in minimizing GHG while still achieving its goals.ICT is theoretically able to reduce the carbon footprint in other fields by a factor of 10, based on the Global e-Sustainability Initiative (GeSI) report.ICT technologies such as Green ICT, the Internet of Things (IoT) and Artificial Intelligence (AI) can play important roles, not just in making our environment smarter, but also greener and more sustainable.
In the era of fourth industrial revolution or Industry 4.0, Computational Intelligence, Data Science and smart Information and Communication technology (ICT) are gaining huge importance in our society.Governments around the world started relying heavily on information and communication technologies to build smart and hyper-connected infrastructures that allow cities to provide better services to people and reduce energy consumption, as the global urban population is growing substantially.These intelligent technologies allow cities to introduce new ways of monitoring the atmosphere, buildings, street lighting, traffic, crowds, crime, etc.In building a smarter and more intelligent city around the world, Computational intelligence and ICT are the underlying technologies; but if not wisely applied, ICT can be a major environmental problem.At present, about 2 percent of global greenhouse gas (GHG) emissions are accounted for by the global ICT industry.This footprint is expected to increase dramatically to about 14 percent by 2040, according to a recent report.Intelligent use of advances in ICT will assist in minimizing GHG while still achieving its goals.ICT is theoretically able to reduce the carbon footprint in other fields by a factor of 10, based on the Global e-Sustainability Initiative (GeSI) report.ICT technologies such as Green ICT, the Internet of Things (IoT) and Artificial Intelligence (AI) can play important roles, not just in making our environment smarter, but also greener and more sustainable.
In the era of fourth industrial revolution or Industry 4.0, Computational Intelligence, Data Science and smart Information and Communication technology (ICT) are gaining huge importance in our society.Governments around the world started relying heavily on information and communication technologies to build smart and hyper-connected infrastructures that allow cities to provide better services to people and reduce energy consumption, as the global urban population is growing substantially.These intelligent technologies allow cities to introduce new ways of monitoring the atmosphere, buildings, street lighting, traffic, crowds, crime, etc.In building a smarter and more intelligent city around the world, Computational intelligence and ICT are the underlying technologies; but if not wisely applied, ICT can be a major environmental problem.At present, about 2 percent of global greenhouse gas (GHG) emissions are accounted for by the global ICT industry.This footprint is expected to increase dramatically to about 14 percent by 2040, according to a recent report.Intelligent use of advances in ICT will assist in minimizing GHG while still achieving its goals.ICT is theoretically able to reduce the carbon footprint in other fields by a factor of 10, based on the Global e-Sustainability Initiative (GeSI) report.ICT technologies such as Green ICT, the Internet of Things (IoT) and Artificial Intelligence (AI) can play important roles, not just in making our environment smarter, but also greener and more sustainable.
Gene co-expression analysis is an important research problem in molecular biology that helps to identify co-occurring genes in potential biological function. Clustering methods have been widely employed for this problem and hierarchical clustering based gene expression analysis has made tremendous progress in the past years. However, these methods heavily rely on proximity measures used in the clustering process. One of the major issues of hierarchical clustering is their inability to detect arbitrary shaped clusters in high dimensional spaces. Another issue is their pre-requisite of distance matrix calculation, which is not computationally efficient for large datasets. To address these issues, this paper proposes approximate similarity measures based on local neighborhood representation using minimum spanning tree. The effectiveness of proposed similarity measures is tested using hierarchical clustering algorithm. Experimental results on microarray gene expression datasets reveal that the proposed similarity measures achieve improved results in terms of clustering accuracy as well as reduced time complexity as compared to conventional distance measures.
In the era of fourth industrial revolution or Industry 4.0, Computational Intelligence, Data Science and smart Information and Communication technology (ICT) are gaining huge importance in our society.Governments around the world started relying heavily on information and communication technologies to build smart and hyper-connected infrastructures that allow cities to provide better services to people and reduce energy consumption, as the global urban population is growing substantially.These intelligent technologies allow cities to introduce new ways of monitoring the atmosphere, buildings, street lighting, traffic, crowds, crime, etc.In building a smarter and more intelligent city around the world, Computational intelligence and ICT are the underlying technologies; but if not wisely applied, ICT can be a major environmental problem.At present, about 2 percent of global greenhouse gas (GHG) emissions are accounted for by the global ICT industry.This footprint is expected to increase dramatically to about 14 percent by 2040, according to a recent report.Intelligent use of advances in ICT will assist in minimizing GHG while still achieving its goals.ICT is theoretically able to reduce the carbon footprint in other fields by a factor of 10, based on the Global e-Sustainability Initiative (GeSI) report.ICT technologies such as Green ICT, the Internet of Things (IoT) and Artificial Intelligence (AI) can play important roles, not just in making our environment smarter, but also greener and more sustainable.
Clustering has been widely applied in interpreting the underlying patterns in microarray gene expression profiles, and many clustering algorithms have been devised for the same. K-means is one of the popular algorithms for gene data clustering due to its simplicity and computational efficiency. But, K-means algorithm is highly sensitive to the choice of initial cluster centers. Thus, the algorithm easily gets trapped with local optimum if the initial centers are chosen randomly. This paper proposes a deterministic initialization algorithm for K-means (DK-means) by exploring a set of probable centers through a constrained bi-partitioning approach. The proposed algorithm is compared with classical K-means with random initialization and improved K-means variants such as K-means++ and MinMax algorithms. It is also compared with three deterministic initialization methods. Experimental analysis on gene expression datasets demonstrates that DK-means achieves improved results in terms of faster and stable convergence, and better cluster quality as compared to other algorithms.
Human activity recognition (HAR) from the time-series data generated by smarlt devices like smartphones and smartwathces is absolutely necessary for future intelligent health-care systems. A number of machine learning approaches have been devised for human activity recognition from such devices. However, most of the existing approaches have considered activity recognition as a supervised learning problem, which requires a training dataset for learning the activities. With ever-growing computing power of smart devices, the amount of data generated by these devises is also increasing in manifold. Thus, it becomes a tedious task to annotate the data collected from the smart devices. Being an unsupervised learning approach, cluster analysis provides meaningful insights on hidden patterns in the huge volume of data without training examples. Hence cluster analysis can be used as a preprocessing step in human activity recognition when there is no sufficient information about the number of activities. This paper presents a comparative study of three clustering algorithms such as K-means, hierarchical agglomerative clustering and Fuzzy C-means for human activity recognition problem. A method to automatically determine the number of activities is also demonstrated in this paper. Experimental results on two UCI activity recognition datasets show that FCM algorithm effectively categorizes the activities.
Software fault prediction is an important task in software development process which enables software practitioners to easily detect and rectify the errors in modules or classes. Various fault prediction techniques have been studied in the past and unsupervised learning methods such as clustering techniques are drawing much attention in the recent years. K-means is a well known clustering algorithm which is applied on various exploratory analysis including software fault prediction. This paper provides a comparative study on software fault prediction using K-means clustering algorithm and its variants. We use five software fault prediction datasets taken from PROMISE repository to evaluate the prediction accuracy of the clustering algorithms. Experimental results indicate that proper initial seed selection enables K-means algorithm to effectively group the faulty modules.
Clustering has become one of the important data analysis techniques for the discovery of cancer disease. Numerous clustering approaches have been proposed in the recent years. However, handling of high-dimensional cancer gene expression datasets remains an open challenge for clustering algorithms. In this paper, we present an improved graph based clustering algorithm by applying edge betweenness criterion on spanning subgraph. We carry out empirical analysis on artificial datasets and five cancer gene expression datasets. Results of the study show that the proposed algorithm can effectively discover the cancerous tissues and it performs better than two recent graph based clustering algorithms in terms of cluster quality as well as modularity index.
Gene expression data clustering is an important biological process in DNA microarray analysis. Although there have been many clustering algorithms for gene expression analysis, finding a suitable and effective clustering algorithm is always a challenging problem due to the heterogeneous nature of gene profiles. Minimum Spanning Tree (MST) based clustering algorithms have been successfully employed to detect clusters of varying shapes and sizes. This paper proposes a novel clustering algorithm using Eigenanalysis on Minimum Spanning Tree based neighborhood graph (E-MST). As MST of a set of points reflects the similarity of the points with their neighborhood, the proposed algorithm employs a similarity graph obtained from k(') rounds of MST (k(')-MST neighborhood graph). By studying the spectral properties of the similarity matrix obtained from k(')-MST graph, the proposed algorithm achieves improved clustering results. We demonstrate the efficacy of the proposed algorithm on 12 gene expression datasets. Experimental results show that the proposed algorithm performs better than the standard clustering algorithms.
Spectral clustering is one of the most popular modern graph clustering techniques in machine learning. By using the eigenvalue analysis, spectral methods partition the given set of points into number of disjoint groups. Spectral methods are very useful in determining non-convex shaped clusters, identifying such clusters is not trivial for many traditional clustering methods including hierarchical and partitional methods. Spectral clustering may be carried out either as recursive bi-partitioning using fiedler vector second eigenvector or as muti-way partitioning using first k eigenvectors, where k is the number of clusters. Although spectral methods are widely discussed, there has been a little attention on which post-clustering algorithm for eg. K-means should be used in multi-way spectral partitioning. This motivated us to carry out an experimental study on the influence of post-clustering phase in spectral methods. We consider three clustering algorithms namely K-means, average linkage and FCM. Our study shows that the results of multi-way spectral partitioning strongly depends on the post-clustering algorithm.
K-means clustering algorithm is rich in literature and its success stems from simplicity and computational efficiency. The key limitation of K-means is that its convergence depends on the initial partition. Improper selection of initial centroids may lead to poor results. This paper proposes a method known as Deterministic Initialization using Constrained Recursive Bi-partitioning (DICRB) for the careful selection of initial centers. First, a set of probable centers are identified using recursive binary partitioning. Then, the initial centers for K-means algorithm are determined by applying a graph clustering on the probable centers. Experimental results demonstrate the efficacy and deterministic nature of the proposed method.
Minimum spanning tree (MST) based clustering algorithms have been employed successfully to detect clusters of heterogeneous nature. Given a dataset of n random points, most of the MST-based clustering algorithms first generate a complete graph G of the dataset and then construct MST from G. The first step of the algorithm is the major bottleneck which takes O(n 2) time. This paper proposes two algorithms namely MST-based clustering on K-means Graph and MST-based clustering on Bi-means Graph for reducing the computational overhead. The proposed algorithms make use of a centroid based nearest neighbor rule to generate a partition-based Local Neighborhood Graph (LNG). We prove that both the size and the computational time to construct the graph (LNG) is O(n 3/2), which is a O(√(n)) factor improvement over the traditional algorithms. The approximate MST is constructed from LNG in O(n^3/2 n) time, which is asymptotically faster than O(n 2). The advantage of the proposed algorithms is that they do not require any parameter setting which is a major issue in many of the nearest neighbor finding algorithms. Experimental results demonstrate that the computational time has been reduced significantly by maintaining the quality of the clusters obtained from the MST.