In this work, we propose an ensemble of classification trees (CT) and artificial neural networks (ANN). Several statistical properties including universal consistency and upper bound of an important parameter of the proposed classifier are shown. Numerical evidence is also provided using various real-life data sets to assess the performance of the model. Our proposed nonparametric ensemble classifier does not suffer from the “curse of dimensionality” and can be used in a wide variety of feature selection cum classification problems. Performance of the proposed model is quite better when compared to many other state-of-the-art models used for similar situations.
Modularity is a widely used goodness metric that effectively measures the strength of the community structures present in a network. However its performance may not be desirable for identifying densely connected communities or clusters of a network. It also often fails to identify communities or clusters that contain very few nodes. Furthermore, modularity is defined based only on the exact node-to-node connectivity of a network while disregarding their neighborhood connectivity. In this paper, we associate the neighborhood connectivity to the modularity function and propose a generalized modularity function based on the node similarity measure which quantifies the quality of a given network partition. Making use of this similarity based modularity function, an effective agglomerative approach for identifying communities is introduced. This agglomerative approach iteratively discovers the final community structure of the network by finding and merging together, at each step, the community pairs which maximize the proposed modularity value. A significant characteristic of the proposed method is that it does not need any prior knowledge about the actual communities of a network. The performance of the proposed method and state-of-the-art algorithms are compared using the value of modularity, normalized mutual information and adjusted variation of information measures on several real world and artificial networks. The empirical results show the effectiveness of the proposed method compared to the state-of-the-art techniques.
The existing biclustering algorithms for finding feature relation based biclusters often depend on assumptions like monotonicity or linearity. Though a few algorithms overcome this problem by using density-based methods, they tend to miss out many biclusters because they use global criteria for identifying dense regions. The proposed method, RelDenClu uses the local variations in marginal and joint densities for each pair of features to find the subset of observations, which forms the bases of the relation between them. It then finds the set of features connected by a common set of observations, resulting in a bicluster. To show the effectiveness of the proposed methodology, experimentation has been carried out on fifteen types of simulated datasets. Further, it has been applied to six real-life datasets. For three of these real-life datasets, the proposed method is used for unsupervised learning, while for other three real-life datasets it is used as an aid to supervised learning. For all the datasets the performance of the proposed method is compared with that of seven different state-of-the-art algorithms and the proposed algorithm is seen to produce better results. The efficacy of proposed algorithm is also seen by its use on COVID-19 dataset for identifying some features (genetic, demographics and others) that are likely to affect the spread of COVID-19.
Extracting an effective feature set on the basis of dataset characteristics can be a useful proposition to address a classification task. For multi-label datasets, the positive and negative class memberships for the instance set vary from label to label. In such a scenario, a dedicated feature set for each label can serve better than a single feature set for all labels. In this article, we approach multi-label learning addressing the same concern and present our work Multi-label learning through Minimum Spanning Tree based subset selection and feature extraction (MuMST-FE). For each label, we estimate the positive and negative class shapes using respective Minimum Spanning Trees (MSTs), followed by subset selection based on the key lattices of the MSTs. We select a unique subset of instances for each label which participates in the feature extraction step. A distance based feature set is extracted for each label from the reduced instance set to facilitate final classification. The classifiers modelled from MuMST-FE is found to possess improved robustness and discerning capabilities which is established by the performance of the proposed schema against several state-of-the-art approaches on ten benchmark datasets.
In this article we propose a general methodology for constructing complex networks. Popular selection schemes in Genetic algorithms are used for this construction. Mathematically, it has been shown that, under some weak constraints, the degree distribution of the resulting networks follow power-law, as seen in real world networks. Power-law degree distribution is one of the most significant structural characteristics observed in many real-world complex networks. The main reason behind the emergence of this phenomenon is the mechanism of preferential attachment which states that in a growing network a node with higher degree is more likely to receive new links. However, degree is not the only key factor influencing the network growth leading to power-law degree distribution. Instead, there must be several other factors whose cumulative effect, called fitness of a node, has a significant role in attracting other nodes and thereby producing power law networks. The concept of fitness can be thought of as a generalization of node degree. Heterogeneity in preferential linking also plays an important role in producing power-law networks in this context. The proposed construction methodology, also leading to power law networks, combines the inherent fitness value of a node, drawn from a particular distribution, with various attachment schemes based on the different selection methods commonly used in Genetic algorithms. Six different selection schemes are used in total. Different well known structural measures like average degree of the nearest neighbors, average path length, clustering coefficient, etc. are calculated for each newly generated network to understand their behavior patterns. It has been found that these six schemes can be divided into two distinct groups of three on the basis of their structural properties, where one of these two groups produces proper power-law networks which possess topological properties similar to observed in the real world. Finally, extensive simulations and experiments over scientific collaboration networks validate the effectiveness of the proposed models.
This chapter investigates theoretical results regarding the behavior of a genetic algorithm-based pattern classification methodology, for an infinitely large number of sample points n, in an N dimensional space RN. It shows that for n → ∞ and for a sufficiently large number of iterations, the performance of this classifier approaches that of the Bayes classifier. Experimental results, for a triangular distribution of points, are also included that conform to this claim. The chapter attempts to show theoretically that the decision surfaces generated by the aforesaid GA-based classifier will approach the Bayes decision surfaces for a large number of training sample points (n) and will consequently provide the optimal decision boundary in terms of the number of misclassified samples. It gives a brief outline of the GA-based classifier. The chapter provides the theoretical treatment to find a relationship between this classifier and the Bayes classifier in terms of classification accuracy.
The use of kernel density estimation is quite well known in large variety of machine learning applications like classification, clustering, feature selection, etc. One of the major issues in the construction of kernel density estimators is the tuning of bandwidth parameter. Most of the bandwidth selection procedures optimize mean integrated squared or absolute error, which require huge computational time as the size of the data increases. Here, the bandwidth has been taken to be a function of inter-point distances of the data set. It is defined as a function of the length of Euclidean Minimal Spanning Tree of the given sample points. No rigorous theory about the asymptotic properties of the EMST based density estimator has been developed in the literature. Theoretical analysis of the asymptotic properties of the EMST based density estimator has been established and proved that the estimator is asymptotically unbiased to the original density at its every continuity point. Moreover, theoretical analysis has been provided for general kernel. Experiments are conducted using both synthetic and real-life data sets to compare the performance of the EMST bandwidth to those of conventional cross-validation and plug-in bandwidth selectors. It is found that the EMST based estimator achieves the comparative performance, while being simpler and faster than the conventional estimators.
Dimensionality reduction is an essential pre-processing technique in many of the data analysis tasks. Popular approaches for dimensionality reduction are Feature Selection (FS) and Feature Extraction (FE). Till now, these approaches are often studied separately or independently so that the final result contains either original or transformed features. In our work, we propose to bridge these two approaches with the aim of finding reduced feature set to contain both kinds (original as well as transformed) of features. A new framework, called Minimum Projection error Minimum Redundancy (MPeMR), is introduced to obtain this result while maintaining orthogonality property among selected original and linear combinations of features. A unified iterative algorithm, for both supervised and unsupervised cases, is also developed under this framework. For each case, the performance of the proposed algorithm is successfully compared with the state-of-the-art methods on real-life data sets.
Measures for uncertainty due to approximation of sets in rough set theory are accuracy and roughness. In determining these quantities, the cardinality of a set is always used and never the numerical values of the attributes (if they exist) of elements in the sets. Therefore, distances between the exact set and the corresponding upper and lower approximations can give a better quantitative measure of the roughness. Here, we propose a measure based on Hausdorff metric which takes into account the distance between two sets, the exact set and its two approximations (lower and upper). Using this measure, we can quantify the uncertainty of a rough set based on the values in the domain of sample points but not on the basis of number of sample points. Also, we propose a new measure for granulation which is again based on the Hausdorff metric. The effectiveness of the proposed measures is demonstrated on a synthetic data.
The similarity based decision rule computes the similarity between a new test document and the existing documents of the training set that belong to various categories. The new document is grouped to a particular category in which it has maximum number of similar documents. A document similarity ba sed supervised decision rule for text categorization is proposed in this article. The similarity measure determine the similarity between two documents by finding their distances with all the documents of training set and it can explicitly identify two dissimilar documents. The decision rule assigns a test document to the best one among the competing categories, if the best category beats the next competing category by a previously fixed margin. Thus the proposed rule enhances the certainty of the decision. The salient feature of the decision rule is that, it never assigns a document arbitrarily to a category when the decision is not so certain. The performance of the proposed decision rule for text categorization is compared with some well known classification techniques e.g., k-nearest neighbor decision rule, support vector machine, naive bayes etc. using various TREC and Reuter corpora. The empirical results have shown that the proposed method performs significantly better than the other classifiers for text categorization.
Different facial feature extraction schemes are available in face recognition literature where the face is cropped and feature points are extracted using mathematical formulae along with probabilistic distance measure between major feature points. In most of the cases the face cropping is done manually and the formulae for extracting a feature point are highly complicated. In this article a facial feature extraction method is proposed where color face images are auto-cropped and control points are extracted, both using the same segmentation mechanism. The segmentation method is first used to auto crop the image and then again applied on this auto cropped image for the detection of major connected components. The feature points are detected simply using the geometrical measurement of location and size of the component without any a priory knowledge of the probabilistic distance between the feature points or using any feature point extraction formula. A T-shaped face image is formed comprising of major feature points. Recognition rate on the unprocessed face images, using PCA, is recorded. PCA is applied on the T-shaped faces and the improvement in rate of recognition is concluded to be statistically significant.
Document clustering refers to the task of grouping similar documents and segregating dissimilar documents. It is very useful to find meaningful categories from a large corpus. In practice, the task to categorize a corpus is not so easy, since it generally contains huge documents and the document vectors are high dimensional. This paper introduces a hybrid document clustering technique by combining a new hierarchical and the traditional k-means clustering techniques. A distance function is proposed to find the distance between the hierarchical clusters. Initially the algorithm constructs some clusters by the hierarchical clustering technique using the new distance function. Then k-means algorithm is performed by using the centroids of the hierarchical clusters to group the documents that are not included in the hierarchical clusters. The major advantage of the proposed distance function is that it is able to find the nature of the corpora by varying a similarity threshold. Thus the proposed clustering technique does not require the number of clusters prior to executing the algorithm. In this way the initial random selection of k centroids for k-means algorithm is not needed for the proposed method. The experimental evaluation using Reuter, Ohsumed and various TREC data sets shows that the proposed method performs significantly better than several other document clustering techniques. F-measure and normalized mutual information are used to show that the proposed method is effectively grouping the text data sets.
International conference on advances in pattern recognition and digital techniques : proceedings of P.C.Mahalanobis birth centenary volume(1993).
The paper addresses the problem of finding top k influential nodes in large scale directed social networks. We propose two new centrality measures, Diffusion Degree for independent cascade model of information diffusion and Maximum Influence Degree. Unlike other existing centrality measures, diffusion degree considers neighbors' contributions in addition to the degree of a node. The measure also works flawlessly with non uniform propagation probability distributions. On the other hand, Maximum Influence Degree provides the maximum theoretically possible influence Upper Bound for a node. Extensive experiments are performed with five different real life large scale directed social networks. With independent cascade model, we perform experiments for both uniform and non uniform propagation probabilities. We use Diffusion Degree Heuristic DiDH and Maximum Influence Degree Heuristic MIDH, to find the top k influential individuals. k seeds obtained through these for both the setups show superior influence compared to the seeds obtained by high degree heuristics, degree discount heuristics, different variants of set covering greedy algorithms and Prefix excluding Maximum Influence Arborescence PMIA algorithm. The superiority of the proposed method is also found to be statistically significant as per T-test.
Degree distribution of nodes, especially a power-law degree distribution, has been regarded as one of the most significant structural characteristics of social and information networks. However it is observed here that for many large scale real world networks, the power-law does not fit properly because of the presence of large fluctuations and sparsity in upper and lower tails of the distribution. Here we have proposed to fit the truncated geometric distribution on three distinct and non-overlapping parts of the degree frequency table. Extensive experiments on twenty three (23) real world networks revealed that the proposed model fitted better than the power-law and other distributions.
With the increasing availability of experimental data on gene interactions, modeling of gene regulatory pathways has gained special attention. Gradient descent algorithms have been widely used for regression and classification applications. Unfortunately, results obtained after training a model by gradient descent are often highly variable. In this paper, we present a new second order learning rule based on the Newton's method for inferring optimal gene regulatory pathways. Unlike the gradient descent method, the proposed optimization rule is independent of the learning parameter. The flow vectors are estimated based on biomass conservation. A set of constraints is formulated incorporating weighting coefficients. The method calculates the maximal expression of the target gene starting from a given initial gene through these weighting coefficients. Our algorithm has been benchmarked and validated on certain types of functions and on some gene regulatory networks, gathered from literature. The proposed method has been found to perform better than the gradient descent learning. Extensive performance comparison with the extreme pathway analysis method has underlined the effectiveness of our proposed methodology.
Thunderstorm forecasting is a challenging job.Machine learning techniques are being applied nowadays in meteorological fields for prediction purpose.This study presents the application of different machine learning tools based on multiple correlation, Multi-layer Perceptron (MLP), K-nearest neighbor (K-nn) method, and modified K-nn method to predict seasonal severe thunderstorms associated with squall occurring in Kolkata, North-East India.The models are trained and tested with the radiosonde data recorded in the early morning at 00:00UTC.The predictors are moisture difference and dry adiabatic lapse rate at different geopotential heights of the atmosphere.Our aim in this paper is to find how much correctly one can nowcast 10 to 14 hours before the 'occurrence'/ 'no occurrence' of evening squall-storms by using a few upper air diagnostic predictors.Modified K-nn method is found to yield very promising prognostic information with high prediction accuracy.The results indicate that forecasting can be done correctly up to 82.02% both for 'squall-storm/no storm' events, and up to 91.11% for 'squall-storm' events using modified K-nn based approach.In this article, modified K-nn method is proved as the best method in comparison with the other methods for the squall-storm prediction.
Objective of the document clustering techniques is to assemble similar documents and segregate dissimilar documents. Unlike document classification, no labeled documents are provided in document clustering. One of the main challenges of any document clustering algorithm is the selection of a good similarity measure. Traditionally, using the vector space model, the number of words common between two documents is used for determining their similarity. This paper introduces a document similarity measure, extensive similarity between the documents. In this approach two documents are considered to be similar if they share a minimum number of common words and they have almost same distance with every other document in the corpus i. e., both are either similar or dissimilar to the other documents. A hierarchical document clustering algorithm, using extensive similarity between the documents is proposed in this article. It is experimentally found on several text data sets that the proposed document clustering algorithm performs significantly better than the traditional document clustering techniques, comparisons for which are based on f-measure and normalized mutual information.
In an extension of previous work, here we introduce a second-order optimization method for determining optimal paths from the substrate to a target product of a metabolic network, through which the amount of the target is maximum. An objective function for the said purpose, along with certain linear constraints, is considered and minimized. The basis vectors spanning the null space of the stoichiometric matrix, depicting the metabolic network, are computed, and their convex combinations satisfying the constraints are considered as flux vectors. A set of other constraints, incorporating weighting coefficients corresponding to the enzymes in the pathway, are considered. These weighting coefficients appear in the objective function to be minimized. During minimization, the values of these weighting coefficients are estimated and learned. These values, on minimization, represent an optimal pathway, depicting optimal enzyme concentrations, leading to overproduction of the target. The results on various networks demonstrate the usefulness of the methodology in the domain of metabolic engineering. A comparison with the standard gradient descent and the extreme pathway analysis technique is also performed. Unlike the gradient descent method, the present method, being independent of the learning parameter, exhibits improved results.
Pabitra Mitra合作论文数Department of Computer Science & Engineering, Indian Institute of Technology3
B. Uma Shankar合作论文数Indian Statistical Institute;Machine Intelligence Unit3