The emotion varies and propagates with the spatial and temporal information of individuals through social media, which uncovers several interaction mechanisms and features the community structure in order to facilitate individuals’ communication and emotional contagion in social networks. Aiming to show the detailed process and characteristics of emotional contagion within social media, we propose an emotional independent cascade model in which individual emotion can affect the subsequent emotion of his/her friends. The transmissibility is introduced to measure the capability of propagating emotion with respect to an individual in social networks. By analyzing the patterns of emotional contagion on Twitter data, we find that the value of transmissibility differs on different layers and on different community structures. Extensive experiments were conducted and the results reveal that, the polar emotion of hub users can lead to the disappearance of opposite emotion, and the transmissibility makes no sense. The final emotional distribution depends on the initial emotional distribution and the transmissibilities. Individuals from a small community are more likely to change their mood by the influence of community leaders. In addition, we compared the proposed model with two other models, the emotion-based spreader–ignorant–stifler model and the standard independent cascade model. The results demonstrate that the proposed model can reflect the real-world situation of emotional contagion for heterogeneous social media while the computational complexities of all these three models are similar.
The publishing freedom of users on Internet poses new challenges in Web content filtering.This paper presents a self-study algorithm,called SAFE(self-study algorithm of filtering Chinese text content),for Chinese content filtering through two layers.It processes texts in the form of data stream.Based on Apriori property,SAFE filters Chinese text content through two layers by mining key characters and keywords without manual dictionary.The performance research of SAFE on the real-world data shows that for the given theme,the recall of SAFE is greater than 93.75% and the precision is 100%.The runtime of SAFE satisfies the real-time requirement of Web applications.
To effectively score pages with uncertainty in web social networks, we first proposed a new concept called transition probability matrix and formally defined the uncertainty in web social networks. Second, we proposed a hybrid page scoring algorithm, called WebScore, based on the PageRank algorithm and three centrality measures including degree, betweenness, and closeness. Particularly, WebScore takes into a full consideration of the uncertainty of web social networks by computing the transition probability from one page to another. The basic idea of WebScore is to: (1) integrate uncertainty into PageRank in order to accurately rank pages, and (2) apply the centrality measures to calculate the importance of pages in web social networks. In order to verify the performance of WebScore, we developed a web social network analysis system which can partition web pages into distinct groups and score them in an effective fashion. Finally, we conducted extensive experiments on real data and the results show that WebScore is effective at scoring uncertain pages with less time deficiency than PageRank and centrality measures based page scoring algorithms.
In order to score Web pages in an effective manner,a new page scoring algorithm,CentralRank,was proposed based on centrality measures,including degree,betweenness and closeness,and the PageRank algorithm.The CentralRank algorithm computes the importance of pages in Web social networks based on the centrality measures and employs the PageRank algorithm to accurately score Web pages.To verify the performance of the CentralRank algorithm,a Web crawler was developed to automatically and effectively crawl Web pages.The Web crawler contains three essential techniques,that is,Web data collection,content analysis and duplicate page detection.Experiments on real data show that the CentralRank algorithm can guarantee less time deficiency and is more exact in scoring Web pages than the centrality measures-based page ranking algorithm and the PageRank algorithm with an average improvement of 14.2% and 7.5%,respectively.
Applying the centrality measures from social network analysis to score web pages may well represent the essential role of pages and distribute their authorities in a web social network with complex link structures. To effectively score the pages, we propose a hybrid page scoring algorithm, called WebRank, based on the PageRank algorithm and three centrality measures including degree, betweenness, and closeness. The basis idea of WebRank is that: (1) use PageRank to accurately rank pages, and (2) apply centrality measures to compute the importance of pages in web social networks. In order to evaluate the performance of WebRank, we develop a web social network analysis system which can partition web pages into distinct groups and score them in an effective fashion. Experiments conducted on real data show that WebRank is effective at scoring web pages with less time deficiency than centrality measures based social network analysis algorithm and PageRank.
Finding relational expressions which exist frequently in one class of data while not in the other class of data is an interesting work. In this paper, a relational expression of this kind is defined as a contrast inequality. Gene Expression Programming (GEP) is powerful to discover relations from data and express them in mathematical level. Hence, it is desirable to apply GEP to such mining task. The main contributions of this paper include: (1) introducing the concept of contrast inequality mining, (2) designing a two-genome chromosome structure to guarantee that each individual in GEP is a valid inequality, (3) proposing a new genetic mutation to improve the efficiency of evolving contrast inequalities, (4) presenting a GEP-based method to discover contrast inequalities, (5) giving an extensive performance study on real-world datasets. The experimental results show that the proposed methods are effective. Contrast inequalities with high discriminative power are discovered from the real-world datasets. Some potential works on contrast inequality mining are discussed.
Cluster analysis in web social networks is an important and challenging problem due to the rapid development of the Internet community, e.g., Facebook and Flickr. To accurately partition web social networks, we proposed a hierarchical clustering algorithm based on blockmodeling, called HCUBE, which employs structural equivalence to measure the similarity of web pages and reduce a large and incoherent network to a set of smaller comprehensible subnetworks. HCUBE uses the inter-connectivity as well as the closeness of clusters to group structurally equivalent pages in an effective fashion. Experiments conducted on real data show that HCUBE is effective at partitioning web social networks compared to the k-means based method.
The trajectory pattern mining problem has recently attracted much attention due to the rapid development of location-acquisition technologies, and parallel computing essentially provides an alternative method for handling this problem. This study precisely addresses the problem of parallel mining of trajectory sequential patterns based on the newly proposed concepts with regard to trajectory pattern mining. We propose an efficient and effective parallel sequential patterns mining (plute) algorithm that includes three essential techniques: prefix projection, data parallel formulation, and task parallel formulation. Firstly, the prefix projection technique is used to decompose the search space as well as greatly reduce the candidate trajectory sequences. Secondly, the data parallel formulation decomposes the computations associated with counting the support of trajectory patterns. Thirdly, the task parallel formulation employs the MapReduce programming model to assign the computations across a set of machines in a scalable and easy-to-use fashion. Based on the properties of parallel trajectory sequences, item pruning and sequence pruning strategies are applied to further prune the candidate sequences. Extensive experiments are conducted to evaluate the performance of plute in terms of parallel computing time and communication cost among processors. Experimental results show that plute outperforms the previously proposed parallel mining strategy (PartSpan) in mining massive trajectory data.
It is a new paradigm to apply data mining technologies to analyze the crime groups and terrorist social networks,there is little work being done on analyzing the communication behavior of criminal and terrorist groups.This paper designed a simulation email system based on personality trait dimensions,called MEP,to model the email users' traffic behavior,proposed a new approach of computing the weight of each dimension in a personality trait vector by using personality trait judge matrix,and simulated the real-world email communication behavior based on normal distribution model satisfying users' personality trait.This paper proposed a social network analysis based algorithm called CNKM(Crime Network Key Member mining) to mine key members of a crime group,and employed time-series analysis techniques to discover the email sending and receiving rules in order to detect the abnormal communication cases.The experimental results show the efficiency and usability of the simulation email analysis system,the average simulation error is less than 10%,and demonstrate that CNKM is efficient.
Objective To provide evidence for the establishment of an essential medicines list,we investigated the institutional medicine supply in rural hospitals and community health service centers in Chengdu.Methods The trained investigators collected medicine sales records and information about the management of institutional pharmacies. Through in-depth interviews with the pharmaceutical personnel,we inquired into the drug supervision and supply networks in rural areas.Then we performed secondary research based on a comparative analysis of drug classification, administration and pharmacies in developed countries.Results Seven township hospitals/community health service centers had pharmacies,facilities,storage,and a clean environment.Three of them used electrical databases to manage medicine sales records.Five township hospitals and 5 village medical rooms purchased medicines from the drug supervision and supply networks every week.In this way,they ensured the quality and accessibility of drugs in rural areas. In the urban community health service centers,medicines were supplied based on the traditional commercial distribution system.Conclusion Rational allocation of health resources to set up institutional pharmacies and village medicine rooms is important.The supervision of village medical rooms must be stricter.We should expand the use of electrical databases and integrate the supervision and supply networks with the supply system of the essential medicines.
In recent years, researchers have paid more and more attention on data mining of practical applications. Aimed to the problem of symptom classification of Chinese traditional medicine, this paper proposes a novel computing model based on the similarities among attributes of high dimension data to compute the similarity between any tuples. This model assumes data attributes as basic vectors of m dimensions and each tuple as a sum vector of all the attribute-vectors. Based on the transcendental concept similarity information among attributes, it suggests a novel distance algorithm to compute the similarity distance of any pair of attribute-vectors. In this method, the computing of similarity between any tuples are turned to the formulas of attribute-vectors and their projections of each other, and the similarity between any pair of tuples can be worked out by computing these vectors and formulas. This paper also presents a novel classification algorithm based on the similarity computing model and successfully applies the algorithm into the symptom classification of Chinese traditional medicine. The efficiency of the algorithm is proved by extensive experiments.
Objective: Traditional Chinese Medicine (TCM) provides an alternative method for achieving and maintaining good health. Due to the increasing prevalence of TCM and the large volume of TCM data accumulated though thousands of years, there is an urgent need to efficiently and effectively explore this information and its hidden rules with knowledge discovery in database (KDD) techniques. This paper describes the design and development of a knowledge discovery system for TCM as well as the newly proposed KDD techniques integrated in this system. Methods: A novel Knowledge dIscovery System for TCM (KISTCM) is developed by incorporating several data mining techniques, primarily including a medicine dependency relationship discovery algorithm, an efficacy dimension reduction algorithm based on neural networks, a method for exploring the relationships between formulae and syndromes using gene expression programming (GEP), and an approach for discovering the properties in terms of nature, taste and meridian based on the herbal dosage by employing the effect degree function to calculate the effect of each property. Results: Representative experimental cases are used to evaluate the system performance. Encouraging results are obtained, including rules previously unknown to algorithm designers and experiment runners. Experiments demonstrate that KISTCM has powerful knowledge discovery and data analysis capabilities, and is a useful tool for discovering the underlying rules in formulae. Our proposed techniques successfully discover hidden knowledge from TCM data, which is a new direction in knowledge discovery. From TCM experts' perspective, the accuracy of data analysis for KISTCM is an improvement, and these results compare favorably to other existing TCM data mining techniques. The system could be expected to be useful in the practice of TCM, e.g., assisting TCM physicians in prescribing formulae or automatically distinguishing between minister and assistant herbs in a formula.
This paper proposes a novel evolution algorithm, which is based on a new concept of chromosome hierarchy network in gene expression programming (CHN-GEP). This new algorithm is efficient for real applications, such as the function finding problem and electric circuit evolving. This paper expatiates four aspects about this algorithm: (1) details the algorithm CHN-GEP, based on CHN; (2) implements both a network-call model and a storage structure for CHN-GEP; (3) creates a novel method for converting artificial neural network (ANN) problems to chromosome network problems, and through this method, CHN-GEP solves quickly these problems; (4) Extensive experimentation shows that CHN-GEP can reduce the average evolution generations by 24–53% of the traditional GEP algorithms for function finding problems.
To solve the problem of reducing prescription effects of traditional Chinese medicine,the fuzzy neuron and the radial basis function was applied to the neural network,a prescription effect reduction algorithm named PERA(Prescription Effect Reduction Algorithm) based on fuzzy neural network was proposed,and a prescription effect reduction system named EFNN(Effect Fuzzy Neural Network) was developed.Experiments demonstrated that the proposed method is better than other traditional attribute reduction algorithms,such as artificial neural network and rough set.The precision of effect reduction of PERA is greater than 90%,the recall is about 40%,greater than traditional attribute reduction algorithms,and its running time is less than traditional neural networks obviously.
The paper proposes a new text similarity computing method based on concept similarity in Chinese text processing. The new method converts text to words vector space model at first, and then splits words into a set of concepts. Through computing the inner products between concepts, it obtains the similarity between words. The new method computes the similarity of text based on the similarity of words at last. The contributions of the paper include: 1) propose a new computing formula between words; 2) propose a new text similarity computing method based on words similarity; 3) successfully use the method in the application of similarity computing of WEB news; and 4) prove the validity of the method through extensive experiments.
In order to efficiently retrieve the k closest pairs between two spatial data sets in a specified space, such as in GIS and CAD applications, we propose a novel algorithm to handle the k-closest-pair range-query problem by progressively augmenting the query window instead of finding all objects in the whole space. We first describe a specific range estimation method to compute the circle query range which helps eliminate the unnecessary distance calculations among spatial objects and improve performance. Then, we use R*-tree to store closest pairs and give algorithms for maintaining this structure. Extensive experiments performed with synthetic as well as with real data sets show that the new algorithm outperforms the existing approaches in most cases. In particular, this technique works well when two spatial data sets are identical.
Carrent management systems of transient population support data query wid simple statistic, and don't have any firaction of deep data analysis. So it is difficult to help people making decision. This paper proposes a new intelligent analysis model, named TP-Miner (Transient Population Miner) that based on intelligence computing in transient population analysis. The system integrates new frontline intelligence computing technologies and creates a transient population, analysis platform, based on. the technologies. Meanwhile, as practice process needs, this study proposed a new classification. alarm, model based on, intelligent analysis, and it gives great help in criminal cases preventing or detecting. The extensive experiments show that this system has high performance and practicality.
Text classification is an important task of data mining. Existing algorithms, which based on vector space models, does not considered concept similarities among words, so the accuracy of traditional text classification cannot guarantee. To solve the problem, this paper proposes a new text classification algorithm in Chinese text processing based on concept similarity. The contributions of the paper include: (1) proposing a new similarity-computing model between words or sentences based on concept similarity; (2) applying the algorithm successfully in the text classification of WEB news; (3). analyzing the similarity computing formulas systematically in theory; (4).proving that the algorithm has much more accurate than traditional k-NN algorithm in text classification problems through extensive experiments.
According to the similarity among data attributes, this paper proposes a reduction method of high dimensional data. Different from the existed algorithms, this method takes attributes as basic vectors in high dimensions space, and data tuples as vector sum of attributes vectors. With the transcendental concept similar information between attributes, the weight computing is defined as formulas of attribute vectors and their projects on each other, and the final result is gotten from three simplifying algorithms which are proposed in this paper. The paper analyzes the new method, and compares reduced results between different algorithms. This method is successfully applied in automatic induction of Chinese traditional medicine prescription. The extension experiments prove the validity of the method.
To deal with the hierarchy coding data structure widely existed in application,this paper proposes a new conception of hierarchy distance and proves its mathematical properties.It also proposes and implements a new clustering algorithm-HDCA(Hierarchy Distance Computing based clustering Algorithm) based on hierarchy distance.The new algorithm overcomes the shortage of traditional algorithm and improves the precision.The paper also proposes a fast algorithm to compute the median of a hierarchy coding data set,and gives a clear proof of the algorithm.Extensive experiments demonstrate that HDCA is much faster than the naive algorithm to compute the median of hierarchy coding data.The new algorithm has been applied in the data analysis of transient population for public security successfully.
Changjie Tang (唐常杰)合作论文数College of Computer Science, Sichuan University34