Erasable-itemset mining is a valuable method of pattern extraction for helping the manager of a factory analyze production planning. The erasable itemsets derived can be considered important production information regarding how to plan the production of a factory during an economic depression or financial shortage for the manager. After the erasable-itemset mining was proposed in 2009, several efficient mining approaches for finding erasable itemsets have been developed. However, these methods require a considerable amount of execution time when the amount of product data is large. Especially, manufacturing small amounts of versatile products has been a trend today, and it will generate a large product database. This paper adopts a bitmap representation for itemsets in an erasable-itemset mining algorithm to speed up the execution. Unlike the traditional bitmap meaning for frequent itemsets, a bitmap for erasable-itemset mining here denotes the relationship that a product includes at least one material (item) in a specified itemset. Using the bitmap representation can easily find the desired products to check, thus decreasing the scans of a database. Experimental evaluation on synthesized and real datasets was used to compare the proposed approach with the other two under different parameter values. The experimental results show that the proposed approach can make a good trade-off between execution time and memory usage.
This paper proposes a bitmap representation approach for modifying the erasable-itemset mining algorithm to increase its efficiency. The proposed approach uses the bitmap concept to save processing time. Through experimental evaluation, simulation datasets were used to compare the traditional erasable itemset mining and the proposed approach under various experimental conditions.
In this paper, we study the properties of partial periodic pattern mining and extend the original problem to high-utility partial periodic pattern mining (HUPPP), which considers not only the occurring time order and periodic length of events but also the quantities and individual profits of the events. Based on the periodic utility function, we have presented a mining algorithm for finding high-utility partial periodic patterns. The algorithm uses the two-phased periodic utility upper-bound (PUUB) model to avoid information loss in the mining process. Finally, the experiments made to show the performance of the proposed algorithm under various parameter settings.
In recent years, fuzzy utility mining has become an area of interest due to advancement of human reasoning. With regards to real applications, transactions in a database often involve things, such as transaction time, stamp, and much more. It is also noted that not all products in a store are displayed on the shelf, especially the seasonal ones. This paper, therefore, addresses these issues by presenting an effective framework called temporal-based fuzzy utility mining to give more attention to the transaction period of given items according to the concept of fuzzy utility mining. The temporal-based fuzzy utility mining proposed here is, however, a more complex approach when compared with the traditional fuzzy utility mining. A more complicated model for non-lost upper-bound fuzzy utility is thus proposed for effective mining. Furthermore, based on this model, a two-phase algorithm is developed for temporal-based fuzzy utility mining. Finally, the difference of fuzzy utility item sets with and without consideration of the lifetime of the items is shown by the experimental results under various experimental conditions.
Different from full periodic patterns, partial periodic patterns could ignore the occurrence of some events in time positions. In this paper, we have presented a gradually pruning algorithm (GPA) for reducing the number of candidate patterns in the mining process. It is based on the two-phased periodic utility upper-bound (PUUB) model and could avoid information loss. Compared to the original approach without gradually pruning, the one proposed here could reduce the execution time but get the same desired results.
Most of the existing studies in temporal data mining consider only lifespan of items to find general temporal association rules. However, an infrequent item for the entire time may be frequent within part of the time. We thus organize time into granules and consider temporal data mining for different levels of granules. Besides, an item may not be ready at the beginning of a store. In this paper, we use the first transaction including an item as the start point for the item. Before the start point, the item may not be brought. A three-phase mining framework with consideration of the item lifespan definition is designed. At last, experiments were made to demonstrate the performance of the proposed framework.
Data mining is the process of extracting desirable knowledge or interesting patterns from existing databases for specific purposes. In real-world applications, transactions may contain quantitative values and each item may have a lifespan from a temporal database. In this paper, we thus propose a data mining algorithm for deriving fuzzy temporal association rules. It first transforms each quantitative value into a fuzzy set using the given membership functions. Meanwhile, item lifespans are collected and recorded in a temporal information table through a transformation process. The algorithm then calculates the scalar cardinality of each linguistic term of each item. A mining process based on fuzzy counts and item lifespans is then performed to find fuzzy temporal association rules. Experiments are finally performed on two simulation datasets and the foodmart dataset to show the effectiveness and the efficiency of the proposed approach. (C) 2016 Elsevier B.V. All rights reserved.
In the past, an FUSP-tree algorithm for inserting customer sequences was proposed for handling the customer sequences insertion. In this paper, the FUSP-tree construction algorithm is thus modified for efficiently handling the deletion of customer sequences. A decremental FUSP-tree algorithm for sequences deletion (FUSP-DEL) is thus proposed for reducing the execution time of re-constructing the tree while the customer sequences are deleted in the original database. Experimental results show that the proposed FUSP-DEL algorithm has a good performance in both of the time and space complexity.
In the past, an incremental algorithm for mining high utility itemsets was proposed to derive high utility itemsets in an incrementally inserted way. In real-world applications, transactions are not only inserted into but also deleted from a database. In this paper, a maintenance algorithm is thus proposed for reducing the execution time of maintaining high utility itemsets due to transaction deletion. Experimental results also show that the proposed maintenance algorithm runs much faster than the batch approach.
Utility mining was proposed as an extension of frequent-itemset mining for concerning various factors from users. In this paper, a fast-updated high-utility-pattern trees for transaction deletion (FUHUP-DEL) algorithm is proposed to handle transaction deletion for efficiently updating discovered high utility itemsets in decremental mining. The HUP-tree structure is adopted in the proposed algorithm for reducing the computations of re-scan database in a level-wise way. Experiments show that the proposed FUHUP-DEL algorithm outperforms the batch two-phase approach.
Weighted itemset mining has been a widely studied topic in data mining. The reason is that weighted itemset mining considers not only the occurrence of items in transactions but also the individual importance of items. The traditional upper-bound model can be used to handle the weighted itemset mining problem, but a large number of candidates are generated by the model. This work thus presents an improved model to enhance the effectiveness of reducing unpromising candidates. Besides, an effective strategy, projection-based pruning, is proposed as well to tighten upper-bounds of weighted supports for itemsets in the mining process, thus reducing the execution time further. Through a series of experimental evaluation, the results on synthetic and real datasets show that the proposed approach has good performance in both pruning effectiveness and execution efficiency under various parameter settings when compared to some other approaches.
Fuzzy utility mining has been an emerging research issue because of its simplicity and comprehensibility. Different from traditional fuzzy data mining, fuzzy utility mining considers not only quantities of items in transactions but also their profits for deriving high fuzzy utility itemsets. In this paper, we introduce a new fuzzy utility measure with the fuzzy minimum operator to evaluate the fuzzy utilities of itemsets. Besides, an effective fuzzy utility upper-bound model based on the proposed measure is designed to provide the downward-closure property in fuzzy sets, thus reducing the search space of finding high fuzzy utility itemsets. A two-phase fuzzy utility mining algorithm, named TPFU, is also proposed and described for solving the problem of fuzzy utility mining. At last, the experimental results on both synthetic and real datasets show that the proposed algorithm has good performance. (C) 2015 Elsevier B.V. All rights reserved.
Weighted sequential pattern mining has recently been discussed in the field of data mining. Different from traditional sequential pattern mining, this kind of mining considers different significances of items in real applications, such as cost or profit. Most of the related studies adopt the maximum weighted upper-bound model to find weighted sequential patterns, but they generate a large number of unpromising candidate subsequences. In this study, we thus propose an efficient approach for finding weighted sequential patterns from sequence databases. In particular, a tightening strategy in the proposed approach is proposed to obtain more accurate weighted upper-bounds for subsequences in mining. Through the experimental evaluation, the results also show the proposed approach has good performance in terms of pruning effectiveness and execution efficiency.
On-shelf utility mining has recently received interest in the data mining field due to its practical considerations. On-shelf utility mining considers not only profits and quantities of items in transactions but also their on-shelf time periods in stores. Profit values of items in traditional on-shelf utility mining are considered as being positive. However, in real-world applications, items may be associated with negative profit values. This paper proposes an efficient three-scan mining approach to efficiently find high on-shelf utility itemsets with negative profit values from temporal databases. In particular, an effective itemset generation method is developed to avoid generating a large number of redundant candidates and to effectively reduce the number of data scans in mining. Experimental results for several synthetic and real datasets show that the proposed approach has good performance in pruning effectiveness and execution efficiency. (C) 2013 Elsevier Ltd. All rights reserved.
Partial periodic patterns are commonly seen in real-life applications and provide useful prediction with uncertainty. Most previous approaches have set a single minimum support threshold for all events to assume they have similar frequencies which is not practical for real-world applications. Instead of setting a single minimum support threshold for all events, Chen et al. proposed an FP-tree-like algorithm to allow multiple minimum supports for reflecting the natures of the events. However, such a tree-based algorithm encountered an efficiency problem while period length is long or event sequential orders in period segments are varied. Under the circumstance, many tree branches are created and much execution time is spent to find partial periodic patterns. In this paper, we thus propose a projection-based algorithm which examines only prefix subsequences and projects only corresponding postfix subsequences with multiple minimum supports to quickly find the partial periodic patterns in a recursive process. Experiments on both synthetic and real-life datasets show that the proposed algorithm is more efficient than the previous one.
Partial periodic patterns are commonly seen in real-world applications. The major problem of mining partial periodic patterns is the efficiency problem due to a huge set of partial periodic candidates. Although some efficient algorithms have been developed to tackle the problem, the performance of the algorithms significantly drops when the mining parameters are set low. In the past, the authors have adopted the projection-based approach to discover the partial periodic patterns from single-event time series. In this paper, the authors extend it to mine partial periodic patterns from a sequence of event sets which multiple events concurrently occur at the same time stamp. Besides, an efficient pruning and filtering strategy is also proposed to speed up the mining process. Finally, the experimental results on a synthetic dataset and real oil price dataset show the good performance of the proposed approach.
Recently, privacy-preserving data mining (PPDM) has become a critical issue to hide the sensitive or private information through data sanitation process, especially for hiding sensitive association rules or frequent itemsets. In this paper, a privacypreserving sequential pattern mining (PPSPM) is thus proposed to hide the sensitive sequences for data sanitization. The side-effects of hiding failure and missing rules are concerned in the proposed algorithm for evaluating the sequences to be deleted. Experiments are also conducted to evaluate the performance of the proposed approach. © 2014 ISSN 1881-803X.
This work presents an efficient approach for deriving itemsets with high fuzzy utility values from quantitative data. Each item in a transaction has its own profit and quantity, and the total fuzzy utility of it is considered. We also design a useful strategy to prune unpromising fuzzy candidate itemsets, thus making the mining process efficient. Through a series of experimental evaluations, the results show the proposed approach could perform well in fuzzy utility mining.
In traditional association rule mining, most algorithms are designed to discover frequent itemsets from a binary database. Utility mining was thus proposed to measure the utility values of purchased items for revealing high utility itemsets from a quantitative database. In the past, a two-phase high utility mining algorithm was thus proposed for efficiently discovering high utility itemsets from a quantitative database. In dynamic data mining, transactions may be inserted, deleted, or modified from a database. In this case, a batch mining procedure must rescan the whole updated database to maintain the up-to-date information. Designing an efficient approach for handling dynamic databases is thus a critical research issue in utility mining. In this paper, an incremental mining algorithm is proposed for efficiently maintaining discovered high utility itemsets based on pre-large concepts. Itemsets are first partitioned into three parts according to whether they have large (high), pre-large, or small transaction-weighted utilization in the original database and in inserted transactions. Individual procedures are then executed for each part. Experimental results show that the proposed incremental high utility mining algorithm outperforms existing algorithms.
Most of the existing studies in utility mining use a single minimum utility threshold to determine whether an item is a high utility item. This way is, however, hard to reflect the nature of items. This work thus presents another viewpoint about defining the minimum utilities of itemsets. The maximum constraint is adopted, which is well explained in the text and suitable to some mining domains when items have different utility values. In addition, an effective two-phase mining approach is proposed to cope with the problem of multi-criteria utility mining under maximum constraints. The experimental results show the performance of the proposed approach.
Wen-Yang Lin合作论文数National University of Kaohsiung;Dept. of Computer Science and Information Engineering, 4