When mining frequent itemsets (abbr. FIs) from dense datasets, too many itemsets are generated and results in the mining task from a large amount of execution time and high memory consumption. Frequent closed itemset (abbr. FCI) is a lossless and concise representation of FIs. Mining FCIs can not only greatly reduce the execution time and memory consumption, but also retain the complete information all of FI. Although many studies have proposed different mining FCI algorithms, but they have less developed methods that can effectively derive all FIs from FCIs. Form this point of view, this study proposes a novel efficient algorithm named DFI-List for efficiently deriving FIS from FCIs. The algorithm adopts the methodology of depth-first-search and divide-and-conquer to derive all FIs from FCIs. DFI-List efficiently derives all the FIs with vertical index structure called Cid List and uses SC Table to quickly find the support count of the derived FI. Experimental results show that the execution speed and memory consumption of the proposed algorithm with the proposed strategy is better than of the state-of-art algorithm.
High utility itemset mining is a popular pattern mining task, which aims at revealing all sets of items that yield a high profit in a transaction database. Although this task is useful to understand customer behavior, an important limitation is that high utility itemsets do not provide information about the purchase quantities of items. Recently, some algorithms were designed to address this issue by finding quantitative high utility itemsets but they can have very long execution times due to the larger search space. This paper addresses this issue by proposing a novel efficient algorithm for high utility quantitative itemset mining, called FHUQI-Miner (Fast High Utility Quantitative Itemset Miner). It performs a depth-first search and adopts two novel search space reduction strategies, named Exact Q-items Co-occurrence Pruning Strategy (EQCPS) and Range Q-items Co-occurrence Pruning Strategy (RQCPS). Experimental results show that the proposed algorithm is much faster than the state-of-art HUQI-Miner algorithm on sparse datasets.
When mining frequent itemsets (abbr. FIs) from dense datasets, it usually produces too many itemsets and results in the mining task to suffer from a very long execution time and high memory consumption. Frequent closed itemset (abbr. FCI) is a compact and lossless representation of FI. Mining FCIs can not only reduce the execution time and memory usage, but also reserve the complete information of FIs derived from FCIs. Although many studies have been proposed with various efficient methods for mining FCIs, few of them have developed algorithms for efficiently deriving FIs from FCIs. In this work, we propose two efficient algorithms named DFI-List and DFI-Growth for efficiently deriving FIs from FCIs. The both algorithms adopt depth-first search and divide-and-conquer methodology to derive all the FIs. DFI-List efficiently derives all the FIs with a vertical index structure called Cid List. DFI-Growth compresses the information of FCIs into tree structures and applies pattern-growth strategy to derive FIs from the trees. Empirical experiments show that DFI-List is the most efficient and scalable algorithm on the dense datasets. For example, when the minimum support threshold is set to 50% on the Chess dataset, DFI-List runs faster than LevelWise (Pasquier et al. Inf Syst 24(1): 25-46, 1999b) over 100 times. As for DFI-Growth, it is the most stable and memory efficient algorithm on the sparse datasets. Both DFI-Growth and DFI-List are superior to the state-of-the-art algorithm (Pasquier et al. Inf Syst 24(1): 25-46, 199b) in terms of execution time.
Partial Periodic itemsets are an important class of regularities that exist in a temporal database. A Partial Periodic itemset is something persistent and predictable that appears in the data. Past studies on Partial Periodic itemsets have been primarily focused on centralized databases and are not scalable for Big Data environments. One cannot ignore the advantage of scalability by using more resources. This is because we deal with large databases in a real-time environment and using more resources can increase the performance. To address the issue we have proposed a parallel algorithm by including the step of distributing transactional identifiers among the machines and mining the identical itemsets independently over the different machines. Experiments on Apache Spark’s distributed environment show that the proposed approach speeds up with the increase in a number of machines.
Ferrite magnetic nanoparticles (MNPs) with spinel structure are of great significance in the study of magnetic induction hyperthermia (MIH). The effect of element doping on the magnetic properties of MNPs is crucial. Here, we report the influence of the Mg2+ substitution on Curie temperature (T-c), magnetic properties, and heating efficiency of MgxZn0.8-xCo0.2Fe2O4 (0.1 <= x <= 0.5) nanoparticles synthesized by hydrothermal method. With the increase of Mg2+ content, T-c increases from 36.7 degrees C to 242.9 degrees C due to the enhancement of the A-B super-exchange interaction. The specific saturation magnetization increases from 35.5 emu.g(-1) to 53.3 emu.g(-1) and then remains constant, which is caused by the effect of Yafet-Kittle angle and magnetic moment. The specific absorption rate (SAR) increases from 3.5 W.g(-1) to 82.7 W.g(-1), which may be ascribed to the effect of size and specific saturation magnetization of the nanoparticles. When x= 0.3, the stable temperature under the alternating magnetic field (32 kA.m(-1),100 kHz) reaches 44.7 degrees C with the SAR of 49.0 W.g(-1). The low toxicity to cells and high heating efficiency endow the MNPs the potential in MIH. (C) 2020 Elsevier B.V. All rights reserved.
Finding Spatial High Utility Itemsets (SHUIs) in a spatiotemporal database is a challenging problem of great importance in many real-world applications. Most previous works focused on the sequential discovery of SHUIs in a database running on a single machine. Consequently, these works are not suitable for big data (or cloud-based) applications as they suffer from the scalability and fault tolerant problems. This paper proposes several novel pruning techniques to reduce the search space and present a more flexible distributed algorithm to find all desired itemsets from the database using Spark in-memory computing architecture. Our algorithm inherits several advantages of Spark, including low communication cost, fault tolerance, and high scalability. Experimental results demonstrate that the proposed algorithm has good scalability and performance on very large databases. Finally, we present a real-world navigation application in which SHUIs generated from the traffic congestion data have been employed to recommend alternative routes to the users.
Mining frequent itemsets (abbr. FIs) from dense databases usually generates a large amount of itemsets, causing the mining algorithms to suffer from long execution time and high memory usage. Frequent closed itemset (abbr. FCI) is a lossless condensed representation of FI. Mining only the FCIs allows to reducing the execution time and memory usage. Moreover, with correct methods, the complete information of FIs can be derived from FCIs. Although many studies have presented various efficient approaches for mining FCIs, few of them have developed efficient algorithms for deriving FIs from FCIs. In view of this, we propose a novel algorithm called DFI-Growth for efficiently deriving FIs from FCIs. Moreover, we propose two strategies, named maximum support selection and maximum support replacement to guarantee that all the FIs and their supports can be correctly derived by DFI-Growth. To the best of our knowledge, the proposed DFI-Growth is the first kind of tree-based and pattern growth algorithm for deriving FIs from FCIs. Experiments show that DFI-Growth is superior to the most advanced deriving algorithm [12] in terms of both execution time and memory consumption.
Sensitivity and linearity are important performance metrics of flexible sensors in the application. An effective approach toward improving the performance of capacitive pressure sensors (CPSs) is the appropriate design of the microstructure of the dielectric layer. Using the finite-element modeling in an integration of Abaqus and COMSOL Multiphysics, we propose a methodology to simulate the deformation and capacitance responses of CPS upon external pressure; the numerical results agree well with the experimental data. With the attempt to improve the performance of widely used pyramidal and cylindrical microstructure-based CPS, the effects of microstructure geometric parameters and mechanical property of materials, such as the elastic modulus, length of hemline, sidewall angle, height, and size on the pressure response, were investigated, and the sensitivity and nonlinear error were also analyzed. It has been found that the sensitivity and linearity are more sensitive to elastic modulus and are less sensitive to height. With the same sensitivity, the cylinders-based CPSs have higher linearity, while the pyramids-based CPSs have larger pressure measuring range. The obtained results could provide reference information for the design of CPS with improved application characteristics.
Recently, queuing recognition becomes an emerging research topic. Its main purpose is to recognize human’s queuing behaviors using sensor devices. However, previous studies take smartphones as the major sensor devices, causing it to suffer from inefficiency and inconvenience problems in supermarket scenarios. In view of this, this work takes account of queuing behaviors in supermarkets and proposes a novel framework for queuing recognition based on smart shopping carts. The proposed framework includes four main modules: (1) queuing micro-activity recognition, (2) queuing state recognition, (3) queuing group recognition (4) queuing time inference. We implement the proposed framework and conduct simulations for queuing scenarios to evaluate the proposed framework. Results show that the proposed framework has good efficiency and achieves high accuracy.
Top-k high utility itemset (abbr. Top-k HUI) mining aims at efficiently mining k itemsets having the highest utility without setting the minimum utility thresholds. Although some studies have been conducted on top-k HUI mining recently, they mainly focus on centralized databases and are not scalable for big data environments. To address the above issues, this paper proposes a novel framework for parallel mining of top-k high utility itemsets in big data. Besides, a new algorithm called PKU (Parallel Top-K High Utility Itemset Mining) is proposed for parallel mining of top-k HUIs on Spark in-memory platform. It adopts MapReduce architecture to divide the whole mining task into several independent subtasks, and takes good use of Spark in-memory computing technology for efficiently processing data in parallel. Moreover, several novel strategies are also proposed for pruning the redundant candidates such that the execution time and memory usage in the mining process are reduced greatly. The proposed PKU algorithm inherits several advantages of Spark, including low communication cost, fault tolerance, and high scalability. Experimental results on both real and synthetic datasets show that PKU has good scalability and performance on large datasets with outperforming several benchmarking algorithms.
Person identification (PID) is a key issue in many IoT applications. It has long been studied and achieved by technologies such as RFID and face/fingerprint/iris recognition. These approaches, however, have their limitations due to environmental constraints (such as lighting and obstacles) or require close contact to specific devices. Therefore, their recognition rates highly depend on use scenarios. To enable reliable and remote PID, in this work, we present EOY (Eye On You) 1 , a data fusion approach that combines two kinds of sensors, a 3D depth camera and wearable sensors embedded with inertial measurement unit (IMU). Since these two kinds of data share common features, we are able to fuse them to conduct PID. Further, the result can be transferred to a mobile platform (such as robot) since we have less constraints on devices. To realize EOY, we develop fusion algorithms to address practical challenges, such as asynchronous timing and coordinate calibration. The experimental evaluation shows that EOY can achieve the recognition rate of 95% and is very robust even in crowded areas.
In this paper, we propose a framework called conversational partner inference using nonverbal information (abbreviated as CFN). We use the wrist-based wearable device that has an accelerometer sensor to detect the user’s hand movement. Besides, we propose three different methods, named leading CFN, trainling CFN and leading-trailing CFN, to integrate the detected movement behaviors with the sound data sensed by microphones to effectively infer conservational partners. In experiments, we collect real data to evaluate the proposed framework. The experimental results show that the accuracy of leading CFN is better than trailing CFN and leading-trailing CFN. Moreover, our approach shows higher accuracy than the state-of-the-art approach for conversational partner inference.
Queuing recognition is a recently new raised research topic, which uses sensors of smartphones to automatically recognize human queuing behaviors. However, existing collaborative approaches need to exchange sensor data among nearby smartphones, causing extra communication overheads and even delay. In view of this, this work proposes a new framework called Qnalyzer for queuing recognition using accelerometer and Wi-Fi signals. It consists of three tiers. The first tier is run by each individual smartphone to identify each user's context without exchanging data with nearby smartphones. A new algorithm called QCF (Queuer and non-queuer ClassiFier) is proposed, which considers mixture features of accelerometer and Wi-Fi signals to effectively identify whether the user is queuing or not. The second tier is an algorithm called QCT (Queuers ClusTering) running at the server side to effectively identify which queuers belong to which queues based on users' movement features. The third tier is an estimation model called QPE (Queue Property Estimation) for measuring waiting time, service time, and queue lengths. The Qnalyzer prototype on Android smartphones and the corresponding performance evaluations under real-life queuing scenarios are implemented. The extensive experiment results show that Qnalyzer achieves good performance with high accuracy.
Detecting students’ attention in class provides key information to teachers to capture and retain students’ attention. Traditionally, such information is collected manually by human observers. Wearable devices, which have received a lot of attention recently, are rarely discussed in this field. In view of this, we propose a multimodal system which integrates a headmotion module, a pen-motion module, and a visual-focus module to accurately analyze students’ attention levels in class. These modules collect information via cameras, accelerometers, and gyroscopes integrated in wearable devices to recognize students’ behaviors. From these behaviors, attention levels are inferred for various time periods using a rule-based approach and a datadriven approach. The former infers a student’s attention states using user-defined rules, while the latter relies on hidden relationships in the data. Extensive experimental results show that the proposed system has excellent performance and high accuracy. To the best of our knowledge, this is the first study on attention level inference in class using wearable devices. The outcome of this research has the potential of greatly increasing teaching and learning efficiency in class. Keywords—Activity Recognition; Attention Sensing; Body-Area Network; Machine Learning; Wearable Computing
Recently, some studies for head gestures recognition have been proposed, but most of them are based on image processing technology. Moreover, there is less focus on the use of wearable devices in recognizing head gestures. In this paper, we apply machine learning techniques to recognize some common head activities using a head-mounted wearable device. We use the wearable device to collect sensor data related to user's head activities. Then, we apply energy-based segmentation method on the collected data to find out the data segments where the activities may occur. Finally, we extract candidate features from the segments and feed them into a pre-trained classifier to identify the type of head gesture. We implement the prototype of above methods on Arduino platform and evaluate the efficiency of the proposed methods on real datasets. The experiment results show that our proposed methods can effectively and efficiently identify different types of head gestures with an average accuracy rate of 95%.
Detecting students' attention in class provides key information to teachers to capture and retain students' attention. Traditionally, such information is collected manually by human observers. Wearable devices, which have received a lot of attention recently, are rarely discussed in this field. In view of this, we propose a multimodal system which integrates a head-motion module, a pen-motion module, and a visual-focus module to accurately analyze students' attention levels in class. These modules collect information via cameras, accelerometers, and gyroscopes integrated in wearable devices to recognize students' behaviors. From these behaviors, attention levels are inferred for various time periods using a rule-based approach and a data-driven approach. The former infers a student's attention states using user-defined rules, while the latter relies on hidden relationships in the data. Extensive experimental results show that the proposed system has excellent performance and high accuracy. To the best of our knowledge, this is the first study on attention level inference in class using wearable devices. The outcome of this research has the potential of greatly increasing teaching and learning efficiency in class.
High utility sequential pattern (HUSP) mining has emerged as a novel topic in data mining. Although some preliminary works have been conducted on this topic, they incur the problem of producing a large search space for high utility sequential patterns. In addition, they mainly focus on mining HUSPs in static databases and do not take streaming data into account, where unbounded data come continuously and often at a high speed. To efficiently deal with both problems, we propose a novel framework for mining high utility sequential patterns over static and streaming databases. In this regard, two efficient data structures named ItemUtilLists (Item Utility Lists) and HUSP-Tree (High Utility Sequential Pattern Tree) are proposed to maintain essential information for mining HUSPs in both offline and online fashions. In addition, a novel utility model called Sequence-Suffix Utility is proposed for effectively pruning the search space in HUSP mining. We propose an algorithm named HUSP-Miner (High Utility Sequential Pattern Miner) to find HUSPs in static databases efficiently. Then, a one-pass algorithm named HUSP-Stream (High Utility Sequential Pattern mining over Data Streams) is proposed to incrementally update ItemUtilLists and HUSP-Tree online and find HUSPs over data streams. To the best of our knowledge, HUSP-Stream is the first method to find HUSPs over data streams. Experimental results on both real and synthetic datasets show that HUSP-Miner outperforms the compared algorithms substantially in terms of execution time, memory usage and number of generated candidates. The experiments also demonstrate impressive performance of HUSP-Stream to update the data structures and discover HUSPs over data streams.
Recently some studies have shown that the major influential factor of our health is not only physical activities, but the states of our emotion that we experience through our daily life, which continuously build our behavior and affect our physical health significantly. Therefore, emotion recognition draws more and more attention for many researchers in recent years. In this paper, we propose a system that uses off-the-shelf wearable sensors, including heart rate, galvanic skin response, and body temperature sensors, to read physiological signals from the user and apply machine learning techniques to recognize emotional states of the user. These states are key steps, toward improving not only the physical health but also emotional intelligence in advanced human-machine interaction. Moreover, we consider three types of emotional states and conduct experiments on real-life scenarios. Experimental results shows that the proposed system has good performance for emotion recognition.
The purpose of time-dependent smart data pricing (abbreviated as TDP) is to relieve network congestion by offering network users different prices over varied periods. However, traditional TDP has not considered applying machine learning concepts in determining prices. In this paper, we propose a new framework for TDP based on machine learning concepts. We propose two different pricing algorithms, named TDP-TR (TDP based on Transition Rules) and TDP-KNN (TDP based on K-Nearest Neighbors). TDP-TR determines prices based on users' past willingness to pay given different prices, while TDP-KNN determines prices based on the similarity of users' past network usages. The main merit of TDP-TR is low computational cost, while that of TDP-KNN is low maintenance cost. Experimental results on simulated datasets show that the proposed algorithms have good performance and profitability.
Roger Nkambou合作论文数Département d'informatique, Faculté des sciences, Université du Québec à Montréal2