Effective tool wear monitoring is of great importance for machining process. With existing deep learning-based methods, the end-to-end model is often combined with sensor data to predict the state of the tool wear. These methods lack an effective mechanism to identify and highlight the important tool wear information in different sensor monitoring data, which limits the monitoring accuracy of the model. Therefore, a new deep prediction framework, called a multi-input parallel convolutional attention network (MPCAN), is proposed in this paper and used for intelligent tool wear monitoring. A multi-input parallel convolutional (MPC) network structure is developed to extract multi-scale features from monitoring data of various sensors. Efficient channel attention (ECA) is used to assign different attention weights to the tool wear information. A bidirectional long short-term memory (Bi-LSTM) is used to obtain the time series features related to tool wear. Finally, the predicted tool wear value is output through the fully connected network. An experimental analysis has been undertaken to illustrate, the effectiveness and superiority of the proposed method in improving the accuracy of the tool wear.
Due to its low stiffness, the boring bar used in deep-hole-boring is prone to violent vibration during the cutting process. It is often inaccurate and inefficient to judge the vibration state of the boring bar through artificial experience. To detect the change of the vibration state of the boring bar over time, guide the adjustment of the processing parameters, and avoid wastage of the workpiece and the loss of equipment, it is particularly important to intelligently monitor the vibration state of the boring bar during processing. In this paper, the boring bar is taken as the research object, and an intelligent monitoring technology of the boring bar’s vibration state based on deep learning is proposed. Based on grouping convolution, channel shuffle, and BiLSTM, a shuffle-BiLSTM NET model is constructed, which is both lightweight and has a high classification accuracy. The boring experiment platform is built, and 192 groups of cutting experiments are carried out. The three-way acceleration and sound pressure signals are collected, and the signals are processed by smoothed pseudo-Wigner–Ville distribution. The original signals are transformed into a 256 × 256 × 3 matrix obtained by a two-dimensional time–frequency spectrum diagram. The matrix is input into the model to recognize the boring bar’s vibration state. The final classification accuracy is 91.2%. A variety of typical deep learning models are introduced for performance comparison, which proves the superiority of the models and methods used in this paper.
The sheer amount of open source codes made available in code repositories and code search engines along with the rapidly increasing releases of Application Programming Interfaces (APIs) have made code devel- opment process easier for programmers. However, learning how to use the elements of an API properly is both challenging and requires learning curve. Mining the available client and test codes can help programmers to iden- tify the best practices in using these APIs. In this paper, we investigate the API usage mining to identify open issues for the researchers. In particular, we make a theoretical comparison of the API usage pattern mining and highlight unresolved issues along with proper suggestions to address them.
The customer evaluation plays an important role for enterprise strategy decision-making. However, which indicators are key indicators for enterprise to know its customer well and how to calculate these indicators are still so challenging. In this paper, we have defined three classes contain ten indicators to evaluate customers who are loyal user of natural gas. We creatively designed the proximate and optimization indicators and imported data envelopment analysis (DEA) method to evaluate customers based on the classified customers by using classification algorithms. A system called Computer Support System for Customer Evaluation (CSSCE) has been developed and deployed for a natural gas company in China to evaluate its customers. The system satisfied user’s requirement. It provides a convenient, fast and objective tool for this company to save customer evaluation time, raises the delicacy manage lever and achieves great commendation from this company.
Big Data has attracted much attention in academia and industry fields in the last few years. The trend of utilizing medical big data has increased tremendously as well. The medical big data is generated from medical records, medicine, research, etc. but it is not well connected due to its long-term development in last decades in most hospitals without high-level strategic guide and plan. The information missing, data dispersivity, information isolated island, redundancy problems are key to the failure of the big data platform which is used to manage the medical data asset. We propose a platform called Data Mid-Platform (DMP) to solve the above problems and serve other big data projects as well as scientific assignments. The feasibility of this mid-platform is highly recommended and certified by several domain experts and it is under construction in a real project. It aims to satisfy at least 3 to 5 years big data applications development under the perspective of big data top-level design.
Recent years, the diabetes mellitus is an important public health problem and has been the top 10 leading causes of death in lower-middle-income countries and upper-middle-income countries in the world. In this study, we tried to use a hybrid feature selection approach to find proper and optimal feature subsets to classify and predict the diabetes mellitus patients in Korea based on the data from Korea National Health and Nutrient Examination Survey (KNHANES). We used the information gain feature selection approach as the filter phase and used the support vector machine with sequential search method as the wrapper phase. To validate the efficiency of the proposed approach, we also compared our proposed approach with several popular feature selection approaches. The results showed that our proposed approach can significantly improve the classification accuracy and outperformed other feature selection approaches.
Hadoop is a globally famous framework for big data processing. Data mining (DM) is the key technique for the discovery of the useful information from massive datasets. In our work, we take advantage of both platforms to design a real-time and intelligent mobile health-care system for chronic disease detection based on IoT device data, government-provided public data and user input data. The purpose of our work is the provision of a practical assistant system for self-based patient health care, as well as the design of a complementary system for patient disease diagnosis. This system was only applied to hypertensive disease during the first research stage. Nevertheless, a detailed design, an implementation, a clear overview of the whole system, and a significant guide for further work are provided; the entire step-by-step procedure is depicted. The experiment results show a relatively high accuracy.
Our world is becoming more interconnected and intelligent, huge amount of data has been generated newly. Home appliances' energy usage is the basis of home energy management and highly depends on weather condition and environment. Using weather in context, it is theorized that usage of home energy would be higher in cold days. Time series and contextual data collected from sensors can be monitored and controlled in home appliances network. The aim of this work is to propose a deep neural network architecture and apply it to a contextual and multivariate time series data. Long short-term memory (LSTM) models are powerful neural networks based on past behaviours in long sequences. LSTM networks have been demonstrated to be particularly useful for learning sequences containing longer-term patterns of unknown length, due to their ability to maintain long-term memory. In this work, we incorporate contextual features into the LSTM model because of ability of keeping context of data for a long-time, and for analysing it we integrated two different datasets; the first dataset contains measurements about house temperature and humidity measured over a period of 4.5 months by a 10 min intervals using a ZigBee wireless sensor network. The second dataset contains measurements about individual household electric power consumptions gathered over a period of 47 months. From the wireless network, the data from the kitchen, laundry and living room were ranked the highest in importance for the energy prediction.
Never before in history is the data growing at such a high volume, variety and velocity. It not only provides multi-sources of information for people to discover useful, important and valuable nuggets of information, but also increases the difficulty in finding such nuggets in almost all fields. Particularly, the field of healthcare is known for its dominical or ontological complexity and variety of clinical data or medical data regarding its variable data standards and data quality and so as the high data dimensionality. In order to effectively use the data at the hand to improve healthcare outcomes and processes, this paper illustrates a model called Risk Factor Detection and Disease Prediction (RFD-DP) model. The model incorporates statistics, data mining and MapReduce techniques on high dimensional clinical data to detect risk factors and generate predicator for a specified disease, hypertension disease. The experimental results indicate that the proposed model outperforms traditional feature selection and classification methods in terms of accuracy, F-score, and AUC. Consequently, the proposed model is promising to be applied to healthcare system.
Disease data provide an abundant source for chronic disease research. Hundreds of applications have been developed to deliver healthcare based on this big data. However, very few applications provide efficient chronic disease data visualization methods to better understand the results. This paper introduces a simple and practical way for visualizing the results of chronic disease detection and prediction. A model called IVIS4BigData has been used to implement the visualization procedure. This model not only demonstrates the historical data but also provides state-of-the-art visualization techniques. An exemplary set of scenarios corresponding to system design as well as visualization evaluation are given at last. Also we consulted several domain experts and common users about our visualization experimental results which satisfied their understanding about our systems. Finally conclusion and overlook of future work complete the paper.
Recently, various studies have shown that meaningful knowledge can be discovered by applying data mining techniques in medical applications, i.e., decision support systems for disease diagnosis. However, there are still several computational challenges due to the high-dimensionality of medical data. Feature selection is an essential pre-processing procedure in data mining to identify relevant feature subset for classification. In this study, we proposed a hybrid feature selection mechanism by combining symmetrical uncertainty and Bayesian network. As a case study, we applied our proposed method to the hypertension diagnosis problem. The results showed that our method can improve the classification performance and outperformed existing feature selection techniques.
Sensor networks integrated with cloud computing platform provides a promising way to develop monitoring system for elderly people especially dementia patients who need particular care for their normal life. In our work, we aim to design a comprehensive, unobtrusive, real-time, low-cost but effective monitoring system to help caregivers for their daily healthcare work. The system design has been finished and the entire experimental environment including sensor network and cloud model has been set up in our lab to collect simulated data and one health care center to collect real data. Though the system has been partially implemented due to time limit, the experiment results are encouraging.
Data Mining (DM) techniques such as classification, clustering, association, regression etc. are widely used in healthcare field in recent years to help improve the quality, efficiency as well as lowering the cost of developing healthcare systems. Especially, with the rapid development of the cloud platform services, which not only reduces the cost (time and expenditure), but also breaks the boundaries of data transactions among different systems and users. Therefore, it provides an effective way to reduce the time, and cost of software development as well as up-to-date services. In our work, we designed and partially implemented a healthcare system based on cloud services for disease detection and prediction using DM techniques in order to provide better services for both patients and health care givers. Although the system has been partially implemented, the experimental results are encouraging.
Data Mining (DM) techniques such as classification, clustering, association, regression etc. provide a promising way to help improving the quality, efficiency and lowing the cost of developing the healthcare systems. Especially with the rapid development of the cloud platform services, like SaaS, it does not only reduce the cost (time and expenditure) but also breaks the boundaries of the data transaction among different systems and users. In this work, we briefly described a healthcare system based on SaaS services for disease detection and prediction by using DM techniques so as to provide better service for both patients and health care givers. Promisingly, our work will provide a guideline for the next stage of research.
Apache Hadoop MapReduce is a well-known software framework for developing applications that process vast amounts of data. Combined with traditional Data Mining (DM) techniques, it provides a more powerful way to handle data with high speed, safety and accuracy. In our work, we took advantages of both Hadoop and DM techniques to design a comprehensive, real-time and intelligent mobile healthcare system for disease detection and prediction. It provides an assistant system for user selfhealthcare as well as a complementary system for doctors’ diagnosis on their daily work. Due to the time limit, the whole system has only been partially implemented, but the whole design work has been finished, the 4-node Hadoop experiment environment has been setup in the lab to do some experiments for further analysis and the experiment result is promising.
The prevention of hypertension is one of the most important topics in health research. In the most of the previous studies used statistical methods for analyzing the association between hypertension prevalence and dietary. However, statistical methods have some limitation which are, it is difficult to interpret variables interaction at a time. Thus we apply the data mining techniques for generation of prognosis factors based on association rule mining. In our experiment, we conducted Korean National Health and Nutrient Examination Survey (KNHANES) data from 2007 to 2014. We used to filter-based feature selection method for find prognosis factors and we generate the rules based on discovered risk factors of prognosis in hypertension. We evaluated discovered rules by support and confidence. In the results shows that, we can find useful rules for prognosis of hypertension. We expected to support medical decision making and easy to interpret prognosis of hypertension.
Token-based source code clones detection provides a promising way to detect the source code duplication and re-dundancy. While preprocessing of clone detection plays an important role in KDD for further processing as the old saying goes: well begun is half done. However, processing unstructured source code files of large software systems is really challenging and time or space consuming. This paper introduces a novel way to clean, tokenize and transform the source code into the appropriate form for mining. A tool called OPP (One Pass Preprocessor) has been developed to preprocess the source code files efficiently and flexibly. The paper experimented on three large open source projects like Wildfly1.02 Linux core-3.6, VTK of different host languages, and the result showed that our tool has great power and flexibility to preprocess the source code files and products high quality output.