In this paper, we introduce EVE, a generic framework for event detection where events can also include outliers, model changes and drifts. Various methods for event detection have been proposed for different types of events. However, many of them make the same or very similar prior assumptions but use different notations and formalizations. EVE provides a general framework for event detection, which allows existing algorithms to be represented using a common basis. The framework includes generic types of time slots, streaming progresses, and measures of similarity between those slots. We demonstrate how existing algorithms fit nicely into this framework by instantiating appropriate window combinations, progress mechanisms, and similarity functions.
In this paper we describe how ensembles can be trained, modified and applied in the open source data analysis platform, KNIME. We focus on recent extensions that also allow ensembles, represented in PMML, to be processed. This way ensembles generated in KNIME can be deployed to PMML scoring engines. In addition ensembles created by other tools and represented as PMML can be applied or further processed (modified or filtered) using intuitive KNIME workflows.
We describe the application of a recently published general event detection framework, called EVE to the challenging task of molecular event detection, that is, the automatic detection of structural changes of a molecule over time. Different types of molecular events can be of interest which have, in the past, been addressed by specialized methods. The framework used here allows different types of molecular events to be systematically investigated. In this paper, we summarize existing molecular event detection methods and demonstrate how EVE can be configured for a number of molecular event types.
We present a novel approach in machine learning by combining naı̈ve Bayes classifiers with tree kernels. Tree kernel methods produce promising results in machine learning tasks containing treestructured attribute values. These kernel methods are used to compare two tree-structured attribute values recursively. Up to now tree kernels are only used in kernel machines like Support Vector Machines or Perceptrons. In this paper, we show that tree kernels can be utilized in a naı̈ve Bayes classifier enabling the classifier to handle tree-structured values. We evaluate our approach on three datasets containing tree-structured values. We show that our approach using tree-structures delivers significantly better results in contrast to approaches using non-structured (flat) features extracted from the tree. Additionally, we show that our approach is significantly faster than comparable kernel machines in several settings which makes it more useful in resource-aware settings like mobile devices. Naı̈ve Bayes Classifier; Tree Kernel; Lazy Learning; Tree-structured Values
An interesting challenge in data stream mining is the detection of events where events are generally defined as anything previously unknown in the data. Therefore outliers, but also model changes or drifts, can be considered as possible events. Various methods for event detection have been proposed for different types of events. In this paper, we describe a more general framework for event detection. The framework enables generic types of time slots and streaming progress through time to be incorporated. It allows measures of similarity to included between those slots, either based directly on the data, or an abstraction, e.g. a model built on the data. We demonstrate that a large number of existing algorithms fit nicely into this framework by choosing appropriate time slots, progress mechanisms, and similarity functions.
In this paper we introduce a modular, highly flexible, open-source environment for data generation. Using an existing graphical data flow tool, the user can combine various types of modules for numeric and categorical data generators. Additional functionality is added via the data processing framework in which the generator modules are embedded. The resulting data flows can be used to document, deploy, and reuse the resulting data generators. We describe the overall environment and individual modules and demonstrate how they can be used for the generation of a sample, complex customer/product database with corresponding shopping basket data, including various artifacts and outliers.
Andreas Nürnberger合作论文数Department for Technical & Operational Information Systems, Faculty of Computer Science, Otto-Von-Guericke-University Magdeburg1
Lars Schmidt-Thieme合作论文数Institute of Computer Science, Department of Mathematics, Natural Science, Economics and Computer Science, University of Hildesheim1
Steffen Oeltze合作论文数Department of Simulation and Graphics
Faculty of Computer Science
University of Magdeburg1
Bernhard Preim合作论文数Department of Simulation and Graphics, University of Magdeburg, Germany1