Big Data are a paradigm through which valuable information is achieved through the analysis of a large amount of data. The sources of these data can be varied, from data streams that will be processed in real time, to the exploitation of transactional data stored in databases. For this last use, due to their scalability, the NoSQL databases, like mongoDB , a DBMS oriented to documents, have been consolidated as a powerful tool for the storage and processing of large volumes of data. On the other hand, information sources for Big Data algorithms can contain imprecise information, and the way to obtain, aggregate and present results can have an imprecise nature as well. For this reason, it is useful to provide fuzzy extensions to these DBMSs. In the case of MongoDB , there are few proposals and not very complete. This paper describes fzMongoDB, a fuzzy database engine that provides the mongoDB database with the capacity to store documents with imprecise information and to retrieve them in a flexible way. It is implemented and integrated on the mongoDB server using the resources it provides. The model and implementation of fzMongoDB also includes an indexing mechanism that accelerates the retrieval process on fuzzy queries. Also, the performance of these indexing mechanisms is evaluated.
Fuzzy association rules (FARs) are a recognized model to study existing relations among data, commonly stored in data repositories. In real-world applications, transactions are continuously processed with upcoming new data, rendering the discovered rules information inexact or obsolete in a short time. Incremental mining methods arise to avoid re-runs of those algorithms from scratch by re-using information that is systematically maintained. These methods are useful for extracting knowledge in dynamic environments. However, executing the algorithms only to maintain previously discovered information creates inefficiencies in real-time decision support systems. In this paper, two active algorithms are proposed for incremental maintenance of previously discovered FARs, inspired by efficient methods for change computation. The application of a generic form of measures in these algorithms allows the maintenance of a wide number of metrics simultaneously. We also propose to compute data operations in real-time, in order to create a reduced relevant instance set. The algorithms presented do not discover new knowledge; they are just created to efficiently maintain valuable information previously extracted, ready for decision making. Experimental results on education data and repository data sets show that our methods achieve a good performance. In fact, they can significantly improve traditional mining, incremental mining, and a naïve approach.
Predicting students’ academic performance is one of the oldest and most popular applications of educational data mining. It helps to estimate the unknown evaluation of a student’s performance. However, a huge amount of data with different formats and from multiple sources may contain a large number of features supposed as not-relevant that could influence the prediction results. The main objective of this paper is to improve the effectiveness of a predictive model for students’ academic performance. For this purpose, we propose a methodology to carry out a comparative study for evaluating the influence of feature selection techniques on the prediction of students’ academic performance. In our study, F-measure parameter is used to evaluate the effectiveness of the selected techniques. Two real data sources are used in this work, Mathematics and language courses. The outcomes are compared and discussed in order to identify the technique that has the best influence for an accurate predictive model.
This work presents an overview of the text mining area, considering the most common techniques, and including proposals based on the application of fuzzy sets. Besides, some of the most frequent text mining applications are mentioned. We discuss the existing approaches, which we call text data mining, in relation to the recently proposed paradigm of text knowledge mining, and we conclude that both are different and complementary, in the sense that they are able to extract different knowledge pieces from text by using different reasoning mechanisms. Future challenges related to text knowledge mining are also briefly outlined.
A wide spectrum of methods for knowledge extraction have been proposed up to date. These expensive algorithms become inexact when new transactions are made into business data, an usual problem in real-world applications. The incremental maintenance methods arise to avoid reruns of those algorithms from scratch by reusing information that is systematically maintained. This paper introduces a software tool: Data Rules Incremental Maintenance System (DRIMS) which is a free tool written in Java for incrementally maintain three types of rules: association rules, approximate dependencies and fuzzy association rules. Several algorithms have been implemented in this tool for relational databases using their active resources. These algorithms are inspired in efficient computation of changes and do not include any mining technique. We operate on discovered rules in their final form and sustain measures of rules up-to-date, ready for real-time decision support. Algorithms are applied over a generic form of measures allowing the maintenance of a wide rules’ metrics in an efficient way. DRIMS software tool do not discover new knowledge, it has been designed to efficiently maintain interesting information previously extracted.
Association Rules (ARs) and Approximate Dependencies (ADs) are significant fields in data mining and the focus of many research efforts. This knowledge, extracted by traditional mining algorithms becomes inexact when new data operations are executed, a common problem in real-world applications. Inc remental mining methods arise to avoid re-runs of those algorithms from scratch by re-using information that is systematically maintained. These methods are useful to extract knowledge in dynamic environments. However, the implementation of algorithms only to maintain previously discovered information creates inefficiencies. In this paper, two active algorithms are proposed for incremental maintenance of previous discovered ARs and ADs, inspired by efficient computation of changes. These algorithms operate over a generic form of measures to efficiently maintain a wide range of rule metrics simultaneously. We also propose to compute data operations at real-time, in order to create a reduced relevant instance set. The algorithms presented do not discover new knowledge; they are just created to efficiently maintain previously extracted valuable information. Experimental results in real education data and repository datasets show that our methods achieve a good performance. In fact, they can significantly improve traditional mining, incremental mining, and a naïve approach.
Tóm tắt. Các phụ thuộc hàm xấp xỉ và luật kết hợp là những tri thức thực sự có ý nghĩa trong khai phá dữ liệu. Trong bài báo này, đầu tiên, chúng tôi nhắc lại một số khái niệm cơ bản của lý thuyết tập thô, các độ đo lỗi g1, g2, g3 của phụ thuộc hàm. Sau đó, chúng tôi đề xuất độ đo lỗi g4 dựa trên phân hoạch và kỳ vọng trong lý thuyết xác suất. Phần tiếp theo chúng tôi xây dựng ma trận phân biệt theo một cách khác và biểu diễn các độ đo lỗi g1, g2, độ phụ thuộc γ và ý nghĩa thuộc tính σ theo ma trận phân biệt được. Cuối cùng, chúng tôi đưa ra mối liên hệ giữa phụ thuộc hàm xấp xỉ và luật kết hợp thông qua độ đo lỗi g4 và độ tin cậy Confidence. Từ khóa. Phụ thuộc hàm xấp xỉ, luật kết hợp
espanolIberdrola Ingenieria y Construccion participa desde hace cinco anos en cinco proyectos sobre el EPR de Flamanville 3, tanto en la isla nuclear como en la isla convencional y la estacion de bombeo. Estos proyectos representan un reto desde el punto de vista tecnico debido a las altas exigencias aplicables al proyecto por el retorno de experiencia del explotador EDF asi como al cumplimiento de las nuevas reglamentaciones surgidas desde la realizacion de la ultima central nuclear en Francia. Este articulo presenta la descripcion de dichos proyectos, asi como su estado actual de avance. EnglishIberdrola Engineering & Construction is participating during the last 5 years in 5 projects on the Flamanville 3 EPR, both in the nuclear island and conventional island and the pump house. These projects represent a challenge from the technical point of view due to the high requirements applicable to the project because of the experience feedback of the operator EDF and of compliance with new regulations that have emerged since the completion of the last nuclear power station in France. This paper presents the description of these projects, as well as its current status.
Several applications to represent classical or fuzzy data in databases have been developed in the last two decades. However, these representations present some limitations specially related with the system portability and complexity. Ontologies provides a mechanism to represent data in an implementation-independent and web-accessible way. To get advantage of this, in this paper, an ontology, that represents fuzzy relational database model, has been redefined to communicate users or applications with fuzzy data stored in fuzzy databases. The communication channel established between the ontology and any Relational Database Management System (RDBMS) is analysed in depth throughout the text to justify some of the advantages of the system: expressiveness, portability and platform heterogeneity. Moreover, some tools have been developed to define and manage fuzzy and classical data in relational databases using this ontology. Even an application that performs fuzzy queries using the same technology is included in this proposal together with some examples using real databases.
Two main data models are currently used for representing knowledge and information in computer systems. Database models, especially relational databases, have been the leader in last few decades, enabling information to be efficiently stored and queried. On the other hand, ontologies have appeared as an alternative to databases in applications that require a more `enriched' meaning. However, there is controversy regarding the best information modeling technique, as both models present similar characteristics. In this paper, we present a review of how ontologies and databases are related, of what their main differences are and of the mechanisms used to communicate with each other.
This document is the result of a Strategic Workshop held in Leuven by the Task Force e-Learning (TF eL) of the Coimbra Group (CG), followed by a series of meetings with the TF members on this issue. It provides short information on the current state of play on learning technologies and related practices at the different institutions of TF eL. This brief focuses on the current knowledge and experience regarding the good use of educational technology at our institutions and describes what it is, why it matters, how it works and where it is going.
Fuzzy data management in databases is a complex process because of flexible data nature and heterogeneous database systems. A solution to this problem has been solved using an ontology which isolates the fuzzy database representation of their management platform making fuzzy schemas Relational Database Management System (RDBMS)-independent and Web-accessible. However, queries performed on fuzzy databases present similar problems. In this proposal an ontology which represents a query structure regardless of any RDBMS implementation is defined. This ontology allows generating and executing any query on fuzzy or classical data according to the system where it is executed.
Different communication mechanisms between ontologies and database (DB) systems have appeared in the last few years. However, several problems can arise during this communication, depending on the nature of the data represented and their representation structure, and these problems are often enhanced when a Fuzzy Database (FDB) is involved. An architecture that describes how such communication is established and which attends to all the particularities presented by both technologies, namely ontologies and FDB, is defined in this paper. Specifically, this proposal tries to solve the problems that emerge as a result of the use of heterogeneous platforms and the complexity of representing fuzzy data.
The Semantic Web has resulted in a wide range of information (e.g., HML, XML, DOC, PDF documents, ontologies, interfaces, forms, etc.) being made available in semantic queries, and the only requirement is that these are described semantically. Generic Web interfaces for querying databases (such as ISQLPlus©) are also part of the Semantic Web, but they cannot be semantically described, and they provide access to one or many databases. In this chapter, we will highlight the importance of using ontologies to represent database schemas so that they are easier to access. The representation of the fuzzy data in fuzzy databases management systems (FDBMS) has certain special requirements, and these characteristics must be explicitly defined to enable this kind of information to be accessed. In addition, we will present an ontology which allows the fuzzy structure of a fuzzy database schema to be represented so that fuzzy data from FDBMS can also be available in the Semantic Web.
In this paper, an ontology system is proposed to represent the knowledge structure enabling fuzzy information to be stored in fuzzy databases. This proposal allows users or applications to simplify the metadata definition process that is necessary for representing and managing imprecise and classic information in these databases. This ontology then acts as an interface that formalizes the representation of such structures and allows access to them. The instances obtained from this ontology represent the schemas that describe domain information in a database. The description of fuzzy and classic database schemas allows access to online public databases for which no other semantic description is associated. This paper also presents another ontology to represent these schemas as instances. Not only does this ontology allow fuzzy data values to be stored (because of the definition of fuzzy data types as classes of the ontology) but it also enables schema tables and attributes to be defined. © 2008 Wiley Periodicals, Inc.
In this paper we introduced an alternative view of text mining and we review several alternative views proposed by different authors. We propose a classification of text mining techniques into two main groups: techniques based on inductive inference, that we call text data mining (TDM, comprising most of the existing proposals in the literature), and techniques based on deductive or abductive inference, that we call text knowledge mining (TKM). To our knowledge, the TKM view of text mining is new though, as we shall show, several existing techniques could be considered in this group. We discuss about the possibilities and challenges of TKM techniques. We also discuss about the application of existing theories in possible future research in this field.
This paper presents GDB, a system to build, debug and test rules in a Fuzzy Relational Deductive Database System (FRDDS) in a visual way avoiding the associated syntax which could be quite complex for a non-expert user. The rules will be used to compute a Measure of Quality (MoQ) of the scientific data generated by the GIADA instrument (ROSETTA space mission). The MoQ will be defined and assigned taking into account the knowledge of the engineers who built the instrument and the formal documentation. Together with this MoQ, the tool will provide a semantic explanation about how the measure has been assigned. With GDB is possible to work with imprecise data, to store knowledge using rules and also to deduce new information. The MoQ and the explanation will provide to the scientists and engineers a better understanding about the instrument behavior and, therefore, about physical phenomena under study.
In this paper we deal with the problem of mining for approximate dependencies (AD) in relational databases. We introduce a definition of AD based on the concept of association rule, by means of suitable definitions of the concepts of item and transaction. This definition allow us to measure both the accuracy and support of an AD. We provide an interpretation of the new measures based on the complexity of the theory (set of rules) that describes the dependence, and we employ this interpretation to compare the new measures with existing ones. A methodology to adapt existing association rule mining algorithms to the task of discovering ADs is introduced. The adapted algorithms obtain the set of ADs that hold in a relation with accuracy and support greater than user-defined thresholds. The experiments we have performed show that our approach performs reasonably well over large databases with real-world data.
Daniel Sánchez合作论文数University of Granada4
Juan-Carlos Cubero合作论文数UNIVERSITY OF GRANADA;AND ARTIFICIAL INTELLIGENCE;DEPARTMENT OF COMPUTER SCIENCE2