The large availability of tabular Open Data sources with hundreds of attributes and relations makes the query development a difficult task, where analytic queries are common. When writing such queries, often called SPJG (Select-Project-Join-GroupBy), it is necessary to understand a data model and to write JOIN operations. The most common approach is to use business intelligence frameworks, or recent solutions based on keywords or examples. However, they require the utilization of specific applications and there is a lack of support for web-based APIs. We present a solution that eases the task of query development for tabular Open Data analytics through an API, using a simplified query representation where it is not allowed to specify the data relations, and consequently neither the joins over them, called Relation-Free Query. We define a single virtual schema that captures the database structure, which allows the use of relation-free queries in existent DBMS's. The concrete queries are exposed by a RESTful API, which is then translated into a database query language using known query generation solutions. The API is available as a microservice. We present a case study to describe solution, using a real world scenario to query in an integrated database of several Brazilian open databases with hundreds of attributes.
We propose a stochastic framework to evaluate the impact of missing data on the performance of predictive models. The framework allows full control of important aspects of the data set structure. These include the number and type of the input variables, the correlation between the input variables and their general predictive power, and sample size. The missing process is generated from a multivariate Bernoulli distribution, which allows us to simulate missing patterns corresponding to the MCAR, MAR and MNAR mechanisms. Although the framework may be applied to virtually all types of predictive models, in this article, we focus on the logistic regression model and choose the accuracy as the predictive measure. The simulation results show that the effects of missing data disappear for large sample sizes, as expected. On the other hand, as the number of input variables increases, the accuracy decreases mainly for binary inputs.
Several studies confirm the greater efficiency of estimators based on ranked set sampling (RSS) when compared to the corresponding ones obtained via simple random sampling (SRS). Recently, ranked set sampling has been considered in statistical process control. In this work, we propose the construction of adaptive control charts based on ranked set sampling. The adaptive strategy uses variable sample sizes and multiple dependent state sampling. Through an extensive simulation study, we verified that the proposed control charts have better performance (lower average number of samples until an out-of-control signal) when compared to the non-adaptive framework via RSS and adaptive via SRS. An illustration with simulated data complements this study.
While several public institutions provide its data openly, the effort required to access, integrate and query this data is too high, reducing the amount of possible dataset users. The Blended Integrated Open Data (BIOD) project has as objective to ease the access to public Open Data. It integrates and makes available more than 300Gb of data, containing billions of records from different Open Data Sets, allowing to query over them, and thus to retrieve related information from originally disconnected data sets. This paper presents the set of open data available, how to access it and how produce new compatible data to improve the existing data set.
Luiz Oliveira合作论文数PPGIA, Pontifical Catholic University of Parana, Rua Imaculada Conceicao, 1155, PR 80215-901, Curitiba, Brazil1