
Most large public datasets containing cyclists for training detectors based on Deep Learning have annotations for bicycles and people, but not for cyclists. Even when it is not the case, the quality and quantity of the images are limited. To overcome these limitations, we propose the new OpenImages Cyclists dataset, built through the pre-selection of images from the OpenImages set and a new algorithm for semiautomatic generation of cyclist annotation aided by people and bicycle detectors. A cyclist detector trained with this dataset achieved identification rates up to 78% and 89% in two different sets of images obtained from security cameras at USP, Campus São Paulo - Capital.
The concept of Smart Cities has gained relevance, especially in the last decade, due to the availability of data associated with cities, e.g., car traffic, public transportation, crime data, etc. The purpose of using these data is to improve the services offered to the citizens. Most of these applications manipulate spatiotemporal data. These data are processed in a dataflow that starts with the collection, integration, and aggregation and ends with visualization. This way, specialized data services for smart city applications are most welcome. However, many of the existing data services in this context, are either specific to a particular application/domain or do not consider the entire data life cycle. In this article, we present Hurricane, a dataflow-oriented data service for smart city applications. Hurricane executes multiple dataflows to gather, pre-process, integrate, and public data. Hurricane was evaluated with an application in the area of public security and results reinforced the importance of this type of data service.
A wide range of applications has used semi-structured data. A characteristic of this type of data is its flexible structure, i.e., it does not rely on schema-based constraints to define its entities. Usually entities of a same kind (i.e, class) do not present the same attribute set. However, some data processing and management applications rely on a data schema to perform their tasks. In this context, the lack of structure is a challenge for these applications to use this data. In this paper, we propose CoFFee, an approach to class schema discovery. Given a set of heterogeneous entity schemata, found within a class, CoFFee provides a summarized set with core attributes. To this end, CoFFee applies a strategy combining attributes co-occurrence and frequency. It models a set of entity schemata as a graph and uses centrality metrics to capture the co-occurrence between attributes. We evaluated CoFFee using data from 12 classes extracted from DBpedia and e-Commerce datasets. We benchmarked it against two other state-of-the-art approaches. The results show that: i) CoFFee effectively provides a summarized schema, minimizing non-relevant attributes without compromising the data retrieval rate; and ii) CoFFee produces a summarized schema of good quality, outperforming the baselines by an average of 19% of F1 score.
Based on openness and transparency for good governance, unimpeded and verifiable access to legal and regulatory information is essential. With such access, we can monitor government actions to ensure that public financial resources are not improperly or inconsistently used. This facilitates, for example, the detection of unlawful behavior in public actions, such as bidding processes and auctions. However, different public agencies have their own criteria for standardizing the models and formats used to make information available, as exemplified in the varying styles observed in municipal, state, and union (federal) documents. In this context, we aim to minimize the effort to deal with public documents, notably official gazettes. For this, we propose a structure-oriented heuristic for extracting relevant excerpts from their texts. We then characterize these excerpts through morphosyntactic analysis and entity recognition. Subsequently, we semantically classify the extracted fragments into "sections of interest" (e.g., bids, laws, personnel, budget) using an active learning strategy to reduce the manual labeling effort. We also improve the classification process by incorporating transformers, stacking, and by combining different types of representations (e.g., frequentist, static, and contextual semantic embeddings). Furthermore, we exploit oversampling based on semi-supervised learning to deal with (labeled) data scarceness and skewness. Finally, we combine all these contributions in a real-time annotation tool with active learning support that achieves 100% accuracy in extraction and an overall accuracy of 85% in classification with very little labeling effort.
In recent years, vast volumes of data are constantly being made available on the Web, and they have been increasingly used as decision support in different contexts. However, for these decisions to be more assertive and reliable, it is necessary to ensure data quality. Although there are several definitions for this area, it is a consensus that data quality is always associated with a specific context. This work aims to analyze data quality in a data warehouse with governmental information of the Brazilian state of Minas Gerais. We first present a brief comparison of eight open-source data quality tools and then choose the Great Expectations tool for analyzing such data in two real applications: public bids and public expenditure. Our analyses show that the chosen tool has relevant characteristics to generate good data quality indicators to reveal data quality issues that may directly impact the construction of final applications using such data.
Transformer architectures have become the main component of various state-of-the-art methods for natural language processing tasks, such as Named Entity Recognition and Relation Extraction (NER+RE). As these architectures rely on semantic (contextual) aspects of word sequences, they may fail to accurately identify and delimit entity spans when there is little semantic context surrounding the named entities. This is the case of entities composed only by digits and punctuation, such as IDs and phone numbers, as well as long composed names. In this article, we propose new techniques for contextual reinforcement and entity delimitation based on pre- and post-processing techniques to provide a richer semantic context, improving SpERT, a state-of-the-art Span-based Entity and Relation Transformer. To provide further context to the training process of NER+RE, we propose a data augmentation technique based on Generative Pretrained Transformers (GPT). We evaluate our strategies using real data from public administration documents (official gazettes and biddings) and court lawsuits. Our results show that our pre- and post-processing strategies, when used co-jointly, allows significant improvements on NER+ER effectiveness, while we also show the benefits of using GPT for training data augmentation.
The popularization of sensoring and connectivity technologies like 5G and IoT are boosting the generation of data streams. Such kinds of data are one of the last frontiers of data mining applications. However, data streams are massive and unbounded sequences of non-stationary data objects that are continuously generated at rapid rates. To deal with these challenges, the learning algorithms should analyze the data just once and update their classifiers to handle the concept drifts. The literature presents some algorithms to deal with the classification of multiclass data streams. However, most of them have high processing time. Therefore, this work proposes a XGBoost-based classifier called AFXGB-MC to fast classify non-stationary data streams with multiple classes. We compared it with the six state-of-the-art algorithms for multiclass classification found in the literature. The results pointed out that AFXGB-MC presents similar accuracy performance, but with faster processing time, being twice faster than the second fastest algorithm from the literature, and having fast drift recovery time.
Temporal networks have been widely used to model instances of a domain of interest and their time-evolving interaction, including modeling individuals and face-to-face contacts throughout time. In the context of infection spread, such individuals can, e.g., remain susceptible, recovered, or be infected at a particular time. Understanding the infection spread behavior (its speed and magnitude, for instance) is crucial for quick and reliable decision making. Network visualization strategies can help in this task as they allow easy identification of who infected whom and when, epidemics outbreak, and other relevant aspects. This paper presents a visualization approach for the simulation and analysis of infection spread dynamics that considers different infection probabilities and different levels of social distancing (inter-group interaction). We performed quantitative and visual experiments using three real-world social networks with distinct characteristics and from two different environments. Our findings reveal the overall influence of different levels of inter-group interaction and infection probabilities in the infection spread dynamics and also demonstrate the usefulness of our approach for enhanced local (individual- or group-level) investigations.
The growing availability of data in digital media has contributed to the creation of a large number of data ecosystems. However, having successful Data Ecosystem is still a challenge. In order to prevent the failure of a Data Ecosystem and ensure its survival, evaluating its health becomes fundamental. In a general way, the health of a Data Ecosystem can be defined as its ability to grow and survive over time. Indicators such as productivity, robustness, niche creation and sustainability can be employed to evaluate the health of a Data Ecosystem. In this paper, we propose a framework for data Ecosystem health evaluation composed of a set of indicators and metrics, which assess the Data Ecosystem’s current state and its ability to stay healthy over time. The results obtained when using the proposed framework offers evidence to assist in decision making on how data has being published and consumed in a Data Ecosystem, as well as to evaluate which ecosystems are more prosperous or need more investments.
This article introduces EERCASE, a Computer Aided Software Engineering tool that is based on the best practices of the Model Driven Development paradigm to provide a consistent environment for relational database design. EERCASE follows the graphical notation of the Enhanced Entity–Relationship model according to Elmasri and Navathe, implements the EERMM metamodel to avoid syntactically invalid constructs, shows and describes static semantic errors, and generates data definition code that takes into account advanced structural validations. The theoretical and technical framework used for the implementation of EERCASE is discussed, with emphasis on the restrictive and informative validations performed by it. In addition, considering feedbacks on modeling errors and code generation, EERCASE is also presented as a computational environment that favors active learning.
This work presents the design and implementation of two web-based search systems, Busc@NIMA and Quem@PUC. Both systems allow the identification of research and development projects, besides existing competencies in laboratories and departments involving professors and researchers at PUC-Rio University. Our applications are based on a list of search-related terms that are matched to the dataset composed of PUC-Rio’s Lattes CVs offered courses, information from administrative systems, and specific keywords that are input by the professors/researchers themselves. To integrate all the needed data, we consider multiple database and search technologies, such as XML, RDF, TripleStores, and Relational Databases. Search results include professor’s name, academic papers, teaching activities, contact links, keywords, and laboratories of those involved with the subject represented by the set of keywords input. We describe the main features that show how our systems work.
Scientific collaboration networks can present different views of researchers’ interactions. This work presents SCI-synergy, an online navigable artifact aiming to promote mechanisms and views of scientific collaboration networks. The artifact focuses on the researchers’ interaction in the co-authorship of publications considering intra- and interprogram relationships. SCI-synergy is developed upon the design science research paradigm using scientific publication data available on the large Digital Bibliography & Library Project (DBLP) repository. Official data from the Sucupira repository of six Brazilian graduate program members including Federal University of Minas Gerais (UFMG), State University of São Paulo (USP), Federal University of Rio Grande do Norte (UFRN), Federal University of Amazonas (UFAM), University of Brasília (UnB), and University of Vale do Rio dos Sinos (UNISINOS) is used. Data from these graduate programs illustrate the artifact usage regarding the scientific collaboration network of each program, how each researcher cooperates, and what relationship patterns exist in intra- and inter-programs views. We advocate that, even though it is necessary to consider data from each program’s history and current contextualization regarding politics, economics, and administration, the collaboration network views provided by SCI-synergy might help to understand collaboration network patterns.
Due to the exploratory nature of DNNs, DL specialists often need to modify the input dataset, change a filter when preprocessing input data, or fine-tune the models’ hyperparameters, while analyzing the evolution of the training. However, the specialist may lose track of what hyperparameter configurations have been used and tuned if these data are not properly registered. Thus, these configurations must be tracked and made available for the user’s analysis. One way of doing this is to use provenance data derivation traces to help the hyperparameter’s fine-tuning by providing a global data picture with clear dependencies. Current provenance solutions present provenance data disconnected from W3C PROV recommendation, which is difficult to reproduce and compare to other provenance data. To help with these challenges, we present Keras-Prov, an extension to the Keras deep learning library to collect provenance data compliant with PROV. To show the flexibility of Keras-Prov, we extend a previous Keras-Prov demonstration paper with larger experiments using GPUs with the help of Google Colab. Despite the challenges of running a DBMS with virtual environments, DL analysis with provenance has added trust and persistence in databases and PROV serializations. Experiments show Keras-Prov data analysis, during training execution, to support hyperparameter fine-tuning decisions, favoring the comparison, and reproducibility of such DL experiments. Keras-Prov is open source and can be downloaded from https://github.com/dbpina/keras-prov.
User reviews are readily available on the Web and widely used for sentiment analysis tasks. Sentiment lexicons plays an important role in sentiment analysis, where each sentiment word is given a sentiment label (positive or negative) or score (1 or -1). However, a sentiment lexicon may express different sentiment polarity according different domain. In addition, only a few studies on Portuguese sentiment analysis are reported due to the lack of resources including domain-specific sentiment lexical corpora. In this paper, we present an effective methodology, called SentiLexBR, using probabilities of the Bayes’ Theorem for building a set of sentiment lexicons. An unsupervised algorithm is proposed to automatically identify sentiment lexicons with their polarities for the Portuguese language. Experimental results on user reviews datasets in 12 different domains indicate the effectiveness of our methodology in domain-specific sentiment lexicon generation for Portuguese. In addition, the sentiment lexicon produced by SentiLexBR also significantly outperforms several alternative approaches of building domain-specific sentiment lexicons.
Competitiveness in the Oil and Gas (O&G) sector has required high technological investments for datacentric decisions. One of the trends is the adoption of Digital Twins (DTs), which use virtual spaces and advanced analytical services to monitor and improve physical spaces. Central to the interconnection of these systems is a Data Fusion Core (DFC) component, which provides data management capabilities. Although the literature has proposed data management functionality in the scope of specific O&G DT applications, different joint efforts towards standardization can be found to deal with data integration and interoperability in the industry. The Open Subsurface Data Universe (OSDU) data platform is an initiative by several partners members of The Open Group consortium created to eliminate data silos in the O&G ecosystem and leverage innovation through a data-driven approach. In this article, we look at the convergence of this effort in providing data management functionalities for digital twins, highlighting strengths, gaps, and opportunities. We investigated the extent to which the OSDU data platform meets the needs of a DFC implementation, with a focus on interoperability, integration, governance, and data lineage. We also propose additional resources for data management in this context, namely data enrichment, workflows, and data lineage. Our main contributions are: (i) analysis of possible data management capabilities for creating a working DFC for an O&G DT and (ii) initial ideas on the complementary role of OSDU data representation and ontologies and how this semantic enrichment can be leveraged in a DFC of a DT.
The evolution of technology has enabled scientists to advance the automation of scientific experiments. Many programming languages have become popular in the scientific environment, especially scripting languages, due to their high abstraction level and simplicity, allowing the specification of complex tasks in fewer steps than traditional programming languages. Due to these features, lots of scientists model their scientific experiments in scripting languages to ensure data management and results control. However, this type of experiment usually generates large volumes of data, making data analysis and threat mitigation difficult. To fill in this gap, we propose P+RProv, an approach to aid scientists in understanding the structure of Python scripts and their results.
Topic modeling approaches extract the most relevant sets of words (grouped into so-called topics) from a document collection. The extracted topics can be used for analyzing the latent semantic structure hiding in the collection. This task is intrinsically unsupervised (without information about the labels), so evaluating the quality of the discovered topics is challenging. To address that, different unsupervised metrics have been proposed, and some of them are close to human perception, e.g., coherence metrics. Moreover, metrics behave differently when facing noise (i.e., unrelated words) in the topics. This article presents an exploratory analysis to evaluate how state-of-the-art metrics are affected by perturbations in the topics. By perturbation, we mean that intruder words are synthetically inserted into the topics to measure the metrics’ ability to deal with noises. Our findings highlight the importance of overlooked choices in the metrics sensitiveness context. We show that some topic modeling metrics are highly sensitive to disturbing; others can handle noisy topics with minimal perturbation. As a result, we rank the chosen metrics by sensitiveness, and as the contribution, we believe that the results might be helpful for developers to evaluate the discovered topics better.
With the emergence of Big Data and the continuous growth of massive data produced by web applications, smartphones, social networks, and others, organizations began to invest in alternative solutions that would derive value from this amount of data. In this context, this article evaluates three factors that can significantly influence the performance of Big Data Hive queries: data modeling, data format and processing tool. The objective is to present a comparative analysis of the Hive platform performance with the snowflake model and the fully denormalized one. Moreover, the influence of two types of table storage file types (CSV and Parquet) and two types of data processing tools, Hadoop and Spark, were also comparatively analyzed. The data used for analysis is the open data of the Brazilian Army in the Google Cloud environment. Analysis was performed for different data volumes in Hive and cluster configuration scenarios. The results yielded that the Parquet storage format always performed better than when CSV storage formats were used, regardless of the model and processing tool selected for the test scenario.
Automatic Speech Recognition (ASR) is essential for many applications like automatic caption generation for videos, voice search, voice commands for smart homes, and chatbots. Due to the increasing popularity of these applications and the advances in deep learning models for transcribing speech into text, this work aims to evaluate the performance of commercial solutions for ASR that use deep learning models, such as Facebook Wit.ai, Microsoft Azure Speech, Google Cloud Speech-to-Text, Wav2Vec, and AWS Transcribe. We performed the experiments with two real and public datasets, the Mozilla Common Voice and the Voxforge. The results demonstrate that the evaluated solutions slightly differ. However, Facebook Wit.ai outperforms the other analyzed approaches for the quality metrics collected like WER, BLEU, and METEOR. We also experiment to fine-tune Jasper Neural Network for ASR with four datasets different with no intersection to the ones we collect the quality metrics. We study the performance of the Jasper model for the two public datasets, comparing its results with the other pre-trained models.
The amount of data daily generated by different sources grows exponentially and brings new challenges to the information technology experts. The recorded data usually include heterogeneous attribute types, such as the traditional date, numerical, textual, and categorical information, as well as complex ones, such as images, videos, and multidimensional data. Simply posing similarity queries over such records can underestimate the semantics and potential usefulness of particular attributes. In this context, the Exploratory Data Analysis (EDA) technology is well-suited to understand data and perform knowledge extraction and visualization of existing patterns. In this paper, we propose Sketch+ , a technique and a corresponding supporting tool to compare electronic health records (provided by hospitals) by similarity, supporting correlation-based exploratory analysis over attributes of different types and allowing data preprocessing tasks for visualization and knowledge extraction. Sketch+ computes partial and overall data correlation considering distance spaces induced by the attributes. It employs both ANOVA and association rules with lift correlations to study relationships between variables, allowing extensive data analysis. Among the tools provided, a pixel-oriented one drives the analysts to observe visual correlations among dates, categorical and numerical attributes. As a running case study, we employed three open databases of COVID-19 cases, showing that specialists can benefit from the inference modules of Sketch+ to analyze electronic records. The study highlights how Sketch+ can be employed to spot strong correlations among tuples and attributes, with statistically significant results. The exploratory analysis has been shown to be an essential complement for similarity search tasks, identifying and evaluating patterns from heterogeneous attributes.