
The growth of the mobile app market and the fitness market led to some interesting projects combining both areas to be developed over the last decade. In this paper we present PersonalFit an intelligent and adaptive prototype that helps gym clients to manage and adapt their training plans in a dynamic and personalized way, envisaging the achievement of better results. Three main modules compose PersonalFit: communications, Android Application and an intelligent plan generator that implements clustering techniques.
Automatic topic detection in document collections is an important tool for various tasks. In particular, it is valuable for studying and understanding socio-political phenomena. A currently relevant example is the automatic analysis of streams of posts issued by different activist groups in the current Brazilian turmoil, through the analysis of the generated streams of texts published on the web. It is useful to determine the relative importance of the different topics identified. We can find in the literature proposals for measuring topic relevance. In this paper, we adopt two of such measures and apply them to data sets extracted from Facebook pages related to Brazilian political activism. On top of the analysis, we then carry an experimental evaluation of the human interpretability for these two measures by comparing their outcomes with the opinion of three Brazilian professionals from the field of Communication Science and media-activists.
The automatic content analysis of mass media in the social sciences has become necessary and possible with the raise of social media and computational power. One particularly promising avenue of research concerns the use of sentiment analysis in microblog streams. However, one of the main challenges consists in aggregating sentiment polarity in a timely fashion that can be fed to the prediction method. We investigated a large set of sentiment aggregate functions and performed a regression analysis using political opinion polls as gold standard. Our dataset contains nearly 233 000 tweets, classified according to their polarity (positive, negative or neutral), regarding the five main Portuguese political leaders during the Portuguese bailout (2011-2014). Results show that different sentiment aggregate functions exhibit different feature importance over time while the error keeps almost unchanged.
Weather and sea-related forecasts provide crucial insights for the practice of nautical sports such as surf and kite surf, and mobile devices are appropriate interfaces for the visualization of meteorology and operational oceanography data. Data are collected and processed by several agencies and are often obtained from forecast models. Their use requires adaptation and refinement prior to visualisation. We describe a set of semantic data services using standard common vocabularies and interoperable interfaces following the recommendations of the INSPIRE directive. NautiCast, a mobile application for forecast delivery illustrates the adaptation of data at two levels: 1) semantic, with the integration of data from different sources via standard vocabularies, and 2) syntactic, with the manipulation of the spacial and temporal resolution of data to get effective mobile communication.
With the emergence of online Music Streaming Services (MSS) such as Pandora and Spotify, listening to music online became very popular. Despite the availability of these services, users face the problem of finding among millions of music tracks the ones that match their music taste. MSS platforms generate interaction data such as users' defined playlists enriched with relevant metadata. These metadata can be used to predict users' preferences and facilitate personalized music recommendation. In this work, we aim to infer music tastes of users by using personal playlist information. Characterizing users' taste is important to generate trustable recommendations when the amount of usage data is limited. Here, we propose to predict the users' preferred music feature's value (e.g. Genre as a feature has different values like Pop, Rock, etc.) by modeling, not only usage information, but also music description features. Music attribute information and usage data are typically dealt with separately. Our method FPMF (Feature Prediction based on Matrix Factorization) treats music feature values as virtual users and retrieves the preferred feature values for real target users. Experimental results indicate that our proposal is able to handle the item cold start problem and can retrieve preferred music feature values with limited usage data. Furthermore, our proposal can be useful in recommendation explanation scenarios.
Anurans (frogs or toads) are closely related to the ecosystem and they are commonly used by biologists as early indicators of ecological stress. Automatic classification of anurans, by processing their calls, helps biologists analyze the activity of anurans on larger scale. Wireless Sensor Networks (WSNs) can be used for gathering data automatically over a large area. WSNs usually set restrictions on computing and transmission power for extending the network's lifetime. Deep Learning algorithms have gathered a lot of popularity in recent years, especially in the field of image recognition. Being an eager learner, a trained Deep Learning model does not need a lot of computing power and could be used in hardware with limited resources. This paper investigates the possibility of using Convolutional Neural Networks with Mel-Frequency Cepstral Coefficients (MFCCs) as input for the task of classifying anuran sounds.
Currently, information security is a significant challenge in the information era because businesses store critical information in databases. Therefore, databases need to be a secure component of an enterprise. Organizations use Intrusion Detection Systems (IDS) as a security infrastructure component, of which a popular implementation is Snort. In this paper, we provide an overview of Snort and evaluate its ability to detect SQL Injection attacks.
Nowadays, online learning seems to be more and more in demand, with several offers being presented to the students. Online learning platforms that can be at any time that is convenient for them at any pace, using any technology available in any place. This also confers to the students the responsibility of been responsible for their own learning. To assist the students in this task the Emotional Adaptive Platforms (EAP) can be used. The EAP can perceived the student emotional state, personality and learning preference and adjust the learning course to it. So a question is asked how do we build an EAP? In this paper, is described how to build an EAP highlighting the major problems and concern issues from the initial research phase to the test of a working prototype. The research was carried out based on the assumption that emotion can influence several aspects of human life, knowledge acquisition, perception, learning process to the way people communicate and the way rational decisions are made.
Baseball is one of the best popular sports in Japan. A large number of baseball spectators are much interested in various predictions related to the game, such as starting players and outcome of the games. In this paper, we propose a heuristic method for predicting a starting pitcher. Predicting a starting pitcher using computers is difficult as as various aspects need to be considered and the volume of information is limited. The accuracy of prediction is low even by ardent followers of baseball. Our proposed method is modeled on the human method of prediction. The preliminary evaluation results show that the proposed method attains a prediction ratio which is equal to or higher than that obtained by the human method.
The Oil and Gas Exploration & Production (E&P) field deals with high-dimensional heterogeneous data, collected at different stages of the E&P activities from various sources. Over the years different soft-computing algorithms have been proposed for data-driven oil and gas applications. The most popular by far are Artificial Neural Networks, but there are applications of Fuzzy Logic systems, Support Vector Machines, and Evolutionary Algorithms (EAs) as well. This article provides an overview of the applications of EAs in the oil and gas E&P industry. The relevant literature is reviewed and categorised, showing an increasing interest amongst the geoscience community.
Social Media users tend to mention entities when reacting to news events. The main purpose of this work is to create entity-centric aggregations of tweets on a daily basis. By applying topic modeling and sentiment analysis, we create data visualization insights about current events and people reactions to those events from an entity-centric perspective.
Business Intelligence (BI) is a field where most of its tools and procedures have gradually changed with time: the advent rise of information circulation, and the more valuable it gets, combined with the sophistication of the technology surrounding the area, makes BI technology advancements seem much more of a reality. With the positive impact BI has had in businesses such as the improvement of the product's quality and the better customer behavior pattern analysis it brought to companies, BI is considered a standard way of analyzing data and of acquiring knowledge, so businesses know how to operate in the short, medium and long-term. This paper will discuss some of the new paradigms and perspectives being debated in the BI industry, as well as to bring into discussion new concepts that appear to be fully attainable, taking into account the current technology advancement, as well as the current discussion topics in the BI community.
Personalization can be defined as the customization of the outputs of a system based on the collected personal information of its users. Personalization techniques rely on user information, such as interests, preferences, geographic location, etc. The data being collected is used to create a profile, and improve the relevance of the outputs presented to the user. Google's search engine, or Facebook's suggestions are examples of personalization. This paper intent to provide an overview of the concept, and pretend to answer the question: Which level of personalization can Big Data Analytics Platforms support?
In the recent past, there has been a tremendous increase of large repositories of data, examples being in healthcare data, consumer data from retailers, and airline passenger data. These data are continually being shared with interested parties, either anonymously -- for research purposes, or openly by financial or insurance companies, for decision-making purposes. When is shared anonymously, there is still the possibility of de-anonymizing the data. Privacy Preserving Data Publishing (PPDP) is a way to allow one to share secure data while ensuring protection against identity disclosure of an individual. Generalization of attributes is a technique of data anonymization where an attribute is replaced with a more generalized value. Differential privacy is a technique that ensures the highest level of privacy for a record owner while providing actual information about the data set. This research develops a framework by generalizing attributes of a data set that satisfy differential privacy principles for publishing secure data for sharing. The proposed algorithm is a non-interactive method to publish anonymize data set, and the decision tree classifier showed better results compared to other existing classification works on anonymized data set. In this paper differential privacy refers to ϵ-differential privacy.
The growing number of different models and approaches for Geographic Information Systems (GIS) brings high complexity when we want to develop new approaches and compare a new GIS algorithm. In order to test and compare different processing models and approaches, in a simple way, we identified the need of defining uniform testing methods, able to compare processing algorithms in terms of performance and accuracy regarding large image processing, algorithms for GIS pattern-detection. Taking into account, for instance, images collected during a done flight or a satellite, it is important to know the processing cost to extract data when applying different processing models and approaches, as well as their accuracy (compare execution time vs. extracted data quality). In this work, we propose a GIS Benchmark (GPII), a benchmark that allows evaluating different approaches to detect/extract selected features from a GIS dataset. Considering a given dataset (or two data-sets, from different years, of the same region), it provides linear methods to compare different performance parameters regarding GIS information, making possible to access the most relevant information in terms of features and processing efficiency. Moreover, our approach to test algorithms makes possible to change the data-set in order to support different purpose algorithms.
Dealing with increasing amounts of data creates the need to deal with redundant, inconsistent and/or complementary repositories which may be different in their data models and/or in their schema. Current data cleaning techniques developed to tackle data quality problems are just suitable for scenarios were all repositories share the same model and schema. Recently, an ontology-based methodology was proposed to overcome this limitation. In this paper, this methodology is briefly described and applied to a real scenario in the health domain with data quality problems.
This paper describes a proposal to develop a Tourism Recommendation System based in users enhanced profiles (composed by basic user information, relations between user and a set of stereotypes and user functionality levels). The main focus of this work is to evaluate if user's physical and psychological functionality levels considered in user's profiles creation, will produce significant changes in the recommendation results. This work aims also to contribute with a different way to classify points-of-interest (POI) considering their capacity to receive tourists with certain levels of physical and psychological issues that will be described in this paper sections.
Research institutions are considering data repositories to manage their outputs and ensure their visibility. In many domains, purpose-built tools can help collect data and metadata as they are created. LabTablet is such a tool, designed to provide the functions of a laboratory notebook, and being able to accompany users in either experimental sessions or field trips. In these contexts, the interaction with the device can be problematic, so we experimented with a speech recognition extension for two purposes: to provide commands, such as requesting readings from the built-in sensors, and to record observations such as a dictated note in a field trip.
Music is important in our daily life not only for entertainment but also for mental health. When listening to music, playlists are used to eliminate the need for individual selection. The creation of playlist is difficult and tedious for users and has been the topic of research in many studies. However, many proposed playlist generation methods are based on either similar acoustic features or meta-data similarities. In this study, we propose a new method for music playlist recommendation using acoustic feature transitions where the next song will be selected such that it naturally transitions from the current song. Our preliminary evaluations show that the proposed method is more effective compared with other methods such as random selection and nearest neighbor methods
An electrocardiogram (ECG) system deals with several challenges related with noise sources. The denoising process is a challenge due to ECG signal amplitude with respect to noise amplitude. The filtering process can be executed by analogue filters or digital filters, however, modern systems uses digital filters mainly because their flexibility, efficiency and hardware costs reduction. This paper compares the ability of two digital techniques used in ECG denoising, namely Adaptive Filter (AF) and Singular Value Decomposition (SVD). The techniques were applied to real ECG signals contaminated with the most common noise sources; artefacts caused by electromyography (EMG) and power line noise (Hum).