
The Web contains vast amounts of semi-structured data in the form of HTML tables found on Web pages which may serve for various applications. One prominent application, which is often referred to Semantic Table Interpretation, is to exploit the semantics of a widely recognized knowledge bases (KB) by matching tabular data, including column headers and cell contents, to semantically rich descriptions of classes, entities and properties in Web KBs. In this paper, we focus on relational tables which are valuable sources of facts about real-world entities (persons, locations, organizations, etc.) and we propose a robust and efficient approach for bridging the gap between millions of Web tables and large-scale Knowledge graphs such as DBpedia. Our approach is holistic and fully unsupervised for semantic interpretation of Web tables based on the DBpedia Knowledge graph. Our approach covers three phases that heavily rely on word and entity pre-trained embeddings to uncover semantics of Web tables. Our experimental evaluation is conducted using the T2D gold standard corpus. Our results are very promising compared to several existing approaches of annotation in web tables.
The paper focuses on calculating suitable place names and descriptive tags for large photo collections of visually interesting sights. The core dataset analyzed contains 45 million crowd-sourced geotagged pictures of the Panoramio database. We present several methods for analysis along with machine learning experiments for tag recommendation and suggest a manually built taxonomy of tag categories, based on the analysis of most widely used taglike words in the photo titles, along with their popularities. The methods, selected tags and the taxonomy can be used for building different tourism applications for visually interesting sights.
Ontology visualization is an important component in the support of human-ontology interaction, as it amplifies cognition and offloads cognitive efforts to the human perceptual system. While a significant amount of research efforts has focused on designing and developing various visual layouts and improve performance of large-scale visualizations, the differences in user preferences and cognitive abilities have been largely overlooked. This provides an opportunity to investigate ways to potentially provide more personalized visual support in human-ontology interaction. To this end, this paper demonstrates successful predictions on an individual user's likelihood to succeed in a given task, based on this person's gaze data collected during interaction. Specifically, we show several statistically significant predictions against a baseline classifier when inferring users' success before a given task is actually completed. Moreover, we present results showing that accurate predictions of user success can be achieved early on during user interaction, such as after just a few minutes in some cases. These findings suggest there are ample opportunities throughout various stages of human-ontology interaction where the underlying visual system may adapt in real time to the user's visual needs to provide the most appropriate visualization with the overall goal of possibly increasing user success in a given task.
The management of aviation data is a great challenge in the aviation industry, as they are complex and can be derived from heterogeneous data sources. To handle this challenge, ontologies can be applied to facilitate the modelling of the data across multiple data sources. This paper presents an aviation domain ontology, the ICARUS ontology, which aims at facilitating the semantic description and integration of information resources that represent the various assets of the ICARUS platform and their use. To present the functionality and usability of the proposed ontology, we present the results of querying the ontology using SPARQL queries through three use case scenarios. As shown from the evaluation, the ICARUS ontology enables the integration and reasoning over multiple sources of heterogeneous aviation-related data, the semantic description of metadata produced by ICARUS, and their storage in a knowledge-base which is dynamically updated and provides access to its contents via SPARQL queries.
Music recommender systems (RS) aim to aid people with finding relevant enjoyable music without having to sort through the enormous amount of available content. Music RS often rely on collaborative filtering methods, which however limits predicting capabilities in cold-start situations or for users who deviate from main-stream music preferences. Therefore, this paper evaluates various content-based music recommendation methods that may be used in combination with collaborative filtering to overcome such issues. Specifically, the paper focuses on the ability of lyrics-based embedding methods such as tf-idf, word2vec or bert to estimate songs similarity compared to state-of-the-art audio and meta-data based embeddings. Results indicate that both audio and lyrics methods perform similarly, which may favor lyrics-based approaches due to the much simpler processing. We also show that although lyrics-based methods do not outperform meta-data based approaches, they provide much more diverse, yet reasonably relevant recommendations, which is suitable in exploration-oriented music RS.
Whilst the CIA have been using psychometric profiling for decades, Cambridge Analytica showed that peoples psychological characteristics can be accurately predicted from their digital footprints, such as their Facebook or Twitter accounts. To exploit this form of psychological assessment from digital footprints, we propose machine learning methods for assessing political personality from Twitter. We have extracted the tweet content of Prime Minster Boris Johnsons Twitter account and built three predictive personality models based on his Twitter political content. We use a Multi-Layer Perceptron Neural network, a Naive Bayes multinomial model and a Support Machine Vector model to predict the OCEAN model which consists of the Big Five personality factors from a sample of 3355 political tweets. The approach vectorizes political tweets, then it learns word vector representations as embeddings from spaCy that are then used to feed a supervised learner classifier. We demonstrate the effectiveness of the approach by measuring the quality of the predictions for each trait per model from a classification algorithm. Our findings show that all three models compute the personality trait "Openness" with the Support Machine Vector model achieving the highest accuracy. "Extraversion" achieved the second highest accuracy personality score by the Multi-Layer Perceptron neural network and Support Machine Vector model.
Forecasting or predicting errors can dramatically reduce the downtime of machines in industrial settings and even allow to take counteractions long before the error affects the production system. A forecast system to predict upcoming critical values for identical production lines under different environmental circumstances is proposed. We focus on errors that result in multiple erroneous work pieces. These error patterns need manual corrections by a machine controller. An analysis of the system observed gathered the information about the types of errors that are observable. 30% of errors are measurement errors or single faulty work-pieces which are not influenced by previous work-pieces and do not show any indication to preceding work-pieces. These errors do not need any type of action by the machine controller. 70% of the observed errors are continuous system deviations which lead to multiple erroneous work-pieces in order or a high percentage of erroneous work-pieces in an observed time frame. We observe multiple production lines which consist of identical machines and produce the same product type. For the forecast of errors, we use the ARIMA, Holt and Holt-Winter method. Each production line and product type combination showed different results for the different forecast methods. We implemented a dynamic system that automatically detects the seasonality and trend of the specific combination to assign a correct forecast method and model. For 40 combinations of production line and product type the holt-winter algorithm performed best for 14, the holt-winter without seasonal or trend component performed best for 13 combinations and the holt-winter with only a trend component performed best for 10 setups. 3 combinations did not have a distinct best method for all observed results. By selecting the correct forecast methods, we were able to boost the forecast accuracy for the overall system over each single forecast method.
Merlyn TRN-1 is a set of precision machined parts so designed that they can be configured into interlinked linear and rotary axis as per your choice. The kit contains all the necessary mechanical structural elements, motors, motor driver electronics, microprocessor controllers and software. Each of them can be individually changed to upgrade your capacity and requirement. The modularity is an integral part of the kit. Merlyn TRN-1 modular design, high precision rugged parts with innovative tolerances and versatile interconnection of parts allow the user to add axis as per their design.
With the expanded use of social media such as Twitter in recent years, it has become easy to add various information such as location data using mobile devices. Using those data, one can observe the real world without using physical sensors. Therefore, social media have high operational value as social sensors. As described herein, we aim to support decision-making for people who intend to visit a specific place at which an event or some trouble recently occurred. After proposing a method of real-time extraction of data reflecting a burst state showing people's concentration, their inactivity, and continuous flow and dispersion, we confirm the method's effectiveness.
Tourism information collection using the web has become popular in recent years. Moreover, tourists are increasingly using the web to obtain tourist information. Particularly because of the spread of social network services (SNSs), various tourism information is available. Various studies are being conducted using Twitter, which is one of SNS. A low-cost moving average method using geotagged tweets posted location information has been proposed to estimate the best time (peak period) for phenological observation. Geotagged tweets are also useful for estimating and acquiring local tourist information in real time, as a social sensor, because the information can reflect real-world situations. We have been working on, we are pursuing an estimation of the best time to view cherry blossoms. Our earlier studies have improved methods of estimating cherry blossom viewing times. The research so far can estimate the spot that the user knows. However, we cannot estimate the cherry blossoms that the users do not know. Therefore, a user requires a system that is independent of the amount of knowledge. It is possible to provide useful information to all users. We propose a prototype system that estimates the best time without prior knowledge of tourist destinations. In the early stages, the purpose is to use tweets to find spots already featured in magazines and the web. As described herein, we detected spots automatically using a geotagged tweet by visualization with a heat map and setting conditions. The proposed method achieved it in about 80%.
There is an increasingly pressing need, by several applications in diverse domains, for developing techniques able to analyze very large collections of static and streaming sequences (a.k.a. data series), predominantly in real-time. Examples of such applications come from Internet of Things installations, neuroscience, astrophysics, and a multitude of other scientific and application domains that need to apply machine learning techniques for knowledge extraction. It is not unusual for these applications, for which similarity search is a core operation, to involve numbers of data series in the order of hundreds of millions to billions, which are seldom analyzed in their full detail due to their sheer size. Such application requirements have driven the development of novel similarity search methods that can facilitate scalable analytics in this context. At the same time, a host of other methods have been developed for similarity search of high-dimensional vectors in general. All these methods are now becoming increasingly important, because of the growing popularity and size of sequence collections, as well as the growing use of high-dimensional vector representations of a large variety of objects (such as text, multimedia, images, audio and video recordings, graphs, database tables, and others) thanks to deep network embeddings. In this work, we review recent efforts in designing techniques for indexing and analyzing massive collections of data series, and argue that they are the methods of choice even for general high-dimensional vectors. Finally, we discuss the challenges and open research problems in this area.
Nudging is about influencing people to make decisions that are beneficial to society and individuals. We are in particular concerned with using nudges to cause a behavioral change for persons, where healthier or environmentally friendlier behavior may be the goal. As people make more and more decisions in a digital context, digital nudging has steadily become more relevant. With today's technology, it is feasible to dynamically generate highly personalized nudges, using information on the person receiving the nudge, such as their intention and the situation they are in. This paper presents a new nudge model, designed with personalization in mind. We propose to use personal and situational data to generate the most suitable nudge designed from nudge components. The presented nudges should be transparent, helpful and effective to the user.
The construction of ontologies from texts in Spanish is a challenge since this language lacks conceptual databases to validate abstract ontology structures as concepts and relations between them. The preceding generates the necessity of using manual evaluation by human experts; carrying high expenses that limit the calibration of algorithm parameters and large-scale evaluations. This document presents a proposal to evaluate abstract ontology structures through the task of semantic clustering of documents, without the expensive necessity of using manual evaluation or conceptual databases. The proposal is not only affordable but also applicable to model data and domains that lack structured knowledge resources. The experiments lead to the extraction and validation of the ontology structures from texts in Spanish regarding the domain of the Colombian armed conflict.
This article discusses a multi-objective business process optimization. The authors present an approach for an evolutionary combinatorial multi-objective optimization of business process designs with a specified genetic algorithm based on multiple populations. The results show that the optimization approach is capable of producing a satisfactory number of optimized designs alternatives.
When searching the internet e.g. for a person, solution to a problem, or some topic of interest, the wanted outcome is usually specific answers. The result quality for this kind of search is reasonably precise, most of the time we get the answers we need. However, searching a second or third time with the same query, the outcome seems to be minor variations on the same results. So what if the search for information is of a different nature, more like exploring. A typical case would be when a person has a hobby, and time after time wants to search for information about it. Very soon all the quickly accessed information has already been seen, and is not that interesting in the context of new information. This paper presents an approach to Incremental Information Retrieval, where each repeated search with a given query, will provide the user with previously obscured (i.e. unseen) results. We have implemented a prototype system, called IIR, where we demonstrate and test our approach. The system targets situations where users have a continuous information need, that cannot be satisfied through a single search on the Internet, but where the user may want to see new results on the same subject over a period of days, months, or even years. A detailed description of the IIR system and results of our tests are presented.
Analyzing music notations is found useful for musicology purposes. This can be applied by retrieving semantic information from digitally annotated music scores. In this paper, we propose an ontology that structures the knowledge extraction process of a music pattern analysis algorithm. In addition to mandatory elements that describe music scores, the proposed ontology relies on contextual elements and attributes for pattern analysis. The ontology then supports the semantic information retrieval and analysis processes of music score contents. We illustrate the whole mechanism by explaining the workflow of the ontology integrated inside a music encoding platform for eastern music.
Illumination variance is one of the largest real-world problems when deploying face recognition systems. Over the last few years much work has gone into the development of novel 3D face recognition methods to overcome this issue. Photometric stereo is a well-established 3D reconstruction technique capable of recovering the normals and albedo of a surface. Although it provides a way to obtain 3D data, the amount of training data available captured using photometric stereo often does not provide sufficient modelling capacity for training state-of-the-art feature extractors, such as deep convolutional neural networks, from scratch. In this work we present a novel approach to utilising the lighting apparatus commonly used for photometric stereo to synthesise data that can act as a biometric. Combining this with deep learning techniques not only did we achieve near state-of-the-art results, but it gave insight into the possibility of using photometric stereo without the need of reconstruction. This could not only simplify the face recognition process but avoid unnecessary error that may arise from reconstruction. Additionally, we utilise the active lighting from photometric stereo to evaluate the effect of illumination on face recognition. We compare our method to the state-of-the-art 3D methods and discuss potential use cases for our system.
Recently, tourism has become a development emphasis for many countries because international tourism can bring huge revenues; it can also positively affect increased long-run economic growth. However, in this era of complex information, it is hard to get integrated tourist information on the Internet. Consequently, tourists might spend a lot of time to search and compare different information and then decided their travel itinerary. To deal with this issue, we propose a formula for ranking tourist attractions by analyzing geo-tagged photographs on Flickr in this paper. In this way, tourists can save their time to find their interest tourist attractions readily. Moreover, our proposed method includes different aspects such as image quality assessment (IQA), the sentiment of comment, and the popularity of tourist attraction which can evaluate the attractive level of tourist attraction. Especially, we provide different ranking results for local residents and foreign visitors.
Mouse activity is known as an important indicator of user attention and interest on a web page. Many modern commercial web analytics services record and report mouse activity of users on websites. The position of the mouse cursor on the screen is the main source of information, as studies show a correlation between the cursor position during mouse activity and the user's eye gaze. This study focuses on mouse movement directions and speeds, and what they indicate, rather than on the mouse cursor position. Statistical analysis of mouse movements on a technical-educational website, which was selected for this study, sheds light on several interesting patterns. For example, most mouse movements in the examined usage data are either approximately horizontal or approximately vertical, horizontal mouse movements are more frequent than vertical mouse movements, and horizontal movements to the left and to the right are not equivalent in terms of moving time and speed. As this study shows, these statistical findings are related to the reading patterns and behaviors of web users. Associating mouse movements with text reading may potentially highlight content that most users tend to skip, and therefore, might not interest the website audience, and content that many readers read more than once or slowly, meaning it is possibly unclear. This could be useful in locating issues in textual content, in websites in general, and especially in online learning and educational technology applications.
The paper investigates the applications of cooperative Multi-Agent Reinforcement Learning (MARL) schemes to Cognitive Radio Networking (CRN), which in turn can facilitate spectrum utilization for wireless (ad hoc) networks within the Internet of Things (IoT). These schemes provide the ability of wireless transceivers to learn the optimal control and configuration in unknown environmental and application conditions, exploiting potential for cooperation among spectrum secondary users. An overview of the existing MARL approaches to the CRN is provided, with an analysis of their advantages and weaknesses compared to the rest of CRN approaches. We argue that in typical CRN practical scenarios including IoT systems, it is of essential importance that the cooperative algorithms are completely decentralized and distributed, having also a capability that the agents/nodes together can successfully calculate the optimal strategy even if the individual agents cannot. Hence, we propose a new scheme for cooperative spectrum sensing and selection within CRN, based on an adaptation of a recently proposed cooperative MARL scheme, provide detailed analysis of its properties and potential performance, indicating its superiority compared to the existing schemes.