This paper proposes an Opinion Mining model, parameterized according to the reviewer profile. The work aims to highlightand resolve some issues resulting from previous activities in evaluating the goodness of the results obtained by the analysisof
Specific linguistic resources, syntactically annotated and distinctive for each language, related to the affective sphere are important in discovering terms or phrases associated with emotions in order to detect expressed emotions. The paper proposes the initial version of a linguistic resource for the Italian language, mapped on WordNet, where each concept, whose meaning falls into the sphere of emotions, is enriched by a category, allowing to better specify the type of emotion expressed by the term, and by a polarity value, whether the emotion is positive or negative. The resource is based on the model of emotions proposed by Robert Plutchik and has been developed, within a national project of Work-School Alternation, in collaboration with some high school students. The work has a twofold value. On one hand, the development of a linguistic resource, on the other the educational and didactic aspect of students’ involvement. Working on the analysis of literary texts with the task of elaborating and defining the emotions described, the students, assisted by their teachers and two researchers, had to face with their feelings and talk more freely about their affective states, recognizing the emotions and giving them a name.
This paper describes the process of building SardaNet, a linguistic resource for Sardinian language including the different linguistic varieties in Sardinia. SardaNet aims at identifying the semantic relations between Sardinian terms, by manually mapping existing WordNet entries to Sardinian word senses. The work, still in progress, has been developed in collaboration with the University of Cagliari. After discussing some linguistic peculiarities, the paper presents the basic steps of the construction process, the method and the tools involved, the issues encountered during the development and the current version of SardaNet.
Many approaches to Opinion Mining are based on linguistic resources, lexicons or lists of words. The lack of suitable and/or available resources is one of the main problems in the process of opinion extraction and in general in the analysis of textual resources based on a linguistic approach. In this paper we describe FreeWordNet, a linguistic resource based on WordNet and useful in the automatic method we propose for the extraction of features in a general domain. In FreeWordNet each synset is enriched with a set of properties related to adjectives and adverbs and has a positive, negative or objective value associated. The properties associated to each synset support a better identification of the sentiment expressed in relation to the domain and give more details about the relevant terms or the expressions having an opinion associated.
This paper proposes an evaluation method for the performance measurement of an Opinion Mining system, parameterized according to the reviewer's point of view. The work aims to highlight and resolve some issues resulting from previous activities in evaluating the goodness of the results obtained by the analysis of the reviews. The evaluation method is based on a model of Opinion Mining system able to identify and assess the aspects included in a collection of reviews and the weighted importance of such aspects for their authors. A user profiling system will work together with the Opinion Mining system, providing the set of parameters to associate with the aspects and allowing the Opinion Mining system to configure itself according to the user preferences. For the preliminary experiments, a narrower sub-set of Yelp dataset limited to restaurants has been used.
This paper extends a previous work done by the same authors [1] having the aim of improving the predictions coming from a matrix factorization based on latent factor models through an ensemble with the predictions obtained by an Opinion Mining methodology based on a linguistic approach. The experimental analysis was carried out on the Yelp business dataset, limited to the Restaurant category. An hypothesis of influence of the restaurant average rating on the number of stars given by the users is tested. An analysis of the meaning of some of the latent factors is shown.
In this paper, we describe our approach and its results for the MediaEval 2015 Retrieving Diverse Social Images task. The main strength of the proposed approach is its exibility that permits to lter out irrelevant images, and to obtain a reliable set of diverse and relevant images. This is done by rst clustering similar images according to their textual descriptions and their visual content, and then extracting images from dierent clusters according to a measure of user’s credibility. Experimental results shown that it is stable and has little uctuation in both single-concept and multi-concept queries.
An experimental analysis of a combination of Opinion Mining and Collaborative Filtering algorithms is presented. The analysis used the Yelp dataset in order to have both the textual reviews and the star ratings provided by the users. The Opinion Mining algorithm was used to work on the textual reviews, while the Collaborative Filtering worked on the star ratings. The research activity carried out shows that most of the Yelp users provided star ratings corresponding to the related textual review, but in many cases an inconsistence was evident. A set of thresholds and coefficients were applied in order to test a hypothesis about the influence of restaurant popularity on the user ratings. Interesting results have been obtained in terms of Root Mean Squared Error (RMSE).
An integration of an Opinion Mining approach with a Collaborative Filtering algorithm has been applied to the Yelp dataset to improve the predictions through the information provided by the user-generated textual reviews. The research, still in progress, based the Opinion Mining approach on the syntactic analysis of textual reviews and on a beginning polarity evaluation of the sentences. The predictions produced in this way was blended with the predictions coming from a Biased Matrix Factorization algorithm obtaining interesting results in terms of Root Mean Squared Error (RMSE), with potential enhancements. We intend to improve these results in a further phase of activity by including in the Opinion Mining approach the semantic disambiguation and by using better criteria of evaluation of the reviews taking into account a set of 12 business aspects. The Opinion Mining approach will be evaluated comparing the output in terms of predictions with the values manually assigned by a small group of people to a sample of the same reviews.
Online users are talking across social media sites, on public forums and within customer feedback channels about products, services and their experiences, as well as their likes and dislikes. The continuous monitoring of reviews is ever more important in order to identify leading topics and content categories and to understand how those topics and categories are relevant to customers according to their habits. In this context, the chapter proposes an Opinion Mining model to analyze and summarize reviews related to generic categories of products and services. The process, based on a linguistic approach to the analysis of the opinions expressed, includes the extraction of features terms from the reviews in generic domains. It is also capable to determine the positive or negative valence of the identified features exploiting Free-WordNet, a WordNet-based linguistic resource of adjectives and adverbs involved in the whole process.
The pervasive diffusion of social networks as common way to communicate and share information is becoming a valuable resource for analysts and decision makers. Reviews are used every day by common people or by companies who need to make decisions. It is evident that even the opinion monitoring is essential for listening to and taking advantage of the conversations of possible customers in a decision making process. Opinion Mining is a way to analyse opinions related to specific topics: products, services, tourist locations, etc. In this paper we propose an automatic approach to the extraction of feature terms, applying our experience in the semantic analysis of textual resources to Opinion Mining task and performing a contextualisation by means of semantic categorisation, and by a set of qualitics associated to the sense expressed by adjectives and adverbs.
Reviews are used every day by common people or by companies who need to make decisions. Such amount of social data can be used to analyze the present and to predict the near future needs or the probable changes. Mining the opinions and the comments is a way to extract knowledge by previous experiences and by the feedback received. In this chapter we propose an automatic linguistic approach to Opinion Mining by means of a semantic analysis of textual resources and based on FreeWordNet, a new developed linguistic resource. FreeWordNet has been defined by the enrichment of the meanings expressed by adjectives and adverbs in WordNet with a set of properties and the polarity orientation. These properties are involved in the steps of distinction and identification of subjective, objective or factual sentences with polarity valence and contribute in a basic way to the task of features contextualization.
Understanding the meaning of a text depends on the knowledge the reader has about the topic addressed in a document, starting from the most complex concept to the simplest one. The representation of the knowledge is generally performed by ontologies, semantic networks, or typified by statistical algorithms able to organize the contents according to rules based on frequency of terms or synsets. The Opinion Mining is a way to go beyond text categorization through the analysis of the opinions related to a specific topic: a product, a service, a tourist location, etc. In this paper we propose to apply our experience in the semantic analysis of textual resources to the Opinion Mining task, with the aim to propose a different approach to the extraction of feature terms, performing a contextualisation by means of semantic categorisation, a semantic net of concept and by a set of qualities associated to the sense expressed by adjectives and adverbs.
In this chapter we illustrate our vision about the evolution of search engines, dealing with some emerging questions related to the social role of the user on the Web and to the actual approach to access the information. In this scenario, is ever more evident the need to redefine the information paradigm bringing the information to the user and not more the user to the information, with search engines able to provide results without direct questions from users, anticipating their needs. A Web in service of the user, automatically informed by the system with suggested resources related with his life style and his common behavior without the need to ask for them. This approach will be applied to a project named A Semantic Search Engine for a Business Network where the development of a business network creates a point of contact between the academic and the research world and the productive one by the introduction of Natural Language Processing, user profiling, automatic information classification according to users’ personal schemas, contributing in such a way to redefine the vision of information and delineating processes of Human-Machine Interaction.
The Web's evolution during the last few years shows that the advantages from the users' point of view are not so macroscopic. Despite information is still the primal element, is ever more evident the need to redefine the information paradigm so that the net and the information can become really user-centric by an inverse process that brings the information to the user and not more the user to information. Define new tools is needed to create a privileged window of observation on information and knowledge: each user with his specific interest. Not more a single available space of information but shared data for everyone. What each user needs is a specific private space of information according to his point of view, his way to classify and manage the information, related to his network of contacts in the way each person choose to live the Web, the net and the knowledge. In this paper we illustrate a part of a project named A Semantic Search Engine for a Business Network where the introduction of Natural Languages, user profiling, automatic information classification according to users' personal schemas will contribute to redefine the vision of information and delineate processes of Human-Machine Interaction.
The work illustrated in this paper is part of the DART search engine. Its main goal is to index and retrieve information both in a generic and in a specific context where documents can be mapped or not on ontologics, vocabularies and thesauri. To achieve this goal. a semantic analysis process on structured and unstructured parts of documents is performed. The unstructured parts need a linguistic analysis and a semantic interpretation performed by means of Natural Language Processing (NLP) techniques. while the structured parts need a specific parser. Semantic keys are extracted from documents starting from the semantic net of WordNet and enriching it of new nodes, links and attributes.
ABSTRACT Available document collections are more and more required for supervised text categorization tasks. They typically are collections of documents classified by domain engineers. In this paper, we propose a semantic text categorization approach able to automatically create document collections in which documents are classified according to WordNet Domains taxonomy. Experiments have been performed by training a classifier with an automatic document collection and comparing results with those obtained by training the same classifier on a hand-made document collection. Experimental results point out that, on average, the performances of the automatic approach are quite similar to those obtained on a document collection classified by domain engineers. KEYWORDS Text Categorization, Document Collections, Intelligent Software Systems, Machine Learning. 1. INTRODUCTION Text categorization can be defined as the task of determining and assigning topical labels to content. The more the amount of available data (e.g., in digital libraries), the greater the need for high-performance text categorization algorithms. In particular, text categorization is a key technology in several information processing tasks, including controlled vocabulary indexing, routing and packaging of news and other text streams, content filtering, information security, help desk automation, and others. In the literature, many machine learning approaches have been proposed, both in the field of supervised [Sebastiani02] and unsupervised [Ghahramani04] learning. In particular, supervised approaches use only labeled data during the training phase. On the contrary, unsupervised approaches use unlabeled data, which may be easy to collect, but difficult to use. Semi-supervised learning in part resolves this problem by using large amount of unlabeled data, together with labeled data, to build better classifiers [Zhu05]. In this scenario, available text categorization document collections are more and more required. They are typically standard collections to which humans have assigned categories from a predefined set ([Lewis96], [Yang99], and [Lewis04]), so that researchers are able to test their algorithms in a controlled benchmarking context. Unfortunately, existing document collections suffer from one or more of the following drawbacks: (i) few documents, (ii) lack of the document full text, (iii) inconsistent or incomplete category assignment, (iv) peculiar textual properties, and (v) limited availability. Moreover, often researchers do not have documentation on how collections were produced, and on the nature of the underlying categories. To our knowledge, so far, only few attempts to automatically create document collections have been proposed [Ko00]. In particular, semantic approaches to text categorization have not been applied to this specific task. In this paper, we illustrate a method to create document collections by adopting a fully-automated semantic approach. Each text document is suitably labeled according to a predefined taxonomy of classes, namely WordNet Domains [Magnini00]. Experimental results point out that the proposed method allows to create reliable document collections.
The automatic creation of a conceptual knowledge map using documents coming from the Web is a very relevant problem because of the difficulty to distinguish between valid and invalid contents documents. In this paper we present an improved search engine GUI for displaying and organizing alternative views of data, by the use of a 3D graphical interface, and a method for organizing search results using a semantic approach during the storage and retrieval of information. The presented work deals with two main aspects. The first one regards the semantic aspects of knowledge management, in order to support the user during the query composition and to supply to him information strictly related to his interests. The second one argues the advantages coming from the adoption of a 3D user interface, to provide alternative views of data.