
The number of documents available into Internet moves each day up. For this reason, processing this amount of information effectively and expressibly becomes a major concern for companies and scientists. Methods that represent a textual document by a topic representation are widely used in Information Retrieval (IR) to process big data such as Wikipedia articles. One of the main difficulty in using topic model on huge data collection is related to the material resources (CPU time and memory) required for model estimate. To deal with this issue, we propose to build topic spaces from summarized documents. In this paper, we present a study of topic space representation in the context of big data. The topic space representation behavior is analyzed on different languages. Experiments show that topic spaces estimated from text summaries are as relevant as those estimated from the complete documents. The real advantage of such an approach is the processing time gain: we showed that the processing time can be drastically reduced using summarized documents (more than 60\% in general). This study finally points out the differences between thematic representations of documents depending on the targeted languages such as English or latin languages.
With the increasing number of annotated corpora for supervised Named Entity Recognition, it becomes interesting to study the combination and augmentation of these corpora for the same annotation task. In this paper, we particularly study the combination of heterogeneous corpora for Medical Entity Recognition by using a meta-learning classifier that combines the results of individual Conditional Random Fields (CRFs) models trained on different corpora. We propose selective data augmentation approaches and compare them with several meta-learning algorithms and baselines. We evaluate our approach using four sub-classifiers trained on four heterogeneous corpora. We show that despite the high disagreements between the individual models on the four test corpora, our selective data augmentation approach improves performance on all test corpora and outperforms the combination of all training corpora.
Sentiment analysis is the process of identifying the subjective information in the source materials towards an entity. It is a subfield of text and web mining. Web is a rich and progressively expanding source of information. Sentiment analysis can be modelled as a text classification problem. Text classification suffers from the high dimensional feature space and feature sparsity problems. The use of conventional representation schemes to represent text documents can be extremely costly especially for the large text collections. In this regard, data reduction techniques are viable tools in representing document collections. Latent Dirichlet allocation (LDA) is a popular generative probabilistic model to represent collections of discrete data. In this regard, this paper examines the performance of LDA in text sentiment classification. In the empirical analysis, five classification algorithms (Naïve Bayes, support vector machines, logistic regression, radial basis function network and K-nearest neighbor algorithms) and five ensemble methods (Bagging, AdaBoost, Random Subspace, voting and stacking) are evaluated on four sentiment datasets.
In this paper, we present an innovative method for multi-label text classification. Our method uses Lucene to index texts and then assigns one or more classes to a new text based on its similarity relative to an annotated corpus. For finer granularity, we split the text into phrases, and then we focus on the noun phrases. Instead of classifying the entire text, we classify each noun phrase. The result of classifying the text is then assembled as the set of classes allocated to its noun phrases.
Unlike the written texts, discourse segmentation of the Arab oral dialogues is a challenging task that is held back in most cases by the spontaneous character of oral speech. Like any segmentation task, segmentation in minimum discursive units (UDM) aims to cut the different statements of a speech into simple proposals easily usable in subsequent treatment. The majority of the work on the Arabic language was based on extensive syntactic analysis approaches. In this article, we try to show the effectiveness of hybrid approaches combining linguistic and probabilistic processes over purely linguistic approaches. The performance of our segmentation was evaluated on a relatively large size corpus. We built this corpus by using the method of the wizard of Oz.
A crucial component of text-to-speech systems is the one responsible for the transcription of the written text to its phonemic representation. although the complexity of the relation between the written and spoken form of languages varies, most languages have their regular and irregular phonological set of rules. In this paper, we present a system for the phonemic transcription of Hungarian. Beside the implementation of rules describing default letter-to-phoneme correspondences and morphophonological alternations, the tool incorporates the knowledge of a Hungarian morphological analyzer in order to be able to detect compound and other morpheme boundaries, and it contains a rich lexicon of entries with irregular pronunciation. It is shown that the system performs well even on texts containing a high number of foreign names.
The amount of user generated contents from various social medias allows analyst to handle a wide view of conversations on several topics related to their business. Nevertheless keeping up-to-date with this amount of information is not humanly feasible. Automatic Summarization then provides an interesting mean to digest the dynamics and the mass volume of contents. In this paper, we address the issue of tweets summarization which remains scarcely explored. We propose to automatically generated summaries of Micro-Blogs conversations dealing with public figures E-Reputation. These summaries are generated using key-word queries or sample tweet and offer a focused view of the whole Micro-Blog network. Since state-of-the-art is lacking on this point we conduct and evaluate our experiments over the multilingual CLEF RepLab Topic-Detection dataset according to an experimental evaluation process.
The indexed Web increases every day, making the development of automatic methods for knowledge extraction more relevant. The area of Sentiment Analysis or Opinion Mining aims to extract opinions from the user-generated content and to define the semantic orientation of each individual opinion. This work proposes an approach to estimate the degree of importance of comments generated by web users by using a Fuzzy system. The system has three inputs: author reputation, number of tuples 〈feature , quality word〉, and percentage of correctly spelled words and one output: importance degree of the comment. The importance degree was used to select the best comments in a Corpus. The paper also describes two experiments: the first was used to fit the system and was conducted with 350 reviews about smartphones (168 positives and 182 negatives). It achieved 63.17% in f-measure in the top 50 positive reviews, and 43.75% in f-measure in top 50 negative reviews. The second was used to compare the results of a sentiment orientation method before and after the selection of the best comments. It was conducted with 1620 reviews also about smartphones (982 positives and 594 negatives) and our approach improved the results of sentiment orientation method up to approximately 10% in f-measure in positive reviews and 7% in f-measure in negative reviews.
In this study, we propose a method to automatically correct redundant sentences using patterns and machine learning and propose methods that combine pattern-based and machine-learning methods. We conducted experiments to correct redundant sentences containing “kanou” (possible or possibility), “toiu” (“that is” or called), and “surukoto” (to do). The results demonstrate that the proposed method can correct redundant sentences at an accuracy of 0.6 and estimate corrected expressions against redundant parts at an accuracy of 0.7; furthermore, we created a method to support a user’s writing. In this method, a system displays redundant parts and provides candidate expressions.
Update Summarization aims to produce summaries under the assumption that the reader had some knowledge about the topic from the source texts. Usually, traditional approaches of summarization use sentential ranking functions in order to find the most relevant and updated sentences from source-texts. We pro-pose the enriching of these methods with the using of subtopic representation, which are coherent textual segments with one or more sentences in a row. The results of our experiments show that our text representation improves the quality of produced summary and show high recall values.
Several methods for ontology development have been proposed. However, the development of domain ontologies is still carried out in an ad-hoc manner. This paper explores the use of a microgenetic algorithm with a seeding scheme based on hierarchical clustering for ontology class hierarchy construction. The microgenetic algorithm (μGA) is composed of an inner loop and an outer loop. The inner loop consists of: the evaluation of the fitness of each member of the population; the selection of parent chromosomes; the generation of a new population by using crossover and mutation operations; and the separation of the best-fit individual after convergence. The outer loop consists of creating a new random population, transferring the best individual from the inner loop, and restarting the inner loop. The fitness function is based on the correlation between the pair-wise similarities based on the semantic similarity measure of Wu-Palmer and those obtained using Internet and the normalized Google distance (NGD). The proposed approach was tested on the construction of a class hierarchy of machining processes. The results indicate that accurate class hierarchies can be obtained and convergence can be achieved fast with little memory to store the population.
Discourse connectives (e.g. however, because) are terms that can explicitly convey a discourse relation within a text. While discourse connectives have been shown to be an effective clue to automatically identify discourse relations, they are not always used to convey such relations, thus they should first be disambiguated between discourse-usage non-discourse-usage. In this paper, we investigate the applicability of features proposed for the disambiguation of English discourse connectives for French. Our results with the French Discourse Treebank (FDTB) show that syntactic and lexical features developed for English texts are as effective for French and allow the disambiguation of French discourse connectives with an accuracy of 94.2%.
In this paper, we present hybrid approaches for pronominal reference type (abstract or concrete) identification and event anaphora resolution for Hindi. Pronominal reference type identification is one of the important parts for any anaphora resolution system as it helps anaphora resolver in optimal feature selection based on pronominal reference types. We use language specific rules and features in set of classifiers (ensemble learning) for pronominal type identification. We discuss event referring anaphors (pro-nouns) and their resolution using Paninian dependency grammar, language syntax, proximity of events, etc. We achieved around 9̃0% accuracy in the pronominal reference type identification and around 7̃1% F-score in the event anaphora resolution on Hindi dependency tree-bank corpus.
In this paper, we present a survey and comparative studies on semantic textual similarity methods, those are based on WordNet taxonomy. We also proposed a new method for measuring semantic similarity between sentences. This proposed method, uses the advantages of taxonomy methods and merge these information to a language model. It considers the WordNet synsets for lexical relationships between nodes/words and uni-gram language model is implemented over a large corpus to assign the information content value between the two nodes of different classes. Finally, a similarity score is generated by considering the maximum weight and shortest distance of the graph. To evaluate and compare the method, SemEval 2015 English STS task 2 training dataset is considered.
. There are several methods and available tools for terminology extraction, but the quality of the extracted terms is not always high. Hence, an important consideration in terminology extraction is to assess the quality of the extracted terms. In this paper, we propose and make available a tool for annotating the correctness of terms extracted by three term-extraction tools. This tool facilitates term annotation by using a domain-specific dictionary, a set of filters, and an annotation memory, and allows for post-hoc evaluation. We present a study in which two human judges used the developed tool for term annotation. Their annotations were then analyzed to determine the efficiency of term extraction tools by measures of precision, recall, and F-score, and to calculate the inter-annotator agreement rate.
. In this paper, we explore the use of a statistical machine translation system for optical character recognition (OCR) error correction. We investigate the use of word and character-level models to support a translation from OCR system output to correct french text. Our experiments show that character and word based machine translation correction make significant improvements to the quality of the text produced through digitization. We test the approach on historical data provided by the National Library of France. It shows a relative Word Error Rate reduction of 60% at the word-level, and 54% at the character level.
Building a parallel Treebank anticipates alignment of linguistic information represented by diverse structures on different layers of a bilingual text. In this paper, we describe our observations for inference translation equivalents in parallel texts of languages with diverse structures - German and Georgian. They belong to the different language families and as a consequence enjoy different typological features manifested by diverse morphological structures, word and phrase order in a clause. In the bilingual German-Georgian Treebank development process it has been given a try to cluster the tolerant syntactic structures and classify phrase conventional translations that could be considered as equivalent units in the bilingual text alignment issue.
Corpora and web texts can become a rich language learning resource if we have a means of assessing whether they are linguistically appropriate for learners at a given proficiency level. In this paper, we aim at addressing this issue by presenting the first approach for predicting linguistic complexity for Swedish second language learning material on a 5-point scale. After showing that the traditional Swedish readability measure, L\"asbarhetsindex (LIX), is not suitable for this task, we propose a supervised machine learning model, based on a range of linguistic features, that can reliably classify texts according to their difficulty level. Our model obtained an accuracy of 81.3% and an F-score of 0.8, which is comparable to the state of the art in English and is considerably higher than previously reported results for other languages. We further studied the utility of our features with single sentences instead of full texts since sentences are a common linguistic unit in language learning exercises. We trained a separate model on sentence-level data with five classes, which yielded 63.4% accuracy. Although this is lower than the document level performance, we achieved an adjacent accuracy of 92%. Furthermore, we found that using a combination of different features, compared to using lexical features alone, resulted in 7% improvement in classification accuracy at the sentence level, whereas at the document level, lexical features were more dominant. Our models are intended for use in a freely accessible web-based language learning platform for the automatic generation of exercises.