Abstract In the field of natural language understanding (NLU), a fundamental element is the representation of word meanings, a process that is integral to a wide range of applications. These applications span numerous domains including machine translation, question answering systems, text summarization, information retrieval, and supporting the functionality of virtual assistants. The increasing demand for reasoning in environments that are multilingual and knowledge transfer in cross-lingual systems has led to the development of cross-lingual semantic spaces. These spaces provide a representation of words from different languages in a shared space. With the increased emphasis on cross-lingual representations, several methods have been developed. Related works usually differ in training strategies and evaluate only limited aspects of semantic spaces. This lack of meaningful comparison made us write this study as we think it is crucial for the following research. As a basis for our comparison, we project semantic spaces into a shared space using both linear transformations supervised by bilingual dictionaries and transformations with no supervision at all. To allow comparison from different points of view, our evaluation includes both intrinsic tasks, such as cross-lingual word similarities, cross-lingual word analogies, and word machine translation, and extrinsic tasks like sentiment analysis and topic classification. Additionally, we also explored hubness to investigate the internal relationships within the semantic space. Our experiments include six languages from three different language families: English, German, Italian, Spanish, Croatian and Czech. Finally, we show that different preprocessing steps can have a significant impact on the performance of cross-lingual semantic spaces.
The rapid development in the field of natural language processing (NLP) and the increasing complexity of linguistic tasks demand the use of efficient and effective methods. Cross-lingual linear transformations between semantic spaces play a crucial role in this domain. However, compared to more advanced models such as transformers, linear transformations often fall short, especially in terms of accuracy. It is thus necessary to employ innovative approaches that not only enhance performance but also maintain low computational complexity. In this study, we propose Kernel Least Squares (KLS) for linear transformation between semantic spaces. In our comprehensive analysis involving three intrinsic and two extrinsic experiments across six languages from three different language families and a comparative evaluation with nine different linear transformation methods, we demonstrate the superior performance of KLS. Our results show that the proposed method significantly improves word translation accuracy, thereby standing out as the most efficient method for transforming only the source semantic space.
Cross-lingual semantic textual similarity systems estimate the degree of the meaning similarity between two sentences, each in a different language. State-of-the-art algorithms usually employ machine translation and combine vast amount of features, making the approach strongly supervised, resource rich, and difficult to use for poorly-resourced languages. In this paper, we study linear transformations, which project monolingual semantic spaces into a shared space using bilingual dictionaries. We propose a novel transformation, which builds on the best ideas from prior works. We experiment with unsupervised techniques for sentence similarity based only on semantic spaces and we show they can be significantly improved by the word weighting. Our transformation outperforms other methods and together with word weighting leads to very promising results on several datasets in different languages.
We generalize the word analogy task across languages, to provide a new intrinsic evaluation method for cross-lingual semantic spaces. We experiment with six languages within different language families, including English, German, Spanish, Italian, Czech, and Croatian. State-of-the-art monolingual semantic spaces are transformed into a shared space using dictionaries of word translations. We compare several linear transformations and rank them for experiments with monolingual (no transformation), bilingual (one semantic space is transformed to another), and multilingual (all semantic spaces are transformed onto English space) versions of semantic spaces. We show that tested linear transformations preserve relationships between words (word analogies) and lead to impressive results. We achieve average accuracy of 51.1%, 43.1%, and 38.2% for monolingual, bilingual, and multilingual semantic spaces, respectively.
In this paper we evaluate our new approach based on the Continuous Bag-of-Words and Skip-gram models enriched with global context information on highly inflected Czech language and compare it with English results. As a source of information we use Wikipedia, where articles are organized in a hierarchy of categories. These categories provide useful topical information about each article. Both models are evaluated on standard word similarity and word analogy datasets. Proposed models outperform other word representation methods when similar size of training data is used. Model provide similar performance especially with methods trained on much larger datasets.
In this paper we extend Skip-Gram and Continuous Bag-of-Words Distributional word representations models via global context information. We use a corpus extracted from Wikipedia, where articles are organized in a hierarchy of categories. These categories provide useful topical information about each article. We present the four new approaches, how to enrich word meaning representation with such information. We experiment with the English Wikipedia and evaluate our models on standard word similarity and word analogy datasets. Proposed models significantly outperform other word representation methods when similar size training data of similar size is used and provide similar performance compared with methods trained on much larger datasets. Our new approach shows, that increasing the amount of unlabelled data does not necessarily increase the performance of word embeddings as much as introducing the global or sub-word information, especially when training time is taken into the consideration.
We demonstrate several ways to use morphological word analogies to examine the representation of complex words in semantic vector spaces. We present a set of morphological relations, each of which can be used to generate many word analogies. 1. We show that the difference-vectors for pairs which have the same relation to each other are similarly aligned. 2. We suggest that addition of difference-vectors is a useful phrase-building operator. 3. We propose that pairs in the same relation may have similar relative frequencies. 4. We suggest that homographs, which necessarily have the same semantic vectors, can sometimes be separated into different vectors for different senses, using frequency estimates and alignment constraints obtained from word analogies. 5. We observe that some of our analogies seem to be parallel, and might be combined. We use Arabic words as a case study, because Arabic orthography includes verb conjugations, object pronouns, definitive articles, possessive pronouns, and some prepositions in single word-forms. Therefore, a number of short phrases, built up of easily perceived constituents, are already present in stock semantic spaces for Arabic available on the web. Similar phrases in English would require including bigrams or trigrams as lemmas in the word embedding, although English derivational morphology allows for other relationships in standard semantic spaces which Arabic does not, for example negation. We make our corpus of morphological relations available to other researchers.
Semantic textual similarity is the core shared task at the International Workshop on Semantic Evaluation (SemEval). It focuses on sentence meaning comparison. So far, most of the research has been devoted to English. In this paper we present first Czech dataset for semantic textual similarity. The dataset contains 1425 manually annotated pairs. Czech is highly inflected language and is considered challenging for many natural language processing tasks. The dataset is publicly available for the research community. In 2016 we participated at SemEval competition and our UWB system were ranked as second among 113 submitted systems in monolingual subtask and first among 26 systems in cross-lingual subtask. We adapt the UWB system for Czech (originally for English) and experiment with new Czech dataset. Our system achieves very promising results and can serve as a strong baseline for future research.
The word embedding methods have been proven to be very useful in many tasks of NLP (Natural Language Processing). Much has been investigated about word embeddings of English words and phrases, but only little attention has been dedicated to other languages. Our goal in this paper is to explore the behavior of state-of-the-art word embedding methods on Czech, the language that is characterized by very rich morphology. We introduce new corpus for word analogy task that inspects syntactic, morphosyntactic and semantic properties of Czech words and phrases. We experiment with Word2Vec and GloVe algorithms and discuss the results on this corpus. The corpus is available for the research community.
We present our UWB system for the task of capturing discriminative attributes at SemEval 2018. Given two words and an attribute, the system decides, whether this attribute is discriminative between the words or not. Assuming Distributional Hypothesis, i.e., a word meaning is related to the distribution across contexts, we introduce several approaches to compare word contextual information. We experiment with state-of-the-art semantic spaces and with simple co-occurrence statistics. We show the word distribution in the corpus has potential for detecting discriminative attributes. Our system achieves F1 score 72.1% and is ranked #4 among 26 submitted systems.
Vector semantic spaces, in which a multi-dimenstional numeric vector is used to represent the meaning of a word, are making new natural language applications possible. Word analogies have become a standard tool to evaluate semantic spaces. They also teach us something about what kinds of information the vectors in the semantic space can embody. Arabic orthography has morphological constructs which are realized in syntax in some other languages: the presence or absence of the article ??; bi-??, ka-?? prepositional prefixes; verbs with object suffixes can constitute an entire sentence. The structured word-forms offer the opportunity to study how vector representations of meaning interact in the semantic space to form verb phrases, noun phrases, and prepositional phrases. We provide a corpus of Arabic analogies focused on the morphological constructs which can participate in some of these phrases. We conducted an examination of ten different semantic spaces to see which of them is most appropriate for this set of analogies, and we illustrate the use of the corpus to examine phrase-building.
We propose a novel metric for evaluating summary content coverage.The evaluation framework follows the Pyramid approach to measure how many summarization content units, considered important by human annotators, are contained in an automatic summary.Our approach automatizes the evaluation process, which does not need any manual intervention on the evaluated summary side.Our approach compares abstract meaning representations of each content unit mention and each summary sentence.We found that the proposed metric complements well the widely-used ROUGE metrics.
We introduce Flames Detector, an online system for measuring flames, i.e. strong negative feelings or emotions, insults or other verbal offences, in news commentaries across five languages.It is designed to assist journalists, public institutions or discussion moderators to detect news topics which evoke flames.We propose a machine learning approach to flames detection and calculate an aggregated score for a set of comment threads.The demo application shows the most flaming topics of the current period in several language variants.The search functionality gives a possibility to measure flames in any topic specified by a query.The evaluation shows that the flame detection in discussions is a difficult task, however, the application can already reveal interesting information about the actual news discussions.
This paper introduces a new unsupervised approach for dialogue act induction. Given the sequence of dialogue utterances, the task is to assign them the labels representing their function in the dialogue. Utterances are represented as real-valued vectors encoding their meaning. We model the dialogue as Hidden Markov model with emission probabilities estimated by Gaussian mixtures. We use Gibbs sampling for posterior inference. We present the results on the standard Switchboard-DAMSL corpus. Our algorithm achieves promising results compared with strong supervised baselines and outperforms other unsupervised algorithms.
Word embeddings are commonly compared either with human-annotated word similarities or through improvements in natural language processing tasks. We propose a novel principle which compares the information from word embeddings with reality. We implement this principle by comparing the information in the word embeddings with geographical positions of cities. Our evaluation linearly transforms the semantic space to optimally fit the real positions of cities and measures the deviation between the position given by word embeddings and the real position. A set of well-known word embeddings with state-of-the-art results were evaluated. We also introduce a visualization that helps with error analysis.
Restaurant Reviews CZ ABSA - 2.15k reviews with their related target and category The work done is described in the paper: https://doi.org/10.13053/CyS-20-3-2469
This paper describes our system used in the Aspect Based Sentiment Analysis (ABSA) task of SemEval 2016.Our system uses Maximum Entropy classifier for the aspect category detection and for the sentiment polarity task.Conditional Random Fields (CRF) are used for opinion target extraction.We achieve state-of-the-art results in 9 experiments among the constrained systems and in 2 experiments among the unconstrained systems.
We present our UWB system for Semantic Textual Similarity (STS) task at SemEval 2016. Given two sentences, the system estimates the degree of their semantic similarity. We use state-of-the-art algorithms for the meaning representation and combine them with the best performing approaches to STS from previous years. These methods benefit from various sources of information, such as lexical, syntactic, and semantic. In the monolingual task, our system achieve mean Pearson correlation 75.7% compared with human annotators. In the cross-lingual task, our system has correlation 86.3% and is ranked first among 26 systems.
We introduce a system focused on solving SemEval 2016 Task 2 ‐ Interpretable Semantic Textual Similarity. The system explores machine learning and rule-based approaches to the task. We focus on machine learning and experiment with a wide variety of machine learning algorithms as well as with several types of features. The core of our system consists in exploiting distributional semantics to compare similarity of sentence chunks. The system won the competition in 2016 in the “Gold standard chunk scenario”. We have not participated in the “System chunk scenario”.