
Pronunciation assessment remains a subjective task which depends on a pronunciation reference hold as canonical. Whether a second language (L2) speaker is able to replicate said reference is decided by an assessor who perceives the identity of the sounds produced. It is known that the assessor has a bias caused by the perception of the speaker, hence the definition of a standard for L2 pronunciation is crucial in a formal assessment. In Computer Assisted Pronunciation Assessment (CAPA), the definition of a pronunciation standard for L2 is not trivial due to limited L2 data annotated for mispronunciations. Inspired on the assessor’s bias, this work explores an alternative to a conventional Automatic Speech Recognition approach for CAPA by using speaker metadata along with acoustic observations for mispronunciation detection. A combination of Bidirectional Long-Short Memory with self-attention was used to detect pronunciation errors in short speech segments. It was found that the use of categorical metadata can have a positive effect in the classification of mispronounced segments depending on the sparsity and balance of the classes. It was also found that different assessors can be influenced differently by information about the speaker’s linguistic background. The effect of the metadata was tested on data from Dutch children learners of English as L2 in schools across the Netherlands. The limited speaker diversity of the corpus made the task a challenge worth keep exploring.
Algorithmic journalism refers to automatic AI-constructed news stories. There have been successful commercial implementations for news stories in sports, weather, financial reporting and similar domains with highly structured, well defined tabular data sources. Other domains such as local reporting have not seen adoption of algorithmic journalism, and thus no automated reporting systems are available in these categories which can have important implications for the industry. In this paper, we demonstrate a novel approach for producing news stories on government legislative activity, an area that has not widely adopted algorithmic journalism. Our data source is state legislative proceedings, primarily the transcribed speeches and dialogue from floor sessions and committee hearings in US State legislatures. Specifically, we create a library of potential events called phenoms. We systematically analyze the transcripts for the presence of phenoms using a custom partial order planner. Each phenom, if present, contributes some natural language text to the generated article: either stating facts, quoting individuals or summarizing some aspect of the discussion. We evaluate two randomly chosen articles with a user study on Amazon Mechanical Turk with mostly Likert scale questions. Our results indicate a high degree of achievement for accuracy of facts and readability of final content with 13 of 22 users in the first article and 19 of 20 subjects of the second article agreeing or strongly agreeing that the articles included the most important facts of the hearings. Other results strengthen this finding in terms of accuracy, focus and writing quality.
While research of sentiment analysis became very popular on the global scope, in Slovak language as an under-resourced language there are still many issues to be tackled, especially the lack of resources. In this paper, we introduce a sentiment analysis game designed to collect sentiment annotations. The game is intended for a single player who, motivated by game score, chooses the sentiment category of each word of a sentence. We describe the annotation collection process during which over 12 500 annotations of individual words and over 1 000 annotations of entire sentences were obtained within a week. The collected annotations were used to construct a sentiment lexicon. Using artificial bee colony algorithm, optimal lexicon construction parameters were discovered. To evaluate the final lexicon's usefulness, we applied simple sentence-level lexicon-based sentiment analysis methods on a manually annotated dataset of mobile phone reviews. The same was done with other existing lexicons for comparison. The results of our experiments show that collecting annotations using our game can be useful as a method for constructing a sentiment lexicon.
The emergence of smart home assistants increased the need for robust Far-Field Speaker Identification models. Speaker Identification enables the assistants to perform personalized tasks. Smart home assistants face very challenging speech conditions, including various room shapes and sizes, various distances of the speaker from the microphone, various types of distractor noises (TV in the background, air conditioner, fridge, babble speech of other speakers, etc.). This paper describes the use of Invariant Representation Learning (IRL) as a method aimed to increase the robustness of Speaker Identification models on Far-Field. We introduce three new versions of IRL: Text-Dependent IRL (TD-IRL), Text Independent IRL (TI-IRL), and Deep Features IRL (DF-IRL). We evaluate the IRL models performance and compare them to the base x-vector model. The various Far-Field scenarios are evaluated using VOiCES dataset - a dataset of simulated Far-Field recordings in four real furnished rooms. TD-IRL and DF-IRL improve the minDCF results on the far-field scenarios by an average of 36%, and TI-IRL improves it by 31% with respect to the baseline model.
Recent developments in Named Entity Recognition (NER) have demonstrated good results for grammatically correct texts, even in low resourced settings. However, when the NER model faces ungrammatical text, it often shows poor performance. In this study, we analyze NER performance on datasets containing errors typical for user-generated texts in the Latvian language. We explore three different strategies to increase the robustness of the named entity recognition: error injection into grammatically correct texts, augmenting grammatically correct texts with erroneous texts and augmenting grammatically correct texts with erroneous texts that contain specific types of errors. We demonstrate that in low resourced settings, the best noise-robust model could be obtained by augmenting training data with datasets containing different error types. Our best model achieves an average F1 score of 83.5 (84.1 for baseline) on grammatically correct text, while keeping good performance (79 F1 vs. 66 for baseline) on noisy texts.
Image captioning is a complex artificial intelligence task that involves many fundamental questions of data representation, learning, and natural language processing. In addition, most of the work in this domain addresses the English language because of the high availability of annotated training data compared to other languages. Therefore, we investigate methods for image captioning in German that transfer knowledge from English training data. We explore four different methods for generating image captions in German, two baseline methods and two more advanced ones based on transfer learning. The baseline methods are based on a state-of-the-art model which we train using a translated version of the English MS COCO dataset and the smaller German Multi30K dataset, respectively. Both advanced methods are pre-trained using the translated MS COCO dataset and fine-tuned for German on the Multi30K dataset. One of these methods uses an alternative attention mechanism from the literature that showed a good performance in English image captioning. We compare the performance of all methods for the Multi30K test set in German using common automatic evaluation metrics. We show that our advanced method with the alternative attention mechanism presents a new baseline for German BLEU, ROUGE, CIDEr, and SPICE scores, and achieves a relative improvement of 21.2% in BLEU-4 score compared to the current state-of-the-art in German image captioning.
In this paper, we discuss some interesting features of training a special acoustic model for only one speaker with a constant acoustic background (acoustic channel). Currently, the LF-MMI method achieves the best results in many speech recognition tasks. A typical LF-MMI training procedure uses a special 1-state HMM topology that has different pdfs at the self-loop and forward transitions. We would like to discuss the replacement of this typical LF-MMI HMM by different types of HMM topologies (1-, 2- and 3-state HMM topologies that have outputs associated with states). Next, we discuss the advantages of using biphone context modeling over using the triphone context or even simpler context-free monophone. We also address the effect of the amount of training data and the context of DNN on WER, and all this with regard to a special acoustic model with one speaker and an almost constant acoustic channel.
We present an investigation into the use of semi-supervised training and content genre adaptation for improved automatic speech recognition (ASR) of diverse user-generated videos in the task of spoken content retrieval (SCR). Previous work has successfully applied semi-supervised training in single domain ASR tasks. Our focus is on the exploration of the effective use of semi-supervised training of ASR systems for transcription of the spoken content stream of user-generated video data in varied domains and acoustic noise conditions for use in SCR systems. We examine all elements of ASR system development including: data segmentation, data selection, genre labels, acoustic modelling and language modelling using semi-supervised training. We evaluate its effectiveness for ASR and a known-item SCR task using the Blip100000 multimedia collection. Our baseline hybrid ASR system trained out-of-domain produced WERs 31.27% and 44.69% on dev and test sets, respectively. By introducing the techniques outlined above, the WERs are reduced to 26.82% and 39.21% respectively. The improved transcripts increased mean reciprocal rank (MRR) results for the SCR task from 15.59% to 39.38% on dev and 20.98% to 37.23% on test sets.
In this paper, we present our progress in pre-training monolingual Transformers for Czech and contribute to the research community by releasing our models for public. The need for such models emerged from our effort to employ Transformers in our language-specific tasks, but we found the performance of the published multilingual models to be very limited. Since the multilingual models are usually pre-trained from 100+ languages, most of low-resourced languages (including Czech) are under-represented in these models. At the same time, there is a huge amount of monolingual training data available in web archives like Common Crawl. We have pre-trained and publicly released two monolingual Czech Transformers and compared them with relevant public models, trained (at least partially) for Czech. The paper presents the Transformers pre-training procedure as well as a comparison of pre-trained models on text classification task from various domains.
Despite the growing popularity of metric learning approaches, very little work has attempted to perform a fair comparison of these techniques for speaker verification. We try to fill this gap and compare several metric learning loss functions in a systematic manner on the VoxCeleb dataset. The first family of loss functions is derived from the cross entropy loss (usually used for supervised classification) and includes the congenerous cosine loss, the additive angular margin loss, and the center loss. The second family of loss functions focuses on the similarity between training samples and includes the contrastive loss and the triplet loss. We show that the additive angular margin loss function outperforms all other loss functions in the study, while learning more robust representations. Based on a combination of SincNet trainable features and the x-vector architecture, the network used in this paper brings us a step closer to a really-end-to-end speaker verification system, when combined with the additive angular margin loss, while still being competitive with the x-vector baseline. In the spirit of reproducible research, we also release open source Python code for reproducing our results, and share pretrained PyTorch models on torch.hub that can be used either directly or after fine-tuning.
Many languages still lack the annotated training data needed for supervised learning. This issue is often addressed by using auxiliary supervision and the so called transfer learning. In this work we focus on the problem of combining two types of auxiliary supervision – cross-lingual and cross-task. Previous work has shown promising results for this combination. Here, we aim to explore various advanced parameter sharing techniques to improve the results. We propose three distinct techniques with various properties and evaluate their performance on four Indo-European languages and four distinct NLP tasks (dependency parsing, language modeling, named entity recognition and part-of-speech tagging). We conclude that the proposed techniques significantly improve the performance for zero-shot learning.
According to Cognitive Grammar (CG) theory, the overall structure of a natural language is motivated by a relatively small set of domain-independent cognitive abilities. In this paper, we draw insights from CG to propose an approach to natural language parsing with little syntactic annotation. A sentence functions as a cohesive whole because its parts are meaningfully linked. We propose that every part of a sentence can be analysed along three axes: composition, interaction and autonomy. When two expressions semantically correspond in all the three axes we call them cohesive. We present an algorithm that reads parts of sentences incrementally, recognises their construction schemas along the three axes, assembles any two component schemas into one composite schema if they are cohesive, parses a span of text as incrementally successive assembly of components into composites, retains multiple running parses within the span and chooses the best parse. The basic construction schema definitions and their patterns of assembly are implemented as dictionary-cum-rules because they are fewer in number, largely language-independent and can be extended to handle language-specific variations. A basic feedforward neural network component was trained to learn all valid patterns of assemblies possible in a span of text and to choose the best parse. A successful parse exhausts all the words in the sentence and ensures local cohesion and assembly at every stage of analysis. We present our approach, parser implementation and evaluation results in Welsh and English. By adding WordNet synsets we are able to show improvements in parser performance.
This paper presents an empirical study that harnesses the benefits of Positional Language Models (PLMs) as key of an effective methodology for understanding the gist of a discursive text via extractive summarization. We introduce an unsupervised, adaptive, and cost-efficient approach that integrates semantic information in the process. Texts are linguistically analyzed, and then semantic information, specifically synsets and named entities, are integrated into the PLM, enabling the understanding of text, in line with its discursive structure. The proposed unsupervised approach is tested for different summarization tasks within standard benchmarks. The results obtained are very competitive with respect to the state of the art, thus proving the effectiveness of this approach, which requires neither training data nor high-performance computing resources.
S-capade (spelling correction aimed at particularly deviant errors) is a phonemic distance based spellchecking tool (Source code repository may be found in the references section[35].) intended for the correction of misspellings made by children. Whilst typographic misspellings typically deviate from the target by only one or two characters, children’s misspellings tend to be more phonetic. They are influenced both by how the child perceives the pronunciation of a word and by the letters they choose to represent that pronunciation. As such, these misspellings are particularly deviant from the target and can negatively impact the performance of conventional spellcheckers. In this paper we demonstrate that S-capade is capable of correcting a significant portion of misspellings made by children where conventional correction tools fail.
Named entity recognition (NER) can be a challenging task, especially in highly inflected languages where each entity can have many different surface forms. We have created the first NER corpus for Icelandic by annotating 48,371 named entities (NEs) using eight NE types, in a text corpus of 1 million tokens. Furthermore, we have used the corpus to train three machine learning models: first, a CRF model that makes use of shallow word features and a gazetteer function; second, a perceptron model with shallow word features and externally trained word clusters; and third, a BiLSTM model with external word embeddings. Finally, we applied simple voting to combine the model outputs. The voting method obtains an \(F_{1}\) score of 85.79, gaining 1.89 points compared to the best performing individual model. The corpus and the models are publicly available.
In this paper, we propose to use the deep metric learning based multi-class N-pair loss, for text-to-speech (TTS) synthesis. We use the proposed loss function in a recurrent conditional variational autoencoder (RCVAE) for transferring expressivity in a French multispeaker TTS system. We extracted the speaker embeddings from the x-vector based speaker recognition model trained on speech data from many speakers to represent the speaker identity. We use mean of the latent variables to transfer expressivity for each emotion to generate expressive speech in the desired speaker’s voice. In contrast to the commonly used loss functions such as triplet loss or contrastive loss, multi-class N-pair loss considers all the negative examples which make each class of emotion distinguished from one another. Furthermore, the presented approach assists in creating a robust representation of expressivity irrespective of speaker identities. Our proposed approach demonstrates the improved performance for transfer of expressivity in the target speaker’s voice in a synthesized speech. To our knowledge, it is for the first time multi-class N-pair loss and x-vector based speaker embeddings are used in a TTS system.
The rapid and pervasive development of methods from Artificial Intelligence ( AI ) affects our everyday life. Its application improves the users’ experience of many daily tasks. Despite the enhancements provided, such approaches have a substantial limitation in the shortfall of people’s trust connected with their lack of explainability. In natural language understanding ( NLU ) and processing ( NLP ), a fundamental objective is to support human interactions using sense-making of the language for communication. Such methods try to comprehend and reproduce the self-evident processes of human communication. This applies either in receiving speech signals or in extracting relevant information from a text. Furthermore, the pervasiveness of AI methods in the workplace and on the free time demands a sustainable and verified support of users’ trust, as a natural condition for their acceptance. The objective of this work is to introduce a framework for the calculation and selection of understandable text features. Such features can increase the confidence placed into adopted NLP solutions. The following work outlines the Text Feature Framework and its text features, based on statistical information coming from a general text corpus. The showcase experiment uses those features to verify them on the concept recognition task. The results shows their capability to explain a model and its predictions. The resulting concept recognition models are competitive with other methods existing in the literature. It has the definitive advantage of being able to externalize the supporting evidence for a choice of concept identification.
In this paper, we present our experiments with BERT (Bidirectional Encoder Representations from Transformers) models in the task of sentiment analysis, which aims to predict the sentiment polarity for the given text. We trained an ensemble of BERT models from a large self-collected movie reviews dataset and distilled the knowledge into a single production model. Moreover, we proposed an improved BERT's pooling layer architecture, which outperforms standard classification layer while enables per-token sentiment predictions. We demonstrate our improvements on a publicly available dataset with Czech movie reviews.
In this paper, speech rhythm metrics were used in classification of native vs. nonnative speakers. The speech corpus exploited is a part of West Point corpus. Nonnative speakers (14) are English participants who read the same set of Arabic text then their Arabic counterpart (15). Seven rhythm metrics from all vowels and consonants were calculated from 145 sentences using two rhythm models: Interval Measures (IM) and Compensation/Control Index (CCI). Rhythm data were use as input vector of ANN-MLP classifier. The classifier was trained and tested using different configurations of the input vectors. The best accuracy of the engine achieved (80.7%) when we used all speech rhythm input vectors.
Automatic speech recognition (ASR) can be deployed in a previously unknown language, in less than 24 h, given just three resources: an acoustic model trained on other languages, a set of language-model training data, and a grapheme-to-phoneme (G2P) transducer to connect them. The LanguageNet G2Ps were created with the goal of being small, fast, and easy to port to a previously unseen language. Data come from pronunciation lexicons if available, but if there are no pronunciation lexicons in the target language, then data are generated from minimal resources: from a Wikipedia description of the target language, or from a one-hour interview with a native speaker of the language. Using such methods, the LanguageNet G2Ps now include simple models in nearly 150 languages, with trained finite state transducers in 122 languages, 59 of which are sufficiently well-resourced to permit measurement of their phone error rates. This paper proposes a measure of the distance between the G2Ps in different languages, and demonstrates that agglomerative clustering of the LanguageNet languages bears some resemblance to a phylogeographic language family tree. The LanguageNet G2Ps proposed in this paper have already been applied in three cross-language ASRs, using both hybrid and end-to-end neural architectures, and further experiments are ongoing.