Introduction: Quantitative methods in health, environmental, and theoretical translation research. Un article de la revue Meta (Pour de nouvelles méthodes en traductologie quantitative) diffusée par la plateforme Érudit.
Abstract Of the commonly-used measures of lexical association or collocation strength, only some directly relate to statistical significance: the t-score, chi-squared, log-likelihood, the z-score and Fisher’s exact test. We describe each of these tests, and also describe a computer simulation by which we can derive confidence limits, and hence the statistical significance, of any measure of lexical association which is derived from the contingency table. We illustrate this approach using pointwise mutual information (PMI). We also describe how the Poisson distribution enables us to find the statistical significance of the raw frequency with which a collocation is found. We compare all these methods using collocates of “take”, namely “take up”, “take place”, “take advantage” and “take stock”.
Tognini-Bonelli (2001) made the following distinction between corpus-based and corpus-driven studies. While corpus-based studies start with pre-existing theories which are tested using corpus data, in corpus driven studies the hypothesis is derived by examination of the corpus evidence. This chapter will give an overview of the two different families of statistical tests which are suited for these two approaches. For corpus-based approaches, we use more traditional statistics, such as the t-test, or ANOVA which return a value called a p-value to tell us to what extent we should accept or reject the initial hypothesis. Multi-level modelling (also known as mixed modelling) is a new technique which shows considerable promise for corpus-based studies, and will also be described here to analyse the ENNTT subset of Europarl corpus. Multi-level modelling is useful for the examination of hierarchically structured or “nested” data, where for example translations may be “nested” together in a class if they have the same language of origin. A multi-level model takes account both of the variation between individual translations and the variation between classes. For example, we might expect the scores (such as vocabulary richness or readability scores) of two translations in the same class to be more similar to each other than two translations in different classes.
This paper follows a recent court case which judged that “Don't despair” (سأيت لا) by Aaidh ibn Abdullah Al-Qarni was plagiarized from Salwa Aladian’s “That is how they defeat Despair” ( اذكه سأيلا اومزه). We use techniques of computational stylometry, Hierarchical Agglomerative Clustering Analysis and Principal Components Analysis, to show that the disputed sections not only resemble Salwa’s work in content, but also
Computer stylometry is the computer analysis of writing style. We use the computer stylometric techniques of Hierarchical Cluster Analysis, Principal Component Analysis and Machine Learning to examine the authorship of “The Liberation of Women” which is normally attributed to Qassim Amin. In particular we examine the assertion by Mohamed Emara that certain chapters of this book were written secretly by Mohammad Abdu, who was the Grand Mufti of Egypt. In our experiments, we consistently find that Qassim Amin is the more likely author of the disputed text. The experiments described in this paper were done using the “Stylometry in R” package of Eder et al. (2016).
Author profiling is the analysis of people’s writing in an attempt to find out which classes they belong to, such as gender, age group or native language. Many of the techniques for author profiling are derived from the related task of Author Identification, so we will look at this topic first. Author identification is the task of finding out who is most likely to have written a disputed document, and there are a number of computational approaches to this. The three main subtasks are the compilation of corpora of texts known to be written by the candidate authors, the selection of linguistic features to represent those texts, and statistics for discriminating between those features which are most indicative of a particular author’s writing style. Plagiarism is the unacknowledged use of another author’s original work, and we will look at software for its detection. The chapter will cover the types of text obfuscation strategies used by plagiarists, commercial plagiarism detection software and its shortcomings, and recent research systems. Strategies have been developed for both external plagiarism detection (where the original source is searched for in a large document collection) and intrinsic plagiarism detection (where the source text is not available, necessitating a search for inconsistencies within the suspicious document). The specific problems of plagiarism by translation of an original in another language, and the unauthorized copying of sections of computer code, are described. Evaluation forums and publicly available test data sets are covered for each of the main topics of this chapter.
This chapter raises and addresses the key question of the relation between empirical translation studies and the descriptive branch of pure translation research initially proposed in the late 1980s. It analyses, explains, and illustrates the rationale and viability of the proposed ‘empirical turn’ in translation studies, which is to better align the central aims and purposes of the discipline with new, practical research needs arisen from our changing social environments in many parts of the world. It is argued in the conclusion that as demonstrated by the diverse chapters in the book, translation studies, especially the descriptive, empirical research branch has benefited and will continue to benefit from the integration of advanced quantitative research methodologies and advances in corpus analytical software development and natural language processing technologies such as machine translation. This bourgeoning translation research field is thus well equipped to play a larger, more significant role in addressing practical, pressing social and research issues such as environment communication, sustainable development and health literacy which is proposed as the much-needed ‘social turn’ in empirical translation studies.
This book in the Edinburgh Textbooks in Empirical Linguistics series is a comprehensive introduction to the statistics currently used in corpus linguistics. Statistica l techniques and corpus applications - whether oriented towards linguistics or language engineering - often go hand in glove, and corpus linguists have used an increasingly wide variety of statistics, drawing on techniques developed in a great many fields.This is the first one-volume introduction to the subject.
In this paper we investigate differences and similarities between dialects using unsupervised learning. We used a binary phonetic representation to cluster utterances from different Arabic and English dialects. This phonetic representation aims to capture phonetic patterns such as vowel and consonant length. We tested this representation on an Arabic dataset containing utterances from speakers of four dialects: Egyptian, Gulf, Levantine, and North African. We validate our approach on an English dataset containing utterances from speakers from Bradford, Cardiff, Dublin, and Liverpool.
In this paper we describe new results of statistical and neural data mining of audiology patient records, with the ultimate aim of looking for factors influencing which patients would most benefit from being fitted with a hearing aid. We describe how a combination of neural and statistical techniques can usefully subdivide a set of patients into clusters, based on their hearing thresholds at six different frequencies, and then label the clusters with meaningful text labels. In our first experiment we cluster the patients based on similarities between their audiograms using k-means clustering, resulting in two main clusters. We then use the chi-squared test to label each cluster with the keywords selected from the text comment, diagnosis and hearing aid type associated with each patient which are most typical (and atypical) of each cluster. In our second experiment, we again cluster the patients based on similarities between their audiograms, but this time using a selforganizing map (SOM). Here the locations in the resulting map, corresponding to individual patients, are labeled with the type of hearing aid selected for each patient. We demonstrate that this automatic textual labeling addresses well the heterogeneous character of medical audiology records, since they consist of numeric, structured and free text data. ————— Appearing in Proceedings of the 19 th Machine Learning conference of Belgium and The Netherlands. Copyright 2010 by the author(s)/owner(s).
This chapter is set in the context of Corpus Pattern Analysis (CPA), a technique developed by Patrick Hanks to map meaning onto word patterns found in corpora. The main output of CPA is the Pattern Dictionary of English Verbs (PDEV), currently describing patterns for over 1,600 verbs, many of which are acknowledged to be multiword expressions (MWEs) such as phrasal verbs or idioms. PDEV entries are manually produced by lexicographers, based on the analysis of a substantial sample of concordance lines from the corpus, so the construction of the resource is very time-consuming. The motivation for the work presented in this chapter is to speed up the discovery of these word patterns, using methods which can be transferred to other languages. This chapter explores the benefits of a detailed contrastive analysis of MWEs found in English and French corpora with a view on English-French translation. The comparative analysis is conducted through a case study of the pair (bite, mordre), to illustrate both CPA and the application of statistical measures for the automatic extraction of MWEs. The approach taken in this chapter takes its point of departure from the use of statistics developed initially by Church & Hanks (1989). Here we look at statistical measures which have not yet been tested for their ability to discover new collocates, but are useful for characterizing verbal MWEs already found. In particular we propose measures to characterize the mean span, rigidity, diversity, and idiomaticity of a given MWE.
—We perform data mining on the publicly available Tinnitus Archive. A number of statistically significant associations with gender were found using the Chi-squared test. These were age, onset rapidity, tinnitus localisation, number of tinnitus sounds heard, sleep interference due to tinnitus, feeling tired and ill because of tinnitus, index of noise exposure and subjective tinnitus pitch. No other associations with gender were statistically significant. In each case where a factor was found to be associated with gender, we analysed the data further by examining the standardised residuals. These showed that more men with tinnitus were younger, and more men had experience of noise exposure. Women were more likely to hear more than three sounds in their tinnitus, hear it in both ears, experience gradual onset of tinnitus, and hear it at lower pitches. Our findings are confirmed by the use of a measure derived from market basket analysis, that of lift.
In this chapter we show that corpora, particularly parallel bilingual corpora, are essential in the development of automatic machine translation (MT) systems, whether translation memories, example-based or statistical. Specific topics examined are the Europarl corpus, similarity measures for sentence matching, the Hofland sentence aligner, automatic generalisation of translation examples through paraphrasing and the discovery of templates, statistical methods of building bilingual dictionaries, the development of MT for less-resourced languages and the evaluation of MT systems.
Arud is the metrical system used in classical Arabic poetry. This paper introduces a method based on an encoding of Arud for the automatic discrimination of authors. This measure was used for both Arabic and English texts in the domain of travel. Perfect classification, using hierarchical cluster analysis (Ward's method) and Burrows' Delta measure of inter-text distances, was obtained when using either the Arud-based encoding or the more common 100 most frequent words as the linguistic feature to represent the texts, and the text samples were of 1000 words in length. For shorter texts of 500 words, the Arudbased measure outperformed the baseline of the 100 most frequent words, as measured by the Rand Index. Both sets of features produced better clustering for the Arabic data set than the English data set for the shorter texts.
The PARSEME shared task aims at identifying verbal MWEs in running texts. Verbal MWEs include idioms (let the cat out of the bag), light verb constructions (make a decision), verb-particle constructions (give up), and inherently reflexive verbs (se suicider 'to suicide' in French). VMWEs were annotated according to the universal guidelines in 18 languages. The corpora are provided in the parsemetsv format, inspired by the CONLL-U format. For most languages, paired files in the CONLL-U format - not necessarily using UD tagsets - containing parts of speech, lemmas, morphological features and/or syntactic dependencies are also provided. Depending on the language, the information comes from treebanks (e.g., Universal Dependencies) or from automatic parsers trained on treebanks (e.g., UDPipe). This item contains training and test data, tools and the universal guidelines file.
This article looks at the provenance of the unfinished novel The Dark Tower, generally attributed to C. S. Lewis. The manuscript was purportedly rescued from a bonfire shortly after Lewis's death by his literary executor Walter Hooper, but the quality of the text is hardly vintage Lewis. Using computer stylometric programs made available by Eder et al.'s (2016: Stylometry with R: A package for computational text analysis. R Journal, 8(1): 107-21) 'stylo' package and a word length analysis, samples of each chapter of The Dark Tower were compared with works known to be by Lewis, two books by Hooper and a hoax letter concerning the bonfire by Anthony Marchington. Initial experiments found that the first six chapters of The Dark Tower were stylometrically consistent with Lewis's known works, but the incomplete Chapter 7 was not. This may have been due to an abrupt change in genre, from narrative to pseudoscientific style. Using principal components analysis, it was found that the first and subsequent components were able to separate genre and individual style, and thus a plot of the second against the third principal components enabled the effects of genre to be filtered out. This showed that Chapter 7 was also consistent with the other samples of C. S. Lewis's writing.
This paper describes a new approach to building the query based relevance sets (qrels) or relevance judgments for a test collection automatically without using any human intervention. The methods we describe use supervised machine learning algorithms, namely the Naïve Bayes classifier and the Support Vector Machine (SVM). We achieve better Kendall's tau and Spearman correlation results between the TREC system ranking using the newly generated qrels and the ranking obtained from using the human-built qrels than previous baselines. We also apply a variation of these approaches by using the doc2vec representation of the documents rather than using the traditional tf-idf representation.