
The digital humanities have been very successful in proposing quantitative methods for the analysis of textual data. However, similar methods are not widespread for the study of artistic expressions that rely on motion (such as theatre). In order to develop a more robust, quantitative approach to the study of motion in theatre performances, we use video processing techniques to analyze a puppet theatre recording from Indonesia. By calculating the average speed of the different scenes, we found that there is a strong correspondence between the narrative structure and the speed of the puppets. We hope this work contributes to a development of quantitative analysis methods for the study of theatre and that it also impacts the way in which theatre documentation projects are carried out in the future.
Variation among human translations is usually invisible, little understood, and under-valued. Previous statistical research finds that translations vary most where the source items are most semantically significant or express most 'attitude' (affect, evaluation, ideology). Understanding how and why translations vary is important for translator training and translation quality assessment, for cultural research, and for machine translation development. Our experimental project began with the intuition that quantitative variation in a corpus of historical retranslations might be used to project quasi-qualitative annotations onto the translated text. We present a web-based system which enables users to create parallel, segment-aligned multi-version corpora, and provides visual interfaces for exploring multiple translations, with their variation projected onto a base text. The system can support any corpus of variant versions. We report experiments using our tools (and stylometric analysis) to investigate a corpus of forty German versions of a work by Shakespeare. Initial findings lead to more questions than answers.
Semi-automated extraction of details corresponding to narratological fabula from a corpus of narrative interviews on a single event provides decontextualized building blocks for transversal, or cross-document, narratives. With information extracted from 503 World Trade Center Task Force Interviews comprising 12,000 pages of testimony and novel visualization techniques, this article proposes a computational method for the emergence of narratives that cross beyond the boundaries of one interview. These assembled narratives, in cases like that of Chief Ganci, can document those who did not survive to tell their own story.
Pearson's chi-squared test is probably the most popular statistical test used in corpus linguistics, particularly for studying linguistic variations between corpora. Oakes and Farrow (2007) proposed various adaptations of this test to allow for the simultaneous comparison of more than two corpora while also yielding an almost correct Type I error rate (i.e. claiming that a word is most frequently found in a variety of English, when in actuality this is not the case). By means of resampling procedures, the present study shows that when used in this context, the chi-squared test produces far too many significant results, even in its modified version. Several potential approaches to circumventing this problem are discussed in the conclusion.
Journal Article The Secret Life of Pronouns. What Our Words Say About Us Get access The Secret Life of Pronouns. What Our Words Say About Us. James Pennebaker. New York: Bloomsbury Press, 2011. xii + 352 pp. ISBN 978-1-608194-80-3. $28.00 (hardback). John Nerbonne John Nerbonne University of Groningenm Search for other works by this author on: Oxford Academic Google Scholar Literary and Linguistic Computing, Volume 29, Issue 1, April 2014, Pages 139–142, https://doi.org/10.1093/llc/fqt006 Published: 05 February 2013
Although the aggregation of many linguistic variables has provided new insights into the structure of language varieties, aggregation studies have been criticized for obscuring the behavior of individual input variables. Previous solutions to this criticism consisted of extensive post-hoc calculations, simple correlation measures, or highly complex algorithms. We think that these solutions can be improved. Therefore, the current article proposes a creative use of Individual Differences Scaling ( INDSCAL) as an alternative, more straightforward solution. INDSCAL is a branch of Multidimensional Scaling, which is currently the preferred dimension reduction technique for most aggregation studies. The link to the existing methodology and the simplicity of its rationale are the main advantages of INDSCAL. The article introduces INDSCAL by means of a non-linguistic example, a discussion of the mathematical properties, and a case study on the lexical convergence between Belgian and Netherlandic Dutch in a corpus of language from 1950 and 1990. The case study shows how INDSCAL reproduces the results of a typical aggregation study, but elegantly keeps open the possibility of investigating the behavior of individual variables.
Most authorship attribution studies have focused on works that are available in the language used by the original author (Holmes, 1994; Juola, 2006) because this provides a direct way of examining an author's linguistic habits. Sometimes, however, questions of authorship arise regarding a work only surviving in translation. One example is 'Constance', the putative 'last play' of Oscar Wilde, only existing in a supposed French translation of a lost English original. The present study aims to take a step towards dealing with cases of this kind by addressing two related questions: (1) to what extent are authorial differences preserved in translation; (2) to what extent does this carry-over depend on the particular translator? With these aims, we analysed 262 letters written by Vincent van Gogh and by his brother Theo, dated between 1888 and 1890, each available in the original French and in an English translation. We also performed a more intensive investigation of a subset of this corpus, comprising forty-eight letters, for which two different English translations were obtainable. Using three different indices of discriminability (classification accuracy, Hedge's g, and area under the receiver operating characteristic curve), we found that much of the stylistic discriminability between the two brothers was preserved in the English translations. Subsidiary analyses were used to identify which lexical features were contributing most to inter-author discriminability. Discrimination between translation sources was possible, although less effective than between authors. We conclude that 'handprints' of both author and translator can be found in translated texts, using appropriate techniques.
MONK is a web-based text mining software application hosted by the University of Illinois Library that enables researchers to analyze encoded digital texts from select databases and digital archives. This study examines sets of quantitative and qualitative data to explore the usage of MONK as a research tool: the author analyzes eighteen months of web analytics data from the MONK website and responses from five interviews with MONK users to examine the ways in which MONK has been most commonly used by researchers. In the paper's analysis, the author considers the implications of MONK's use in digital humanities research and teaching, and how a digital humanities tool such as MONK can be maintained for public use. This study ultimately explores how user studies of digital humanities tools can reveal insights into humanities scholars' needs for using digital tools to pursue new research methodologies, and argues that studying the usability and preservation of digital humanities tools will enable information professionals to address humanities scholars' needs for their digital scholarship.
Firstly, Lidun Hareide and Knut Hofland describe through practical advice the compilation process of The Norwegian Spanish Parallel Corpus (NSPC) created at the University of Bergen (Norway), as well as preliminary findings from ongoing and planned research based on it. The corpus is primarily constructed for research in Translation Studies, and is built to be roughly comparable to the Spanish-English P-ACTRES corpus.
The political negotiation, erection, and fall of national and cultural borders represent an issue that frequently occupies the media. Given the historical importance of boundaries as a marker of cultural identity, as well as their function to separate and unite people, the Body Type Dictionary (BTD; Wilson, 2006) represents a suitable computerized content analysis measure to analyse vocabulary qualified to measure body boundaries and their penetrability. Out of this context, this study aimed to assess the inter-method reliability of the BTD (Wilson, 2006) in relation to Fisher and Cleveland's (1956, 1958) manual scoring system for high and low barrier personalities. The results indicated that Fisher and Cleveland's manually coded barrier and penetration imagery scores showed an acceptable positive correlation with the computerized frequency counts of the BTD's coded barrier and penetration imagery scores, thereby indicating an inter-method reliability. In addition, barrier and penetration imagery correlated positively with primordial thought language in the picture response test, and narratives of everyday and dream memories, thereby indicating correlational validity.
In this article I develop a set of simple algorithms for deriving syllable count information for words from fixed-meter poetry. The focus is on the determination of what features of language or meter might be most useful. I therefore first review what factors might be useful for this, selecting those that require as little information as possible about the language in question and making as few computational demands as possible. We end up with algorithms based on: (i) the number of syllables in each line, (ii) the number of words in each line, (iii) the number of letters in those words, and (iv) the frequency of those words.I test these algorithms on corpora from English and Welsh, getting parallel results in both cases. The results establish that the variables I identify do have significant success in deriving syllable count, but that work remains to be done.
This study investigates data from the BBC Voices project, which contains a large amount of vernacular data collected by the BBC between 2004 and 2005. The project was designed primarily to collect information on vernacular speech around the UK for broadcasting purposes. As part of the project, a web-based questionnaire was created, to which tens of thousands of people supplied their way of denoting thirty-eight variables that were known to exhibit marked lexical variation. Along with their variants, those responding to the online prompts provided information on their age, gender, and-significantly for this study-their location, this being recorded by means of their postcode. In this study, we focus on the relative frequency of the top ten variants for all variables in every postcode area. By using hierarchical spectral partitioning of bipartite graphs, we are able to identify four contemporary geographical dialect areas together with their characteristic lexical variants. Even though these variants can be said to characterize their respective geographical area, they also occur in other areas, and not all people in a certain region use the characteristic variant. This supports the view that dialect regions are not clearly defined by strict borders, but are fuzzy at best.
One of the major challenges in the process of machine translation is word sense disambiguation (WSD), which is defined as choosing the correct meaning of a multi-meaning word in a text. Supervised learning methods are usually used to solve this problem. The disambiguation task is performed using the statistics of the translated documents (as training data) or dual corpora of source and target languages. In this article, we present a supervised learning method for WSD, which is based on K-nearest neighbor algorithm. As the first step, we extract two sets of features: the set of words that have occurred frequently in the text and the set of words surrounding the ambiguous word. In order to improve the classification accuracy, we perform a feature selection process and then propose a feature weighting strategy to tune the classifier. In order to show that the proposed schemes are not language dependent, we apply the suggested schemes to two sets of data, i.e. English and Persian corpora. The evaluation results show that the feature selection and feature weighting strategies have a significant effect on the accuracy of the classification system. The results are also encouraging compared with the state of the art.
Quantifying the similarity or dissimilarity between documents is an important task in authorship attribution, information retrieval, plagiarism detection, text mining, and many other areas of linguistic computing. Numerous similarity indices have been devised and used, but relatively little attention has been paid to calibrating such indices against externally imposed standards, mainly because of the difficulty of establishing agreed reference levels of inter-text similarity. The present article introduces a multi-register corpus gathered for this purpose, in which each text has been located in a similarity space based on ratings by human readers. This provides a resource for testing similarity measures derived from computational text-processing against reference levels derived from human judgement, i.e. external to the texts themselves. We describe the results of a benchmarking study in five different languages in which some widely used measures perform comparatively poorly. In particular, several alternative correlational measures ( Pearson r, Spearman rho, tetrachoric correlation) consistently outperform cosine similarity on our data. A method of using what we call `anchor texts' to extend this method from monolingual inter-text similarity-scoring to inter-text similarity-scoring across languages is also proposed and tested.
We compute the rate of textual signals of risk of war recognizable in series of consecutive political speeches about a disputed issue serious enough to entail an international conflict. The speeches concern Iran's nuclear program. We trace textual signals forewarning of risks of war that reactions to this affair lead to. The thrust of the textual analysis rests on the interplay of affiliation and power words in continuous texts, following D. C. McClelland's model for anticipating wars. The speeches are those of Iranian President Mahmoud Ahmadinejad, US Secretary of State Hillary R. Clinton, Iranian Grand Ayatollah Ali Khamenei, and Israeli Prime Minister Benjamin Netanyahu. Prefiguring a military confrontation before it occurs involves structuring information from unstructured data. Despite such imperfect knowledge, by the end of January 2012, our results show a receding risk of war on the Iranian side, but an increasing risk on the American one, while remaining ambiguous on the Israeli one.
In recent years, great availability of various language resources in different forms as well as rapid development of computer technology and programming skills have made researchers in the fields of linguistics and computer science cooperate in solving different problems of computational linguistics and natural language processing. Building large monolingual as well as bilingual corpora in digital forms and storing them in computer memories has enabled linguists and lan- guage engineers to automatically explore techniques for processing information with the help of various computer programs without any need to manually col- lect and analyze data. One of the main applications of monolingual corpora can be seen in developing automatic spell-checking systems. In such systems, a large monolingual corpus can function as a database instead of a monolingual dictionary. In the present study, it has been tried to demonstrate the effectiveness of a large monolingual corpus of Persian in improving the output quality of a spell-checker developed for this language. In the present spelling correction system, the three phases of error detection, making suggestions, and ranking suggestions are performed in three separate stages. An experiment was carried out to evaluate the performance of the spell-checking system.
We define a model of discourse coherence based on Barzilay and Lapata’s entity grids as a stylometric feature for authorship attribution. Unlike standard lexical and character-level features, it operates at a discourse (cross-sentence) level. We test it against and in combination with standard features on nineteen book-length texts by nine nineteenth-century authors. We find that coherence alone performs often as well as and sometimes better than standard features, though a combination of the two has the highest performance overall. We observe that despite the difference in levels, there is a correlation in performance of the two kinds of features.
Even in academia, where much is done on a voluntary basis, it is generally acknowledged that good work deserves a correct remuneration. A lot of labour is involved in publishing a peer-reviewed scholarly journal and, whether access to the publication is open or subscription-based, there is always a cost involved. In the real world, no one expects anything to be free, except for a smile and the sun, maybe. Over the past couple of years, I have been contacted by a handful of scholars who announced that they do not want to contribute to or review for LLC anymore because they object to ‘giving away their research and peer reviewing for free to a publisher who charges readers and makes a profit’. I regret that this point of view diabolizes LLC by romanticizing the ideal of Open Access publication. In the first part of this editorial, I would like to take the opportunity to explain why this perspective on LLC is false on at least three points and why a decision not to devote any time or effort to LLC directly affects open access publications. In the second part of this editorial, I am presenting a report on the past record-breaking year of publishing ‘LLC. The Journal of Digital Scholarship in the Humanities’.
In this article, the application of Ant-Colony Optimization (ACO) to a morphological segmentation task is described, where the aim is to analyse a set of words into their constituent stem and ending. A number of criteria for determining the optimal segmentation are evaluated comparatively while at the same time investigating more comprehensively the effectiveness of the ACO system in defining appropriate values for system parameters. Owing to the characteristics of the task at hand, particular emphasis is placed on studying the ACO process for learning sessions of a limited duration. Morphological segmentation becomes hardest in highly inflectional languages, where each stem is associated with a large number of distinct endings. Consequently, the present article investigates morphological segmentation of words from a highly inflectional language, specifically Ancient Greek, by combining pattern-recognition principles with limited linguistic knowledge. To weigh these sources of knowledge, a set of weights is used as a set of system parameters, to be optimized via ACO. ACO-based experimental results are shown to be of a higher quality than those achieved by manual optimisation or 'randomised generate and test' methods. This illustrates the applicability of the ACO-based approach to the morphological segmentation task.