The article uses a corpus workbench (Sketch Engine) to investigate practices of evaluation in online book reviews. The reviews were taken from Goodreads, Amazon, bol.com and a number of Dutch online book discussion platforms. We look at tools that have been used to study online book reviews. Then we investigate our own collection of reviews. Findings suggest (1) that online reviews are not just centred on the reviewers’ experiences but include solid discussion of the merits of books; (2) that reviewers of suspense prefer plot and character while reviewers of literary books prefer style and story; (3) that literal and metaphorical phrases referring to the body are often used in describing positive reading experiences; and (4) that positive reviews recount parts of the story, while negative reviews try to explain why the book was a disappointment.
We apply Top2Vec to a corpus of 10,921 novels in the Dutch language. For the purposes of our research we want to understand if our topic model may serve as a proxy for genre. We 昀椀nd that topics are extremely narrowly related to an existing genre classi昀椀cation historically created by publishers. Inter-estingly we also 昀椀nd that, notwithstanding careful vocabulary 昀椀ltering as suggested by prior research, various other signals, such as author signal, stubbornly remain.
Belongs to article ‘A pretty sublime mix of WTF and OMG’. (More data to be added upon acceptance). Dutch online book reviews drawn from ODBR database. See: Boot, P. (2017). A Database of Online Book Response and the Nature of the Literary Thriller. Digital Humanities 2017. Montreal.
Being able to identify and analyse reading impact expressed in online book reviews allows us to investigate how people read books and how books affect their readers. In this paper we investigate the feasibility of creating an English translation of a rule-based reading impact model for reviews of Dutch fiction. We extend the model with additional rules and categories to measure reading impact in terms of positive and negative feeling, narrative and stylistic impact, humour, surprise, attention, and reflection. We created ground truth annotations to evaluate the model and found that the translated rules and new impact categories are effective in identifying certain types of reading impact expressed in English book reviews. However, for some types of impact the rules are inaccurate, and for most categories they are incomplete. Additional rules are needed to improve recall, which could potentially be enhanced by incorporating Machine Learning. At the same time, we conclude that some impact aspects are hard to extract with a rule-based model. When applying the model to a large set of reviews, lists of the top-scoring books in the impact categories show the model's prima-facie validity. Correlations among the categories include some that make sense and others that require further research. Overall, the evidence suggests that for investigating the impact of books, manually formulated rules are partially successful, and are probably best used in a hybrid approach.
Linguistic Inquiry and Word Count (LIWC) is a text analysis program developed by James Pennebaker and colleagues. At the basis of LIWC is a dictionary that assigns words to categories. This dictionary is specific to English. Researchers who want to use LIWC on non-English texts have typically relied on translations of the dictionary into the language of the texts. Dictionary translation, however, is a labour-intensive procedure. In this paper, we investigate an alternative approach: to use Machine Translation (MT) to translate the texts that must be analysed into English, and then use the English dictionary to analyse the texts. We test several LIWC versions, languages and MT engines, and consistently find the machine-translated text approach performs better than the translated-dictionary approach. We argue that for languages for which effective MT technology is available, there is no need to create new LIWC dictionary translations.
Prominent among the social developments that the web 2.0 has facilitated is digital social reading (DSR): on many platforms there are functionalities for creating book reviews, 'inline' commenting on book texts, online story writing (often in the form of fanfiction), informal book discussions, book vlogs, and more. In this article we argue that DSR offers unique possibilities for research into literature, reading, the impact of reading and literary communication. We also claim that in this context computational tools are especially relevant, making DSR a field particularly suitable for the application of Digital Humanities methods. We draw up an initial categorization of research aspects of DSR and briefly examine literature for each category. We distinguish between studies on DSR that use it as a lens to study wider processes of literary exchange as opposed to studies for which the DSR culture is a phenomenon interesting in its own right. Via seven examples of DSR research we discuss the chosen approaches and their connection to research questions in literary studies.
What is the impact of reading fiction? We analyse online Dutch book reviews to detect overall affective impact, narrative feelings, response to style and reflection. We create a set of rules that analyze the reviews and detect the impact aspects. We evaluate the detection by asking raters about the presence of these aspects in reviews and comparing these ratings to our detection. Inter rater agreements are weak to moderate; however, there is a significant correlation between the model's predictions for all impact aspects except reflection. The detected impact correlates with book genres in the way one would expect: narrative feelings are highest for thrillers, stylistic response is highest for literary books. We can thus estimate some aspects of the response books evoke in readers. Initial results suggest that the appreciation of style is linked to reflection in the reader. However, the concepts underlying the impact categories need further exploration.
In online book reviews readers often describe their reading experience and the impression that a book left. The great volume of online reviews makes these reviews a great source for investigating the impact books have on readers. Recently, a reading impact model was introduced that can be used to automatically identify expressions of reading impact in Dutch reviews and that is able categorise them according to emotional impact, aesthetic or narrative feeling, or feelings of reflection. This paper provides an analysis of the characteristics of the book review domain that affect how this computational model identifies impact. We look at features like the length of reviews, the nature of the website on which the review was published, the genre of book and the characteristics of the reviewer. The findings in this paper provide insight in how different selection criteria for reviews can be used to study various aspects of reading impact.
Naturalism is one of the best-studied literary movements in (Dutch) literary history. Ton Anbeek formulated in 1982 eight characteristics of Dutch naturalist fiction. Working within the digital library environment Nederlab, we test these characteristics by applying LIWC (Linguistic Inquiry and Word Count) to a corpus of naturalist fiction and a reference corpus of other fiction from the same period. We confirm some of Anbeek’s claims (naturalist novels are about nervous characters experiencing a process of disenchantment), fail to confirm some (we find no evidence for the role of determinism), and cannot test some others (e.g. the naturalist author despises bourgeois society). We do find some new ‘negative’ characteristics of Dutch Naturalism: subjects that occur significantly less in naturalist than in non-naturalist texts. Among these are words related to work, achievement and money. These findings intuitively fit with our idea of a naturalist character but require further study.
Online book discussion is a popular activity on weblogs, specialized book discussion sites, booksellers’ sites and elsewhere. These discussions are important for research into literary reception and should be made and kept accessible for researchers. This article asks what an archive of online book discussion should and could look like, and how we could describe such an archive in terms of some of the central concepts of textual scholarship: work, document, text, transcription and variant. What could an approach along the lines of textual scholarship mean for such a collection? If such a collection holds many pieces of information that would not usually be considered text (such as demographic information about contributors), could we still call such a collection an edition, and could we call editing the activity of preparing such a collection?The article introduces some of the relevant (Dutch-language) sites, and summarizes their properties (among others: they are dynamic and vulnerable, they contain structured data and are very large) from the perspective of creating a research collection. It discusses the interpretation of some essential terms of textual studies in this context, and briefly lists a number of components that a digital edition of these sites might or should contain. It argues that such a collection is the result of scholarly work and should not be considered as 'just' a web archive.
The words we use in everyday language reveal our thoughts, feelings, personality, and motivations. Linguistic Inquiry and Word Count (LIWC) is a software program to analyse text by counting words in 66 psychologically meaningful categories that are catalogued in a dictionary of words. This article presents the Dutch translation of the dictionary that is part of the LIWC 2007 version. It describes and explains the LIWC instrument and it compares the Dutch and English dictionaries on a corpus of parallel texts. The Dutch and English dictionaries were shown to give similar results in both languages, except for a small number of word categories. Correlations between word counts in the two languages were high to very high, while effect sizes of the differences between word counts were low to medium. The LIWC 2007 categories can now be used to analyse Dutch language texts.
The study of literature has traditionally focused on the literary work, and sometimes its author, rather than on the response that works evoked in their readers. The arrival of the computer in the study of literature has not really changed that – perhaps unsurprisingly, as reader response has never been systematically recorded. The fact that readers have begun to document their reading and reading response on websites is therefore very fortunate (Gruzd and Rehberg Sedo, 2012, Maryl, 2008: 390406). Booksellers’ sites such as Amazon and review sites such as Goodreads, as well as weblogs, forums and general-purpose social media sites provide access to first-hand reading reports. Though most research on these sites focusses on (behavior of) users (Nakamura, 2013: 238-243, Thomas and Round, 2016: 239-253), we are beginning to see them being used in literary research (Finn, 2011) . This paper presents the Online Dutch Book Response (ODBR) database, that was designed to facilitate research into book response. At present, the database holds reviews and other response items from an online bookseller as well as from four Dutchlanguage mass review sites, including, where available, the information about books, reviewers and sites necessary to put the reviews into context. To show one type of research that the database supports, the paper displays a clustering of reviews by genre, based on frequently used words. I discuss the clustering and what it suggests for further research.
In the scholarly domain, annotation is a fundamental activity (Unsworth, 2000). Current webbased annotation facilities enable a specific way of annotation (via note-taking, highlighting or commenting) which are useful when scholars are exploring or gathering an initial set of resources, but more sophisticated support is needed for detailed analysis, close reading, and data enrichment. At this point, it is important to take into account the structural relations between documents and their parts. For example, when annotating a letter, annotation tools should be aware that a targeted text fragment is the name of the sender, or that the annotation of a film targets the intellectual work instead of the specific version or copy on which the annotation is made. In addition, many standalone tools use annotation models with idiosyncratic solutions to enable the relations between different media objects and their parts, which limits the possibilities to exchange those annotations . In general, there is a lack of necessary details for durable access to and interpretation of annotations. For this, detailed information is needed about the annotated object, the annotator and the annotation itself (Melgar et al. 2016, Walkowski & Barker, 2010). In this paper we focus on the requirements for the annotated object, in a web-based environment, and propose a method for making necessary details of objects openly available for any annotation tool.
Young Agents: the Young Author’s Role on the Dutch Republic’s Book Market1 In this article, we investigate the role of young authors on the upcoming and flourishing book market of the Dutch Republic (1550-1800), focusing in particular on their ‘agency’. By investigating the specific contribution of young authors to this market and by discriminating between roles of young adults and adults, we introduce a new approach to early modern authorship. Combining quantitative digital experiments and qualitative textual analyses, we seek to tease out the dynamic relationship between the young authors’ independence, competences, and behaviour on the one hand, and representations, self-images and production figures on the other. The first results of our research reveal that young authors frequently showed themselves indebted to their masters, but also created their own voice, often self-confident and engaged, by appropriating specific genres and topics in new ways. Contrary to what contemporary poetics and well-known forms of (self)reflection suggest, young authors appear to have had a clear, outspoken presence in the Dutch book market which became ever more prominent in both the consumption and production of Dutch books.
Annotation in digital scholarly editions (of historical documents, literary works, letters, etc.) has long been recognized as an important desideratum, but has also proven to be an elusive ideal. In so far as annotation functionality is available, it is usually developed for a single edition and cannot easily be deployed elsewhere. There is a very fundamental reason for this: the interface to the edition that the user gets to see, is only a presentation layer, usually in HTML, perhaps as a pdf or as an e-book. What the user sees is a temporary representation of the edition, the true structure of which remains 'under the hood', usually in an XML file, a set of XML files or a database. Annotation should refer to structural and stable objects (verse lines, words in lines, paragraphs, document sections, etc.) in the edition's source files, that cannot necessarily be identified from the HTML representation of the text. It is important to understand that this internal structure can be very complex, containing different sorts of editorial annotation, perhaps multiple layers of critical apparatus comparing the text to earlier or later versions, in the case of hand-written texts information about deletions in and additions to the text, and often voluminous editorial introductory material, linked to the text. An important challenge for applying a generic annotation tool for scholarly editions is thus to develop a set of agreements about how this highly complex internal structure of the edition can be represented in (or linked from) the HTML display. The annotation tool should be able to read that information and construct durable anchors for the annotations. These anchors should facilitate display of the relevant annotations in different edition display formats and should remain stable while other aspects of the edition's interface may change, depending on changes in the environment, design trends and developing technology. In our talk, we will discuss how the information about the edition's structure can be made available to annotation tools. Technologies that we will consider include RDFa, Shared Canvas/IIIF and CTS URN's. We will discuss their strengths and weaknesses in this context and propose a solution.
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
Since he coined the term in 2000, Franco Moretti’s notion of ‘distant reading’ has become very popular. It is no wonder that, looking for a title for a collection of essays, Moretti or his publisher chose that appealing, programmatic, and polemical term. Distant Reading brings together ten essays, published between 1994 and 2011, and shows both continuities and developments in Moretti’s thought. One prominent example of development is the meaning of the term ‘distant reading’ itself. In its original formulation, in the essay ‘Conjectures on World Literature’, the term had nothing to do with the digital analysis of literature we associate it with today, and not even with the visualizations and mappings of e.g. Graphs, Maps and Trees (Moretti 2005). Instead, ‘distant reading’ originally referred to the study of world literature while relying on studies done by other researchers. The size and linguistic diversity of the world’s literatures make this inevitable: ‘… literary history … will become “second hand”: a patchwork of other people’s research, without a single direct textual reading. Still ambitious … but the ambition is now directly proportional to the distance from the text …’. (p. 48)