
ABSTRACT: This article describes progress to date on the Canadian English Dictionary , a new general dictionary of Canadian English being edited by the recently incorporated, not-for-profit Society for Canadian English, with the support of Editors Canada, the University of British Columbia’s Canadian Word Centre (formerly Canadian English Lab), and Queen’s University’s Strathy Language Unit. The initial fascicle, Q, consisting of more than five hundred entries, is available for online review, and work is continuing alphabetically, now in fascicle R. Issues that arose in the creation of process documentation for definition, pronunciation, and etymology are discussed with specific examples.
ABSTRACT: Johnson’s “single-handed” lexicography remains a commonplace of critical discussion of his work, particularly in relation to his selection of the words and evidence which his Dictionary of the English Language contains. Based on a new examination of primary data, this article instead explores the collective and collaborative realities on which Johnsonian lexicography rested. Focusing on the work of Francis Stewart, and the evidence preserved in the seven extant volumes of the edition of Shakespeare that Johnson (and his assistants) used in making the Dictionary , it examines patterns of shared reading to challenge conventional expectations about Johnson’s data selection (including structural patterns such as phrasal verbs). While the article argues for a significant recuperation of the invisible labor on which Johnson’s Dictionary depended, it also provides a case study of Stewart’s work, between 1746 and 1752, as part of the community of practice in Johnson’s dictionary garret.
ABSTRACT: This paper documents the decision-making and evidence-based practices behind the Dictionary of Jamaican Slang ( DJS ). It presents research and preparation of the dictionary as a practice: iterative, conversational, and grounded in usage evidence drawn from Jamaican popular music, social media, and other public discourse. The paper outlines the team’s working definition of slang, its inclusion and exclusion logic for borderline items, the grammatical analysis, and the definition practices. We also describe quotation hunting as a bibliographic and methodological task, and we discuss how allonymy and colexification structure the slang lexicon and shape cross-referencing decisions.
ABSTRACT: The present article systematically investigates Ludwig Wittgenstein's elementary school dictionary from 1926 from a pluricentric viewpoint (e.g., Clyne 1984) for the first time. The slim volume of 2,920 lexemes, which has largely escaped scholarly attention, was written when Wittgenstein was teaching grade-4 pupils in rural Austria. While the few available statements to date speak of Wittgenstein's inclusion of "Austrian dialect," the present study shows these assessments as misguided and influenced by a "One Standard German Axiom." Analysis identifies fewer than twenty non-standard dialect words, while more than 12% of all entries (364) are Austrianisms, that is, terms of what would later be called Standard Austrian German; this is a rate almost five times higher than commonly offered for present-day Austrian German. This raises the question to what degree, if at all, Wittgenstein anticipated the pluricentricity (multiple standards) of German. After a brief assessment of its pedagogical purpose and key lexicographical features, the dictionary is gauged in relation to Wittgenstein's stated goal to "include only words, but all such words, that are known to Austrian elementary pupils. Therefore it excludes many a good German word unusual in Austria" (Wittgenstein 1925, 3). Embedding the analysis in biographical and philosophical accounts, this study finds that Wittgenstein, in line with his post- Tractatus philosophy that abandoned the view of language as a logical structure for messy usage-based negotiations of meaning, intuitively codified the current norm of Austrian German. This meant excluding the exonormative East Central German standard in an act quite atypical for the interwar period.
ABSTRACT: This article describes how Oxford English Dictionary ( OED ) antedating supports the teaching of digital research methods and fundamental linguistics concepts to undergraduate students. Antedating requires students to find earlier citations for existing OED entries through systematic database searching. Students submit weekly research reports that document both their discoveries and research processes. This method addresses contemporary pedagogical challenges, including student unfamiliarity with digital primary source research and the need for AI-resistant assignments. Drawing on learning research that demonstrates the benefits of "desirable difficulties" (Bjork and Bjork 2011) and "authentic activity" (Brown, Collins, and Duguid 1989), the course demonstrates how antedating can serve broader educational goals in linguistics and the digital humanities.
ABSTRACT: This paper presents a detailed account of the methodology employed in defining the word list for a Georgian–English Learner's Dictionary , which is being created at the Center for Lexicography and Language Technologies of Ilia State University. Although learner's dictionaries have been published in Georgia since the 1930s, none have been based on a word list developed using modern lexicographic methodologies. Furthermore, the absence of a fully representative corpus of the Georgian language limits the possibility of generating a frequency-based word list solely from corpus data. Given these constraints, we analyzed word lists from existing Georgian dictionaries and the frequency lists from the available Georgian language corpora. Additionally, we examined various methods for compiling frequency lists in other languages to inform our own approach.
ABSTRACT: Chinese Muslims have their own history of Arabic–Chinese bilingual lexicography. Traceable back to the Ming (1368–1644), it is a history that was intertwined with madrasa education until the Qing (1644–1911) period. Lexicographical works compiled then were subject lexicons that were intended to be a helpmeet in learning specific Islamic classical works. The first three decades of the twentieth century witnessed the modernization of Arabic–Chinese lexicography, which constituted part of the larger story of the modernization of Chinese Islamic education, and represents Chinese Muslims' efforts of religious self-salvation and self-awakening in a time of the waning of Islam in China. Progressive ahong s valued lexicography as a mode of knowledge production to expand Chinese Muslims' Arabic and Chinese vocabulary, improve their Chinese and Arabic linguistic capabilities in speech, reading, and translation, and reshape their horizons of general knowledge. We know that dictionaries can be a product, form, or tool of power relations and hegemonies of knowledge, and a language facility to standardize language of preaching for cultural and religious conversions. The Arabic–Chinese lexicography conducted by Chinese Muslims by the 1930s, however, highlights the importance of minority groups' own appeals to lexicography.
ABSTRACT: Usage citations in historical dictionaries typically consist of brief snippets that illustrate the lexeme's senses. While perhaps sufficient for the purposes of lexicography and historical linguistic work, such snippets can fail to convey the wider historical, discursive, and social contexts in which a lexeme is used. These contexts are often the primary interest of a dictionary's users, such as historians, sociologists, and sociolinguists. Early uses of hooker , the slang term for a prostitute, demonstrate how dictionary citations can mask these wider contexts and can even mislead historical linguists as to the word's origin. This article looks at the treatment of three texts by the OED Online and Green's Dictionary of Slang that constitute much of the early record of hooker 's use in the 1830s–40s. It examines how the publications and discourse communities in which the word appears provide further insight into how the word developed and its relation to the social framework of American society in the period. The use of hooker is rooted in nineteenth-century patriarchal society and, in the antebellum South, linked to the institution of chattel slavery. Users of historical dictionaries should be advised to use dictionaries' citations only as a jumping-off point and to consult the full text of the works in which they appear.
ABSTRACT: One of the main objectives of valency lexicons is to identify predicate argument structures evoked by verbs. The analysis used in these resources heavily rests on lexical semantics and as such provides an accurate description of event structure in cases in which a verb is used in its prototypical sense, such as when the verb kick is used as a force verb in the example He kicked the door . However, this approach to event structure representation proves inadequate for examples in which a verb is used in a non-prototypical sense in which additional event participants are evoked constructionally, such as the use of kick in a transfer event: He kicked him the ball . In this paper, we propose a two-tier representation of event structure in which lexical semantics and the semantics of argument structure constructions are dealt with independently. We present a detailed model of constructional analysis that is compatible with lexically based analyses of event structure used in valency lexicons, such as Propbank, and show that a mapping between lexical and constructional representations provides a more comprehensive understanding of the event semantics when compared to a lexical representation alone.
ABSTRACT: The opening up of the Internet to large-scale public access from the early 1990s was attended by a proliferation of words associated with, referring to, and circulating on the Internet. A great many dictionaries of such Internet terms were published in the course of the decade, which effectively constituted a distinctive vocabulary. This vocabulary was unusually fluid and quickly socially grounded. Dictionaries of Internet terms worked with somewhat different ideas of authority and usage from those of standard ordinary-language or technical and professional dictionaries. This paper considers the development of such dictionaries in that decade, with some necessary attention to earlier and later trajectories. Categories relevant to dictionaries of Internet terms are proposed and the relationship between them considered in a broadly chronological order. Thus, this paper offers a brief history of a neglected area of dictionary-making and an unusual historicist perspective on a formative period of the Internet. Five categories of dictionaries are considered and the relationships between them noted. Three of these started appearing before the Internet became a widely availed public feature: computer technology dictionaries for professional and pedagogic purposes, dictionaries of computer terms for laypersons, and, influentially, hacker slang/jargon dictionaries. Elements of these converged in the 1990s on dictionaries of Internet terms, within which several subcategories can be discerned. A fifth category taken up here is an offshoot of the latter: cyberdictionaries of Internet terms and slang.
ABSTRACT: In the SynSemClass project, we are building a multilingual ontology for text annotation at the semantic (meaning) level. The SynSemClass ontology consists of entries (hereafter called classes ) that represent one event type each, defined across multiple languages. Each class, representing an eventive type, contains a set of words (synonymous or semi-synonymous verbs) that express that event type, as described by its definition. Each verb is linked to both its syntactic properties and its occurrence in similar semantic lexicons for the particular language. Depending on the resource being referred to, the chain of links maps the particular class member to its sense as defined in that other resource. In addition, the semantic roles associated with the class are linked to the arguments (valency slots) as defined by the external resource, which in turn—in some cases—also contain morphosyntactic information relevant for each of the arguments. Such a mapping allows for the extraction of the correct case, adposition, word-order precedence and/or negation as well as other properties important for the corresponding surface form. In this paper, we present the system of the linked data as present in SynSemClass, together with examples taken from the languages covered.
ABSTRACT: Although semantic role labeling based on a pre-existing valency lexicon is useful in many downstream natural language processing (NLP) tasks, computational versions of these types of lexicons are not available in most languages. For parallel sentences where the source language already has automatic or manual lexical (SRL) annotation, the annotations can be transferred onto the target language text. The Universal PropBanks (UP2.0) system uses unsupervised word alignments, filtering heuristics, and bootstrapping to automatically project English SRL annotations to twenty-three languages. Since this approach uses English-based semantic representations to create annotations in other languages, it may miss language-specific nuances. We provide a case study using English PropBank representations and the Russian PropBank. We evaluate the UP2.0 projection of English annotations onto sentences that have been annotated manually with Russian PropBank, assessing discrepancies. Based on our error analysis and PropBank annotation guidelines, we identify language-specific and language-independent principles that may be violated by the automatic projection and implement these as post-processing filters to improve the precision of the automatic annotation projection.
ABSTRACT: This article charts a path—now made possible by digitization of almost all key texts—through the tangled history of eighteenth-century English/German lexicography, which emerged in response to growing German interest in English language and culture, and which rested entirely in the hands of teachers of English in Germany until the (hitherto almost completely overlooked) intervention of Adelung's English–German dictionary (1783). The dictionaries of Christian Ludwig (1706, 1716) and Theodor[e] Arnold (1736, 1739, 1757) laid the foundation for two partially interconnected lexicographical strands, until they were drawn together by Ernst Klausing (1770, 1771). While Ludwig based his works on existing bilingual dictionaries for other language combinations (French/English, Italian/German and German/French/Italian), Arnold used, besides Ludwig, Nathan Bailey's English orthographical dictionary (1727)—rather than Bailey's 1730 dictionary, as previously believed. Rogler (1763) preferred to use Samuel Johnson's recently published English dictionary (1755) as a source; so too did Adelung (1783), using the fourth, 1773 edition. In the 1790s, Johannes Ebers (whose 1796–1799 dictionary has not before been examined in detail in published work) was the first to draw on Adelung's German dictionary as an authoritative monolingual German source. Ebers was also the first to take into account the needs of English students of German, while also, like Ludwig at the start of the century, relying heavily on another bilingual dictionary, this time the German/French dictionary of Friedrich Schwan (1782–1784).
ABSTRACT: UMR (Uniform Meaning Representation) is a powerful new tool for representing the lexis-semantics interface graphically in a cross-linguistically uniform way. Its primary purpose is to support semantic parsing in natural language processing (NLP), but it is also a useful tool for documenting and analyzing semantic structures in language. In this paper, we demonstrate strengths and weaknesses of UMR in annotation of Arapaho, a highly polysynthetic and agglutinating Algonquian language. We describe lexicographical design tensions that emerge when applying UMR to languages like Arapaho and offer suggestions for refining UMR so that its goal of applying cross-linguistically can be fully realized. These tensions center on the need for a special Arapaho-UMR lexicon to guide annotators in carving up morphologically complex words into distinct predicate and argument structures in UMR graphs. Such a resource is critical for polysynthetic and agglutinating languages in which a single word may encode an entire complex proposition. We propose such a lexicon and demonstrate how it will help annotators stay as they navigate morphological complexity inside and outside the verb stem.
ABSTRACT: Traditionally, dictionaries were the primary tool for translators; however, the advent of AI-driven online translation platforms is revolutionizing information retrieval. While recent research has explored dictionary use among trainee translators, fewer studies have comprehensively addressed the diverse range of resources, including AI-based tools like ChatGPT. In this regard, this study has investigated the types of resources most frequently used by 135 trainee translators in the last two years of the four-year degree program in Translation and Interpreting (BA) and from the MA in Translation, at the University of Granada (Spain). A self-report questionnaire was devised to collect the data. The results revealed that electronic dictionaries are still very popular among trainee translators though their use tended to be higher in Master's students. Machine translation was also often used. However, ChatGPT was less frequently used among trainee translators. The widespread use of machine translation and electronic dictionaries by trainee translators highlights the need for a well-rounded teaching approach that will integrate machine translation tools but also prioritize the essential role of dictionaries.
ABSTRACT: This study investigates dictionaries’ explicit and implicit views on the category of preposition. Current English-language dictionaries, almost across the board, define prepositions as words that must take noun-phrase complements (objects). But, in conflict with these definitions, entries that label words like about, before, except, from, in, until , and with as prepositions include examples where these words have non-NP complements or none at all. I argue that this analysis is empirically inadequate and results in dictionary entries that are more complex, less internally consistent, and harder for dictionary users to navigate than is necessary or justified. Adopting a view of prepositions as characteristically taking complements, but not restricted to NP complements, would result in simpler, more accurate, and more user-friendly dictionary entries.
ABSTRACT: Among the prized possessions of Robert Peattie, a young Scottish schoolteacher who had emigrated to New Zealand in 1875, was an abridged edition of Jamieson's Scots Dictionary . Rather than treating his dictionary as a passive reminder of the life he had left behind, Peattie actively added to its information, filling the margins of his copy with evidence of how Scots words recorded by Jamieson were being used half a century later. His annotations are part memoir, part diary, offering vignettes from Peattie's boyhood in Fife, as well as recording the continuing use of Scots among his fellow emigrants in Otago. They are therefore a form of lexicographical life-writing, where headwords serve as prompts for memory, both from earlier life and recent experience. Peattie's work was given new life in the twentieth century, when his annotated copy was consulted by the editors of the Scottish National Dictionary . This paper analyzes the style and content of Peattie's annotations from his copy, which is now in Glasgow University Library, and draws on newspaper and other contemporary sources to place Peattie's work in the context of attitudes to Scots among the Scottish emigrant community in nineteenth-century New Zealand.