Investigating serious crimes is inherently complex and resource-constrained. Law enforcement agencies (LEAs) grapple with overwhelming volumes of offender and incident data, making effective suspect identification difficult. Although machine learning (ML)-enabled systems have been explored to support LEAs, several have failed in practice. This highlights the need to align system behavior with stakeholder goals early in development, motivating the use of Goal-Oriented Requirements Engineering (GORE). This paper reports our experience applying the GORE framework KAOS to designing an ML-enabled system for identifying suspects in online child sexual abuse. We describe how KAOS supported early requirements elaboration, including goal refinement, object modeling, agent assignment, and operationalization. A key finding is the central role of data elicitation: data requirements constrain refinement choices and candidate agents while influencing how goals are linked, operationalized, and satisfied. Conversely, goal elaboration and agent assignment shape data quality expectations and collection needs. Our experience highlights the iterative, bidirectional dependencies between goals, data, and ML performance. We contribute a reference model for integrating GORE with data-driven system development, and identify gaps in KAOS, particularly the need for explicit support for data elicitation and quality management. These insights inform future extensions of KAOS and, more broadly, the application of formal GORE methods to ML-enabled systems for high-stakes societal contexts.
This article evaluates the reliability, efficiency, and effectiveness of Linguistic Inquiry and Word Count (LIWC; Boyd et al., 2022) for the analysis of a white nationalist forum. This is important because LIWC has been the computational tool of choice for scores of studies generally and many examining extremist content in a forensic or security context. Our purpose, therefore, is to understand whether LIWC can be depended upon for large-scale analyses; we initially examine this here using a small sample of posts from a set of just eight users and manually checking the program's automated codings of a subset of categories. Our results show that the LIWC coding cannot be relied upon - precision falls to as low as 49.6 % and recall as low as 41.7 % for some categories. It would be possible to engage in considerable manual correction of these results, but this undermines its purported efficiency for large datasets.
This paper reports an initial application of contemporary cognitive stylistics to forensic linguistic contexts. In both areas, a need has been identified for robust analyses. An intercoder reliability study was developed using data from a historic authorship analysis case involving single-authored hate mail. Exploring the applicability of Cognitive Grammar’s notion of construal as a reliable framework for describing salient features of the author’s style, this test examined the accuracy and consistency of descriptions of schematicity and specificity within the corpus, as applied by independent coders. Iterative coding and testing demonstrated that reliability was achievable, but depended upon a protocol developed through considerable definitional work, refining the concepts of specificity and elaboration as taken from Cognitive Grammar. Our findings support the idea that the identification of stylistic features can be rigorous, retrievable, and replicable, but also that a fuller system of coding will require a substantial research programme. Such an approach, bringing together contemporary stylistics and forensic authorship analysis, would be a productive collaboration between both disciplines and a valuable research method for verifiability in stylistics more generally. Content: Readers are advised that the letters analysed for this study contain offensive language, and that short quotes within this paper include racist and hateful language directed at particular groups.
Artificial Intelligence (AI) has become an important part of our everyday lives, yet user requirements for designing AI-assisted systems in law enforcement remain unclear. To address this gap, we conducted qualitative research on decision-making within a law enforcement agency. Our study aimed to identify limitations of existing practices, explore user requirements and understand the responsibilities that humans expect to undertake in these systems. Participants in our study highlighted the need for a system capable of processing and analysing large volumes of data efficiently to help in crime detection and prevention. Additionally, the system should satisfy requirements for scalability, accuracy, justification, trustworthiness and adaptability to be adopted in this domain. Participants also emphasised the importance of having end users review the input data that might be challenging for AI to interpret, and validate the generated output to ensure the system's accuracy. To keep up with the evolving nature of the law enforcement domain, end users need to help the system adapt to the changes in criminal behaviour and government guidance, and technical experts need to regularly oversee and monitor the system. Furthermore, user-friendly human interaction with the system is essential for its adoption and some of the participants confirmed they would be happy to be in the loop and provide necessary feedback that the system can learn from. Finally, we argue that it is very unlikely that the system will ever achieve full automation due to the dynamic and complex nature of the law enforcement domain.
This paper sets the stage for our primary objective, which is to identify and examine various forms of claimed expertise in anonymous online interactions. By building upon the findings and incorporating the proposed enhancements, we aim to gain a deeper understanding of the nature and implications of different expertise claims within the context of power hierarchies. A combination of various machine learning techniques is employed in this work, including classical methods, deep learning models, and transformer-based approaches to create classification models, while using three datasets collected by specialists and annotated by linguistics experts. The first experiments’ results in binary classification, indicating whether a given post reflects expertise or not, are particularly promising, especially when utilising transformer-based approaches. The second set of experiments, focusing on the classification of different types of expertise, produced a diverse range of results with the less favourable results primarily caused by an imbalance in labelling between different classes.
The purpose of this paper is to provide both a theoretical foundation and apractical framework for analysing power and authority in online interactions.This is to assist forensic linguists and law enforcement in their understandingof anonymous online criminal networks, and the roles of individuals in theseonline communities. The lack of contextual knowledge present in anonymousonline fora creates a challenge for the analyst in finding a framework totheorise, explore, and describe different types of power performance and thusthe different roles of interactants in these fora. In this paper, we provide aframework to describe the basis on which individuals make claims to powerand use this framework to explore the nature and distribution of power acrossdifferent fora of both criminal and benign intent. This is developed throughan analysis of three online discussion fora, of approximately 160,000 words,resulting in a framework of nine main categories of power resource. Thisallows us to contrast the three fora, showing differences in the nature anddistribution of power resource, and also enables description of individualsas high-resource or low-resource with regards to their claims to power. Thistheory and framework can also be productive in the analysis of language andpower in computer mediated communication (CMC) more widely.
In this Element, the authors introduce and apply a framework for the linguistic analysis of fake news. They define fake news as news that is meant to deceive as opposed to inform and argue that there should be systematic differences between real and fake news that reflect this basic difference in communicative purpose. The authors consider one famous case of fake news involving Jayson Blair of The New York Times, which provides them with the opportunity to conduct a controlled study of the effect of deception on the language of a single reporter following this framework. Through a detailed grammatical analysis of a corpus of Blair's real and fake articles, this Element demonstrates that there are clear differences in his writing style, with his real news exhibiting greater information density and conviction than his fake news. This title is also available as Open Access on Cambridge Core.
The Aston Forensic Linguistic Databank (FoLD) is a permanent,controlled access online repository for forensic linguistic data. We broadlyunderstand forensic linguistics as any academic research with a potential toimprove the delivery of justice through the analysis of language. FoLD thuscomprises a wide range of datasets with relevance to forensic linguistics andlanguage and law, including commercial extortion letters, investigative interviewsin police and other contexts, legal documents, forum posts from far-right onlinegroups, and comment threads from political blogs. This paper outlines how FoLDworks and its potential impact on the general discipline of forensic linguistics.
The paper presents a two-part forensic linguistic analysis of an historic collection of abuse letters, sent to individuals in the public eye and individuals' private homes between 2007 and 2009. We employ the technique of structural topic modelling ( STM) to identify distinctions in the core topics of the letters, gauging the value of this relatively under-used methodology in forensic linguistics. Four key topics were identified in the letters, `Politics A' and `B', `Healthcare' and `Immigration', and their coherence, correlation and shifts in topic were evaluated. Following the STM, a qualitative corpus linguistic analysis was undertaken, coding concordance lines according to topic, with the reliability between coders tested. This coding demonstrated that various connected statements within the same topic tend to gain or lose prevalence over time, and ultimately confirmed the consistency of content within the four topics identified through STM throughout the letter series. The discussion and conclusions to the paper reflect on the findings and also consider the utility of these methodologies for linguistics and forensic linguistics in particular. The study demonstrates real value in revisiting a forensic linguistic dataset such as this to test and develop methodologies for the field.
This Element examines progress in research and practice in forensic authorship analysis. It describes the existing research base and examines what makes an authorship analysis more or less reliable. Further to this, the author describes the recent history of forensic science and the scientific revolution brought about by the invention of DNA evidence. They chart the rise of three major changes in forensic science – the recognition of contextual bias in analysts, the need for validation studies and shift in logic of providing identification evidence. This Element addresses the idea of progress in forensic authorship analysis in terms of these three issues with regard to new knowledge about the nature of authorship and methods in stylistics and stylometry. The author proposes that the focus needs to shift to validation of protocols for approaching case questions, rather than on validation of systems or general approaches. This title is also available as Open Access on Cambridge Core.
Chapter 2 The Starbuck Case Methods for Addressing Confirmation Bias in Forensic Authorship Analysis Tim Grant, Tim GrantSearch for more papers by this authorJack Grieve, Jack GrieveSearch for more papers by this author Tim Grant, Tim GrantSearch for more papers by this authorJack Grieve, Jack GrieveSearch for more papers by this author Book Editor(s):Ria Perkins, Ria PerkinsSearch for more papers by this authorIsabel Picornell, Isabel PicornellSearch for more papers by this authorMalcolm Coulthard, Malcolm CoulthardSearch for more papers by this author First published: 04 April 2022 https://doi.org/10.1002/9781394266661.ch2 AboutPDFPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onEmailFacebookTwitterLinkedInRedditWechat Summary Nearly two years previously, Debbie married Jamie Starbuck following a relatively brief courtship. Since Dror et al.'s work gave the issue prominence, confirmation bias in forensic evidence has received considerable attention. One key feature of how the authors tackled the Starbuck case was a separation of roles between the two analysts Grant and Grieve (TG and JG). One basic distinction between a stylistic approach and a stylometric approach is that the stylistic approach generally involves a data-driven generation of a case-specific feature set, whereas stylometric analysis tends to rely on predesigned feature sets. A more effective part of the strategy to mitigate bias was the restriction in the flow of text to JG as the analyst. TG wrote a formal expert witness report, explaining the method and crediting JG with his role in the analysis. Risks of unconscious confirmation bias can be mitigated but perhaps never avoided altogether. Further reading Argamon , S. ( 2018 ). Computational forensic authorship analysis: Promises and pitfalls . Language and Law/Linguagem e Direito , 5 ( 2 ), 7 – 37 . Google Scholar Dror , I.E. , Peron , A. E. , Hind , S. L. , & Charlton , D. ( 2005 ). When emotions get the better of us: the effect of contextual top-down processing on matching fingerprints . Applied Cognitive Psychology , 19 ( 6 ), 799 – 809 . 10.1002/acp.1130 Web of Science®Google Scholar Grant , T. ( 2020 ). Text messaging forensics: Txt 4n6: Idiolect free authorship analysis? In M. Coulthard , A. May , & R. Sousa-Silva (eds.), The Routledge Handbook of Forensic Linguistics . ( 2nd ed.). Routledge . Google Scholar Grant , T. ( 2012 ). TXT 4N6: method, consistency, and distinctiveness in the analysis of SMS text messages . Journal of Law & Policy , 21 , 467 . Google Scholar Grant , T. ( 2021 ). The idea of progress in forensic authorship analysis . Cambridge University Press . Google Scholar Grieve , J. , & Woodfield , H. ( 2020 ). Investigative linguistics . In M. Coulthard , A. May, & R. Sousa-Silva (eds.), The Routledge Handbook of Forensic Linguistics ( 2nd ed.). Routledge . Google Scholar Suggested Research Questions The susceptibility of authorship analysis to contextual bias could be explored experimentally. i.e. telling authorship analysts alternative stories around a set problem, and seeing if they come up with different answers. This could be applied to both stylistic and computational approaches to authorship analysis. It would be useful to explore ways in which computational and more qualitative approaches to authorship analysis might be combined. e.g. using heavily computational methods to elicit a large set of features but to also to examine those features to rule out ones without linguistic explanation. References Argamon , S. ( 2018 ). Computational forensic authorship analysis: Promises and pitfalls . Language and Law/Linguagem E Direito , 5 ( 2 ), 7 – 37 . Google Scholar Biber , D. ( 1995 ). Dimensions of register variation . Cambridge University Press . 10.1017/CBO9780511519871 Google Scholar Coulthard , M. ( 2004 ). Author identification, idiolect, and linguistic uniqueness . Applied Linguistics , 25 ( 4 ), 431 – 447 . 10.1093/applin/25.4.431 Web of Science®Google Scholar Dror , I. E. , Peron , A. E. , Hind , S. L. , & Charlton , D. ( 2005 ). When emotions get the better of us: The effect of contextual top-down processing on matching fingerprints . Applied Cognitive Psychology , 19 ( 6 ), 799 – 809 . 10.1002/acp.1130 Web of Science®Google Scholar Dror , I. E. , Charlton , D. , & Péron , A. E. ( 2006 ). Contextual information renders experts vulnerable to making erroneous identifications . Forensic Science International , 156 , 74 – 78 . 10.1016/j.forsciint.2005.10.017 PubMedWeb of Science®Google Scholar Dror , I. E. , & Hampikian , G. ( 2011 ). Subjectivity and bias in forensic DNA mixture interpretation . Science & Justice , 51 ( 4 ), 204 – 208 . 10.1016/j.scijus.2011.08.004 CASPubMedWeb of Science®Google Scholar Forensic Regulator . ( 2015 ). Cognitive bias effects relevant to forensic science examinations . Home Office . Google Scholar Grant , T. ( 2012 ). TXT 4N6: Method, consistency, and distinctiveness in the analysis of SMS text messages . Journal of Law & Policy , 21 , 467 . Google Scholar Grant , T. ( 2020 ). Text messaging forensics: Txt 4n6: Idiolect free authorship analysis? In M. Coulthard , A. May, & R. Sousa-Silva (eds.), The Routledge handbook of forensic linguistics ( 2nd ed.). Routledge . Google Scholar Grant , T. , & Baker , K. ( 2001 ). Identifying reliable, valid markers of authorship: A response to Chaski . Forensic Linguistics , 8 , 66 – 79 . 10.1558/sll.2001.8.1.66 Google Scholar Grant , T. , & MacLeod , N. ( 2020 ). Language and online identities: The undercover policing of internet sexual crime . Cambridge University Press . 10.1017/9781108766425 Google Scholar Grieve , J. ( 2007 ). Quantitative authorship attribution: An evaluation of techniques . Literary and Linguistic Computing 22 ( 3 ), 251 – 270 . 10.1093/llc/fqm020 Google Scholar Grieve , J. , Clarke , I. , Chiang , E. , Gideon , H. , Heini , A. , Nini , A. , & Waibel , E. ( 2019 ). Attributing the Bixby Letter using n-gram tracing . Digital Scholarship in the Humanities , 34 ( 3 ), 493 – 512 . 10.1093/llc/fqy042 Google Scholar Grieve , J. , & Woodfield , H. ( 2020 ). Investigative linguistics . In M. Coulthard , A. May , & R. Sousa Silva (eds.), The Routledge handbook of forensic linguistics ( 2nd ed .). Routledge . Google Scholar Luyckx , K. , & Daelemans , W. ( 2011 ). The effect of author set size and data size in authorship attribution . Literary and Linguistic Computing , 26 ( 1 ), 35 – 55 . 10.1093/llc/fqq013 Web of Science®Google Scholar Wagner , S. E. ( 2012 ). Age grading in sociolinguistic theory . Language and Linguistics Compass , 6 ( 6 ), 371 – 382 . 10.1002/lnc3.343 Google Scholar Wright , D. ( 2017 ). Using word n-grams to identify authors and idiolects: A corpus approach to a forensic linguistic problem . International Journal of Corpus Linguistics , 22 ( 2 ), 212 – 241 . 10.1075/ijcl.22.2.03wri Google Scholar Methodologies and Challenges in Forensic Linguistic Casework ReferencesRelatedInformation
individual’s set of psychological traits, or their psychopathology,impacts how they experience the world around them, and language offers resourcesthat allow for that experience to be shared with and communicated to someoneelse. That language can then be analyzed for patterns and their connectionsto the psychological traits. In a forensic context, such connections may givevaluable insights. There already exist psychological and linguistic approachesto the analysis of forensic texts, but the psychological approach largely lacksgrounding in linguistic theory and the linguistic approach does not typically allowconsideration of psychological characteristics. What this paper aims to provideis a step toward bridging that gap. In this paper we examine the system ofattitude from the Appraisal framework developed by Martin and White (2005) andadapted by Gales (2010) and Hurt (2020) applying this to the writings of four serialmurderers with documented mental health diagnoses. Significant patterns in theattitudinal resources were identified quantitatively and examined qualitativelythrough the lens of the psychological traits that comprised the authors’ diagnosesto determine if there was a relationship between them. Despite the obviouslimitation presented by the sample size, the results of this study suggest theapproach presented in this paper warrants further investigation.